Dear Nvidia
I refer to the Web
We are currently planning to replace our existing x86-based solution with GB10.
However, we have encountered several technical bottlenecks during the platform migration process. Regarding GB10βs architectural support, could you provide any recommendations or best practices for our reference?
root@855edf0a6149:/workspace# cd TensorRT-Edge-LLM
root@855edf0a6149:/workspace/TensorRT-Edge-LLM# pip3 install .
Processing /workspace/TensorRT-Edge-LLM
DEPRECATION: Setting PIP_CONSTRAINT will not affect build constraints in the future, pip 26.2 will enforce this behaviour change. A possible replacement is to specify build constraints using --build-constraint or PIP_BUILD_CONSTRAINT. To disable this warning without any build constraints set --use-feature=build-constraint or PIP_USE_FEATURE=βbuild-constraintβ.
Installing build dependencies β¦ done
Getting requirements to build wheel β¦ done
Preparing metadata (pyproject.toml) β¦ done
Collecting torch~=2.10.0 (from tensorrt-edgellm==0.6.1)
Downloading torch-2.10.0-cp312-cp312-manylinux_2_28_aarch64.whl.metadata (31 kB)
Collecting transformers==4.57.6 (from tensorrt-edgellm==0.6.1)
Downloading transformers-4.57.6-py3-none-any.whl.metadata (43 kB)
Requirement already satisfied: nvidia-modelopt==0.39.0 in /usr/local/lib/python3.12/dist-packages (from nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (0.39.0)
Collecting onnx==1.19.0 (from tensorrt-edgellm==0.6.1)
Downloading onnx-1.19.0-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.whl.metadata (7.0 kB)
Collecting datasets==4.4.2 (from tensorrt-edgellm==0.6.1)
Downloading datasets-4.4.2-py3-none-any.whl.metadata (19 kB)
Requirement already satisfied: tqdm~=4.67.1 in /usr/local/lib/python3.12/dist-packages (from tensorrt-edgellm==0.6.1) (4.67.1)
Collecting numpy~=2.2.6 (from tensorrt-edgellm==0.6.1)
Downloading numpy-2.2.6-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.metadata (63 kB)
Collecting pillow==12.1.1 (from tensorrt-edgellm==0.6.1)
Downloading pillow-12.1.1-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl.metadata (8.8 kB)
Collecting torchvision==0.25.0 (from tensorrt-edgellm==0.6.1)
Downloading torchvision-0.25.0-cp312-cp312-manylinux_2_28_aarch64.whl.metadata (5.4 kB)
Collecting peft==0.18.1 (from tensorrt-edgellm==0.6.1)
Downloading peft-0.18.1-py3-none-any.whl.metadata (14 kB)
Collecting backoff==2.2.1 (from tensorrt-edgellm==0.6.1)
Downloading backoff-2.2.1-py3-none-any.whl.metadata (14 kB)
Requirement already satisfied: soundfile==0.13.1 in /usr/local/lib/python3.12/dist-packages (from tensorrt-edgellm==0.6.1) (0.13.1)
Requirement already satisfied: librosa==0.11.0 in /usr/local/lib/python3.12/dist-packages (from tensorrt-edgellm==0.6.1) (0.11.0)
Collecting einops==0.8.2 (from tensorrt-edgellm==0.6.1)
Downloading einops-0.8.2-py3-none-any.whl.metadata (13 kB)
Requirement already satisfied: filelock in /usr/local/lib/python3.12/dist-packages (from datasets==4.4.2->tensorrt-edgellm==0.6.1) (3.20.1)
Requirement already satisfied: pyarrow>=21.0.0 in /usr/local/lib/python3.12/dist-packages (from datasets==4.4.2->tensorrt-edgellm==0.6.1) (22.0.0)
Requirement already satisfied: dill<0.4.1,>=0.3.0 in /usr/local/lib/python3.12/dist-packages (from datasets==4.4.2->tensorrt-edgellm==0.6.1) (0.4.0)
Requirement already satisfied: pandas in /usr/local/lib/python3.12/dist-packages (from datasets==4.4.2->tensorrt-edgellm==0.6.1) (2.3.3)
Requirement already satisfied: requests>=2.32.2 in /usr/local/lib/python3.12/dist-packages (from datasets==4.4.2->tensorrt-edgellm==0.6.1) (2.32.5)
Requirement already satisfied: httpx<1.0.0 in /usr/local/lib/python3.12/dist-packages (from datasets==4.4.2->tensorrt-edgellm==0.6.1) (0.28.1)
Requirement already satisfied: xxhash in /usr/local/lib/python3.12/dist-packages (from datasets==4.4.2->tensorrt-edgellm==0.6.1) (3.6.0)
Requirement already satisfied: multiprocess<0.70.19 in /usr/local/lib/python3.12/dist-packages (from datasets==4.4.2->tensorrt-edgellm==0.6.1) (0.70.18)
Requirement already satisfied: fsspec<=2025.10.0,>=2023.1.0 in /usr/local/lib/python3.12/dist-packages (from fsspec[http]<=2025.10.0,>=2023.1.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (2025.10.0)
Requirement already satisfied: huggingface-hub<2.0,>=0.25.0 in /usr/local/lib/python3.12/dist-packages (from datasets==4.4.2->tensorrt-edgellm==0.6.1) (1.2.3)
Requirement already satisfied: packaging in /usr/local/lib/python3.12/dist-packages (from datasets==4.4.2->tensorrt-edgellm==0.6.1) (25.0)
Requirement already satisfied: pyyaml>=5.1 in /usr/local/lib/python3.12/dist-packages (from datasets==4.4.2->tensorrt-edgellm==0.6.1) (6.0.3)
Requirement already satisfied: audioread>=2.1.9 in /usr/local/lib/python3.12/dist-packages (from librosa==0.11.0->tensorrt-edgellm==0.6.1) (3.1.0)
Requirement already satisfied: numba>=0.51.0 in /usr/local/lib/python3.12/dist-packages (from librosa==0.11.0->tensorrt-edgellm==0.6.1) (0.63.1)
Requirement already satisfied: scipy>=1.6.0 in /usr/local/lib/python3.12/dist-packages (from librosa==0.11.0->tensorrt-edgellm==0.6.1) (1.16.3)
Requirement already satisfied: scikit-learn>=1.1.0 in /usr/local/lib/python3.12/dist-packages (from librosa==0.11.0->tensorrt-edgellm==0.6.1) (1.8.0)
Requirement already satisfied: joblib>=1.0 in /usr/local/lib/python3.12/dist-packages (from librosa==0.11.0->tensorrt-edgellm==0.6.1) (1.5.3)
Requirement already satisfied: decorator>=4.3.0 in /usr/local/lib/python3.12/dist-packages (from librosa==0.11.0->tensorrt-edgellm==0.6.1) (5.2.1)
Requirement already satisfied: pooch>=1.1 in /usr/local/lib/python3.12/dist-packages (from librosa==0.11.0->tensorrt-edgellm==0.6.1) (1.8.2)
Requirement already satisfied: soxr>=0.3.2 in /usr/local/lib/python3.12/dist-packages (from librosa==0.11.0->tensorrt-edgellm==0.6.1) (1.0.0)
Requirement already satisfied: typing_extensions>=4.1.1 in /usr/local/lib/python3.12/dist-packages (from librosa==0.11.0->tensorrt-edgellm==0.6.1) (4.15.0)
Requirement already satisfied: lazy_loader>=0.1 in /usr/local/lib/python3.12/dist-packages (from librosa==0.11.0->tensorrt-edgellm==0.6.1) (0.4)
Requirement already satisfied: msgpack>=1.0 in /usr/local/lib/python3.12/dist-packages (from librosa==0.11.0->tensorrt-edgellm==0.6.1) (1.1.2)
Requirement already satisfied: ninja in /usr/local/lib/python3.12/dist-packages (from nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (1.13.0)
Requirement already satisfied: pydantic>=2.0 in /usr/local/lib/python3.12/dist-packages (from nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (2.12.5)
Requirement already satisfied: nvidia-ml-py>=12 in /usr/local/lib/python3.12/dist-packages (from nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (13.590.44)
Requirement already satisfied: rich in /usr/local/lib/python3.12/dist-packages (from nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (14.2.0)
Requirement already satisfied: pulp in /usr/local/lib/python3.12/dist-packages (from nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (3.3.0)
Requirement already satisfied: regex in /usr/local/lib/python3.12/dist-packages (from nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (2025.11.3)
Requirement already satisfied: safetensors in /usr/local/lib/python3.12/dist-packages (from nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (0.7.0)
Requirement already satisfied: torchprofile>=0.0.4 in /usr/local/lib/python3.12/dist-packages (from nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (0.0.4)
Collecting cppimport (from nvidia-modelopt[onnx]==0.39.0->tensorrt-edgellm==0.6.1)
Downloading cppimport-26.4.17.tar.gz (28 kB)
Installing build dependencies β¦ done
Getting requirements to build wheel β¦ done
Preparing metadata (pyproject.toml) β¦ done
Requirement already satisfied: ml_dtypes in /usr/local/lib/python3.12/dist-packages (from nvidia-modelopt[onnx]==0.39.0->tensorrt-edgellm==0.6.1) (0.5.4)
Collecting onnx-graphsurgeon (from nvidia-modelopt[onnx]==0.39.0->tensorrt-edgellm==0.6.1)
Downloading onnx_graphsurgeon-0.6.1-py2.py3-none-any.whl.metadata (8.2 kB)
Collecting onnxconverter-common~=1.16.0 (from nvidia-modelopt[onnx]==0.39.0->tensorrt-edgellm==0.6.1)
Downloading onnxconverter_common-1.16.0-py2.py3-none-any.whl.metadata (4.8 kB)
Collecting onnxruntime~=1.22.0 (from nvidia-modelopt[onnx]==0.39.0->tensorrt-edgellm==0.6.1)
Downloading onnxruntime-1.22.1-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl.metadata (4.9 kB)
Requirement already satisfied: onnxscript in /usr/local/lib/python3.12/dist-packages (from nvidia-modelopt[onnx]==0.39.0->tensorrt-edgellm==0.6.1) (0.5.7)
Requirement already satisfied: polygraphy>=0.49.22 in /usr/local/lib/python3.12/dist-packages (from nvidia-modelopt[onnx]==0.39.0->tensorrt-edgellm==0.6.1) (0.49.26)
Requirement already satisfied: protobuf>=4.25.1 in /usr/local/lib/python3.12/dist-packages (from onnx==1.19.0->tensorrt-edgellm==0.6.1) (6.33.2)
WARNING: nvidia-modelopt 0.39.0 does not provide the extra βtorchβ
Requirement already satisfied: psutil in /usr/local/lib/python3.12/dist-packages (from peft==0.18.1->tensorrt-edgellm==0.6.1) (7.1.3)
Collecting accelerate>=0.21.0 (from peft==0.18.1->tensorrt-edgellm==0.6.1)
Downloading accelerate-1.13.0-py3-none-any.whl.metadata (19 kB)
Requirement already satisfied: cffi>=1.0 in /usr/local/lib/python3.12/dist-packages (from soundfile==0.13.1->tensorrt-edgellm==0.6.1) (2.0.0)
Requirement already satisfied: setuptools in /usr/local/lib/python3.12/dist-packages (from torch~=2.10.0->tensorrt-edgellm==0.6.1) (80.9.0)
Requirement already satisfied: sympy>=1.13.3 in /usr/local/lib/python3.12/dist-packages (from torch~=2.10.0->tensorrt-edgellm==0.6.1) (1.14.0)
Requirement already satisfied: networkx>=2.5.1 in /usr/local/lib/python3.12/dist-packages (from torch~=2.10.0->tensorrt-edgellm==0.6.1) (3.6.1)
Requirement already satisfied: jinja2 in /usr/local/lib/python3.12/dist-packages (from torch~=2.10.0->tensorrt-edgellm==0.6.1) (3.1.6)
Collecting huggingface-hub<2.0,>=0.25.0 (from datasets==4.4.2->tensorrt-edgellm==0.6.1)
Downloading huggingface_hub-0.36.2-py3-none-any.whl.metadata (15 kB)
Requirement already satisfied: tokenizers<=0.23.0,>=0.22.0 in /usr/local/lib/python3.12/dist-packages (from transformers==4.57.6->tensorrt-edgellm==0.6.1) (0.22.1)
Requirement already satisfied: aiohttp!=4.0.0a0,!=4.0.0a1 in /usr/local/lib/python3.12/dist-packages (from fsspec[http]<=2025.10.0,>=2023.1.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (3.13.2)
Requirement already satisfied: anyio in /usr/local/lib/python3.12/dist-packages (from httpx<1.0.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (4.12.0)
Requirement already satisfied: certifi in /usr/local/lib/python3.12/dist-packages (from httpx<1.0.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (2025.11.12)
Requirement already satisfied: httpcore==1.* in /usr/local/lib/python3.12/dist-packages (from httpx<1.0.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (1.0.9)
Requirement already satisfied: idna in /usr/local/lib/python3.12/dist-packages (from httpx<1.0.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (3.11)
Requirement already satisfied: h11>=0.16 in /usr/local/lib/python3.12/dist-packages (from httpcore==1.*->httpx<1.0.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (0.16.0)
Requirement already satisfied: hf-xet<2.0.0,>=1.1.3 in /usr/local/lib/python3.12/dist-packages (from huggingface-hub<2.0,>=0.25.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (1.2.0)
Collecting coloredlogs (from onnxruntime~=1.22.0->nvidia-modelopt[onnx]==0.39.0->tensorrt-edgellm==0.6.1)
Downloading coloredlogs-15.0.1-py2.py3-none-any.whl.metadata (12 kB)
Collecting flatbuffers (from onnxruntime~=1.22.0->nvidia-modelopt[onnx]==0.39.0->tensorrt-edgellm==0.6.1)
Downloading flatbuffers-25.12.19-py2.py3-none-any.whl.metadata (1.0 kB)
Requirement already satisfied: aiohappyeyeballs>=2.5.0 in /usr/local/lib/python3.12/dist-packages (from aiohttp!=4.0.0a0,!=4.0.0a1->fsspec[http]<=2025.10.0,>=2023.1.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (2.6.1)
Requirement already satisfied: aiosignal>=1.4.0 in /usr/local/lib/python3.12/dist-packages (from aiohttp!=4.0.0a0,!=4.0.0a1->fsspec[http]<=2025.10.0,>=2023.1.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (1.4.0)
Requirement already satisfied: attrs>=17.3.0 in /usr/local/lib/python3.12/dist-packages (from aiohttp!=4.0.0a0,!=4.0.0a1->fsspec[http]<=2025.10.0,>=2023.1.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (25.4.0)
Requirement already satisfied: frozenlist>=1.1.1 in /usr/local/lib/python3.12/dist-packages (from aiohttp!=4.0.0a0,!=4.0.0a1->fsspec[http]<=2025.10.0,>=2023.1.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (1.8.0)
Requirement already satisfied: multidict<7.0,>=4.5 in /usr/local/lib/python3.12/dist-packages (from aiohttp!=4.0.0a0,!=4.0.0a1->fsspec[http]<=2025.10.0,>=2023.1.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (6.7.0)
Requirement already satisfied: propcache>=0.2.0 in /usr/local/lib/python3.12/dist-packages (from aiohttp!=4.0.0a0,!=4.0.0a1->fsspec[http]<=2025.10.0,>=2023.1.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (0.4.1)
Requirement already satisfied: yarl<2.0,>=1.17.0 in /usr/local/lib/python3.12/dist-packages (from aiohttp!=4.0.0a0,!=4.0.0a1->fsspec[http]<=2025.10.0,>=2023.1.0->datasets==4.4.2->tensorrt-edgellm==0.6.1) (1.22.0)
Requirement already satisfied: pycparser in /usr/local/lib/python3.12/dist-packages (from cffi>=1.0->soundfile==0.13.1->tensorrt-edgellm==0.6.1) (2.23)
Requirement already satisfied: llvmlite<0.47,>=0.46.0dev0 in /usr/local/lib/python3.12/dist-packages (from numba>=0.51.0->librosa==0.11.0->tensorrt-edgellm==0.6.1) (0.46.0)
Requirement already satisfied: platformdirs>=2.5.0 in /usr/local/lib/python3.12/dist-packages (from pooch>=1.1->librosa==0.11.0->tensorrt-edgellm==0.6.1) (4.5.1)
Requirement already satisfied: annotated-types>=0.6.0 in /usr/local/lib/python3.12/dist-packages (from pydantic>=2.0->nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (0.7.0)
Requirement already satisfied: pydantic-core==2.41.5 in /usr/local/lib/python3.12/dist-packages (from pydantic>=2.0->nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (2.41.5)
Requirement already satisfied: typing-inspection>=0.4.2 in /usr/local/lib/python3.12/dist-packages (from pydantic>=2.0->nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (0.4.2)
Requirement already satisfied: charset_normalizer<4,>=2 in /usr/local/lib/python3.12/dist-packages (from requests>=2.32.2->datasets==4.4.2->tensorrt-edgellm==0.6.1) (3.4.4)
Requirement already satisfied: urllib3<3,>=1.21.1 in /usr/local/lib/python3.12/dist-packages (from requests>=2.32.2->datasets==4.4.2->tensorrt-edgellm==0.6.1) (2.6.1)
Requirement already satisfied: threadpoolctl>=3.2.0 in /usr/local/lib/python3.12/dist-packages (from scikit-learn>=1.1.0->librosa==0.11.0->tensorrt-edgellm==0.6.1) (3.6.0)
Requirement already satisfied: mpmath<1.4,>=1.1.0 in /usr/local/lib/python3.12/dist-packages (from sympy>=1.13.3->torch~=2.10.0->tensorrt-edgellm==0.6.1) (1.3.0)
Collecting humanfriendly>=9.1 (from coloredlogs->onnxruntime~=1.22.0->nvidia-modelopt[onnx]==0.39.0->tensorrt-edgellm==0.6.1)
Downloading humanfriendly-10.0-py2.py3-none-any.whl.metadata (9.2 kB)
Collecting mako (from cppimport->nvidia-modelopt[onnx]==0.39.0->tensorrt-edgellm==0.6.1)
Downloading mako-1.3.11-py3-none-any.whl.metadata (2.9 kB)
Requirement already satisfied: pybind11 in /usr/local/lib/python3.12/dist-packages (from cppimport->nvidia-modelopt[onnx]==0.39.0->tensorrt-edgellm==0.6.1) (3.0.1)
Requirement already satisfied: MarkupSafe>=2.0 in /usr/local/lib/python3.12/dist-packages (from jinja2->torch~=2.10.0->tensorrt-edgellm==0.6.1) (3.0.3)
Requirement already satisfied: onnx_ir<2,>=0.1.12 in /usr/local/lib/python3.12/dist-packages (from onnxscript->nvidia-modelopt[onnx]==0.39.0->tensorrt-edgellm==0.6.1) (0.1.12)
Requirement already satisfied: python-dateutil>=2.8.2 in /usr/local/lib/python3.12/dist-packages (from pandas->datasets==4.4.2->tensorrt-edgellm==0.6.1) (2.9.0.post0)
Requirement already satisfied: pytz>=2020.1 in /usr/local/lib/python3.12/dist-packages (from pandas->datasets==4.4.2->tensorrt-edgellm==0.6.1) (2025.2)
Requirement already satisfied: tzdata>=2022.7 in /usr/local/lib/python3.12/dist-packages (from pandas->datasets==4.4.2->tensorrt-edgellm==0.6.1) (2025.3)
Requirement already satisfied: six>=1.5 in /usr/local/lib/python3.12/dist-packages (from python-dateutil>=2.8.2->pandas->datasets==4.4.2->tensorrt-edgellm==0.6.1) (1.16.0)
Requirement already satisfied: markdown-it-py>=2.2.0 in /usr/local/lib/python3.12/dist-packages (from rich->nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (4.0.0)
Requirement already satisfied: pygments<3.0.0,>=2.13.0 in /usr/local/lib/python3.12/dist-packages (from rich->nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (2.19.2)
Requirement already satisfied: mdurl~=0.1 in /usr/local/lib/python3.12/dist-packages (from markdown-it-py>=2.2.0->rich->nvidia-modelopt==0.39.0->nvidia-modelopt[torch]==0.39.0->tensorrt-edgellm==0.6.1) (0.1.2)
Downloading backoff-2.2.1-py3-none-any.whl (15 kB)
Downloading datasets-4.4.2-py3-none-any.whl (512 kB)
Downloading einops-0.8.2-py3-none-any.whl (65 kB)
Downloading onnx-1.19.0-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.whl (18.0 MB)
ββββββββββββββββββββββββββββββββββββββββ 18.0/18.0 MB 11.4 MB/s 0:00:01
Downloading peft-0.18.1-py3-none-any.whl (556 kB)
ββββββββββββββββββββββββββββββββββββββββ 557.0/557.0 kB 10.7 MB/s 0:00:00
Downloading pillow-12.1.1-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl (6.3 MB)
ββββββββββββββββββββββββββββββββββββββββ 6.3/6.3 MB 11.5 MB/s 0:00:00
Downloading torchvision-0.25.0-cp312-cp312-manylinux_2_28_aarch64.whl (2.3 MB)
ββββββββββββββββββββββββββββββββββββββββ 2.3/2.3 MB 11.1 MB/s 0:00:00
Downloading torch-2.10.0-cp312-cp312-manylinux_2_28_aarch64.whl (146.0 MB)
ββββββββββββββββββββββββββββββββββββββββ 146.0/146.0 MB 11.5 MB/s 0:00:12
Downloading transformers-4.57.6-py3-none-any.whl (12.0 MB)
ββββββββββββββββββββββββββββββββββββββββ 12.0/12.0 MB 10.5 MB/s 0:00:01
Downloading huggingface_hub-0.36.2-py3-none-any.whl (566 kB)
ββββββββββββββββββββββββββββββββββββββββ 566.4/566.4 kB 9.2 MB/s 0:00:00
Downloading numpy-2.2.6-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (14.0 MB)
ββββββββββββββββββββββββββββββββββββββββ 14.0/14.0 MB 11.3 MB/s 0:00:01
Downloading onnxconverter_common-1.16.0-py2.py3-none-any.whl (89 kB)
Downloading onnxruntime-1.22.1-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl (14.5 MB)
ββββββββββββββββββββββββββββββββββββββββ 14.5/14.5 MB 11.3 MB/s 0:00:01
Downloading accelerate-1.13.0-py3-none-any.whl (383 kB)
Downloading coloredlogs-15.0.1-py2.py3-none-any.whl (46 kB)
Downloading humanfriendly-10.0-py2.py3-none-any.whl (86 kB)
Downloading flatbuffers-25.12.19-py2.py3-none-any.whl (26 kB)
Downloading mako-1.3.11-py3-none-any.whl (78 kB)
Downloading onnx_graphsurgeon-0.6.1-py2.py3-none-any.whl (59 kB)
Building wheels for collected packages: tensorrt-edgellm, cppimport
Building wheel for tensorrt-edgellm (pyproject.toml) β¦ done
Created wheel for tensorrt-edgellm: filename=tensorrt_edgellm-0.6.1-py3-none-any.whl size=186258 sha256=9a74c92ceeb7b76eb9a6cd2678474576d9f8bcff7ce4d1d483097eda5636a4a6
Stored in directory: /root/.cache/pip/wheels/ee/2a/a7/c3dde2da01476e5aee2aa0767630d005a6fb8d33ee59ab2e8f
Building wheel for cppimport (pyproject.toml) β¦ done
Created wheel for cppimport: filename=cppimport-26.4.17-py3-none-any.whl size=18786 sha256=be84e26072d976d1d88c24c565c16a7f1b3a71a6ff6f2655576a0046bb12af91
Stored in directory: /root/.cache/pip/wheels/61/04/ef/0e7d3685e15c25df1f461860424fb12c439c0b93007b3e6ae3
Successfully built tensorrt-edgellm cppimport
Installing collected packages: flatbuffers, pillow, numpy, mako, humanfriendly, einops, backoff, torch, huggingface-hub, cppimport, coloredlogs, torchvision, onnxruntime, onnx, accelerate, transformers, onnxconverter-common, onnx-graphsurgeon, datasets, peft, tensorrt-edgellm
Attempting uninstall: pillow
Found existing installation: pillow 12.0.0
Uninstalling pillow-12.0.0:
Successfully uninstalled pillow-12.0.0
Attempting uninstall: numpy
Found existing installation: numpy 2.1.0
Uninstalling numpy-2.1.0:
Successfully uninstalled numpy-2.1.0
Attempting uninstall: einops
Found existing installation: einops 0.8.1
Uninstalling einops-0.8.1:
Successfully uninstalled einops-0.8.1
Attempting uninstall: torch
Found existing installation: torch 2.10.0a0+b4e4ee81d3.nv25.12
Uninstalling torch-2.10.0a0+b4e4ee81d3.nv25.12:
Successfully uninstalled torch-2.10.0a0+b4e4ee81d3.nv25.12
Attempting uninstall: huggingface-hub
Found existing installation: huggingface_hub 1.2.3
Uninstalling huggingface_hub-1.2.3:
Successfully uninstalled huggingface_hub-1.2.3
Attempting uninstall: torchvision
Found existing installation: torchvision 0.25.0a0+ca221243
Uninstalling torchvision-0.25.0a0+ca221243:
Successfully uninstalled torchvision-0.25.0a0+ca221243
Attempting uninstall: onnx
Found existing installation: onnx 1.18.0
Uninstalling onnx-1.18.0:
Successfully uninstalled onnx-1.18.0
Attempting uninstall: datasets
Found existing installation: datasets 4.4.1
Uninstalling datasets-4.4.1:
Successfully uninstalled datasets-4.4.1
Successfully installed accelerate-1.13.0 backoff-2.2.1 coloredlogs-15.0.1 cppimport-26.4.17 datasets-4.4.2 einops-0.8.2 flatbuffers-25.12.19 huggingface-hub-0.36.2 humanfriendly-10.0 mako-1.3.11 numpy-2.2.6 onnx-1.19.0 onnx-graphsurgeon-0.6.1 onnxconverter-common-1.16.0 onnxruntime-1.22.1 peft-0.18.1 pillow-12.1.1 tensorrt-edgellm-0.6.1 torch-2.10.0 torchvision-0.25.0 transformers-4.57.6
WARNING: Running pip as the βrootβ user can result in broken permissions and conflicting behaviour with the system package manager, possibly rendering your system unusable. It is recommended to use a virtual environment instead: 12. Virtual Environments and Packages β Python 3.14.4 documentation . Use the --root-user-action option if you know what you are doing and want to suppress this warning.
root@855edf0a6149:/workspace/TensorRT-Edge-LLM# # Set up workspace directory
export WORKSPACE_DIR=$HOME/tensorrt-edgellm-workspace
export MODEL_NAME=Qwen3-0.6B
mkdir -p $WORKSPACE_DIR
cd $WORKSPACE_DIR
Step 1: Quantize to FP8 (downloads model automatically)
tensorrt-edgellm-quantize-llm
βmodel_dir Qwen/Qwen3-0.6B
βoutput_dir $MODEL_NAME/quantized
βquantization fp8
Step 2: Export to ONNX
tensorrt-edgellm-export-llm
βmodel_dir $MODEL_NAME/quantized
βoutput_dir $MODEL_NAME/onnx
/usr/local/lib/python3.12/dist-packages/modelopt/torch/utils/import_utils.py:32: UserWarning: Failed to import transformer engine plugin due to: AttributeError(βmodule βtransformer_engineβ has no attribute βpytorchββ). You may ignore this warning if you do not need this plugin.
warnings.warn(
/usr/local/lib/python3.12/dist-packages/modelopt/torch/utils/import_utils.py:32: UserWarning: Failed to import transformer_engine plugin due to: ImportError(β/usr/local/lib/python3.12/dist-packages/transformer_engine/transformer_engine_torch.cpython-312-aarch64-linux-gnu.so: undefined symbol: _ZN3c104cuda20CUDACachingAllocator9allocatorEβ). You may ignore this warning if you do not need this plugin.
warnings.warn(
ModelOpt save/restore enabled for transformers library.
ModelOpt save/restore enabled for peft library.
ModelOpt save/restore enabled for transformers library.
ModelOpt save/restore enabled for peft library.
tokenizer_config.json: 9.73kB [00:00, 86.1MB/s]
vocab.json: 2.78MB [00:00, 29.4MB/s]
merges.txt: 1.67MB [00:00, 28.1MB/s]
tokenizer.json: 100%|βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 11.4M/11.4M [00:01<00:00, 8.76MB/s]
config.json: 100%|ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 726/726 [00:00<00:00, 15.9MB/s]
torch_dtype is deprecated! Use dtype instead!
model.safetensors: 100%|ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 1.50G/1.50G [01:29<00:00, 16.8MB/s]
generation_config.json: 100%|βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 239/239 [00:00<00:00, 5.28MB/s]
AutoModelForCausalLM failed: Torch not compiled with CUDA enabled
Error during model quantization: Could not load model from Qwen/Qwen3-0.6B. Error: Unrecognized configuration class <class βtransformers.models.qwen3.configuration_qwen3.Qwen3Configβ> for this kind of AutoModel: AutoModelForImageTextToText.
Model type should be one of AriaConfig, AyaVisionConfig, BlipConfig, Blip2Config, ChameleonConfig, Cohere2VisionConfig, DeepseekVLConfig, DeepseekVLHybridConfig, Emu3Config, EvollaConfig, Florence2Config, FuyuConfig, Gemma3Config, Gemma3nConfig, GitConfig, Glm4vConfig, Glm4vMoeConfig, GotOcr2Config, IdeficsConfig, Idefics2Config, Idefics3Config, InstructBlipConfig, InternVLConfig, JanusConfig, Kosmos2Config, Kosmos2_5Config, Lfm2VlConfig, Llama4Config, LlavaConfig, LlavaNextConfig, LlavaNextVideoConfig, LlavaOnevisionConfig, Mistral3Config, MllamaConfig, Ovis2Config, PaliGemmaConfig, PerceptionLMConfig, Pix2StructConfig, PixtralVisionConfig, Qwen2_5_VLConfig, Qwen2VLConfig, Qwen3VLConfig, Qwen3VLMoeConfig, ShieldGemma2Config, SmolVLMConfig, UdopConfig, VipLlavaConfig, VisionEncoderDecoderConfig.
Traceback:
Traceback (most recent call last):
File β/usr/local/lib/python3.12/dist-packages/tensorrt_edgellm/llm_models/model_utils.pyβ, line 440, in load_hf_model
trust_remote_code=True).to(device)
^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/transformers/modeling_utils.pyβ, line 4343, in to
return super().to(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.pyβ, line 1381, in to
return self._apply(convert)
^^^^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.pyβ, line 933, in _apply
module._apply(fn)
File β/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.pyβ, line 933, in _apply
module._apply(fn)
File β/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.pyβ, line 964, in _apply
param_applied = fn(param)
^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.pyβ, line 1367, in convert
return t.to(
^^^^^
File β/usr/local/lib/python3.12/dist-packages/torch/cuda/init.pyβ, line 417, in _lazy_init
raise AssertionError(βTorch not compiled with CUDA enabledβ)
AssertionError: Torch not compiled with CUDA enabled
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File β/usr/local/lib/python3.12/dist-packages/tensorrt_edgellm/llm_models/model_utils.pyβ, line 447, in load_hf_model
model = AutoModelForImageTextToText.from_pretrained(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/transformers/models/auto/auto_factory.pyβ, line 607, in from_pretrained
raise ValueError(
ValueError: Unrecognized configuration class <class βtransformers.models.qwen3.configuration_qwen3.Qwen3Configβ> for this kind of AutoModel: AutoModelForImageTextToText.
Model type should be one of AriaConfig, AyaVisionConfig, BlipConfig, Blip2Config, ChameleonConfig, Cohere2VisionConfig, DeepseekVLConfig, DeepseekVLHybridConfig, Emu3Config, EvollaConfig, Florence2Config, FuyuConfig, Gemma3Config, Gemma3nConfig, GitConfig, Glm4vConfig, Glm4vMoeConfig, GotOcr2Config, IdeficsConfig, Idefics2Config, Idefics3Config, InstructBlipConfig, InternVLConfig, JanusConfig, Kosmos2Config, Kosmos2_5Config, Lfm2VlConfig, Llama4Config, LlavaConfig, LlavaNextConfig, LlavaNextVideoConfig, LlavaOnevisionConfig, Mistral3Config, MllamaConfig, Ovis2Config, PaliGemmaConfig, PerceptionLMConfig, Pix2StructConfig, PixtralVisionConfig, Qwen2_5_VLConfig, Qwen2VLConfig, Qwen3VLConfig, Qwen3VLMoeConfig, ShieldGemma2Config, SmolVLMConfig, UdopConfig, VipLlavaConfig, VisionEncoderDecoderConfig.
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File β/usr/local/lib/python3.12/dist-packages/tensorrt_edgellm/scripts/quantize_llm.pyβ, line 100, in main
quantize_and_save_llm(model_dir=args.model_dir,
File β/usr/local/lib/python3.12/dist-packages/tensorrt_edgellm/quantization/llm_quantization.pyβ, line 426, in quantize_and_save_llm
model, tokenizer, processor = load_hf_model(model_dir, dtype, device)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/tensorrt_edgellm/llm_models/model_utils.pyβ, line 451, in load_hf_model
raise ValueError(
ValueError: Could not load model from Qwen/Qwen3-0.6B. Error: Unrecognized configuration class <class βtransformers.models.qwen3.configuration_qwen3.Qwen3Configβ> for this kind of AutoModel: AutoModelForImageTextToText.
Model type should be one of AriaConfig, AyaVisionConfig, BlipConfig, Blip2Config, ChameleonConfig, Cohere2VisionConfig, DeepseekVLConfig, DeepseekVLHybridConfig, Emu3Config, EvollaConfig, Florence2Config, FuyuConfig, Gemma3Config, Gemma3nConfig, GitConfig, Glm4vConfig, Glm4vMoeConfig, GotOcr2Config, IdeficsConfig, Idefics2Config, Idefics3Config, InstructBlipConfig, InternVLConfig, JanusConfig, Kosmos2Config, Kosmos2_5Config, Lfm2VlConfig, Llama4Config, LlavaConfig, LlavaNextConfig, LlavaNextVideoConfig, LlavaOnevisionConfig, Mistral3Config, MllamaConfig, Ovis2Config, PaliGemmaConfig, PerceptionLMConfig, Pix2StructConfig, PixtralVisionConfig, Qwen2_5_VLConfig, Qwen2VLConfig, Qwen3VLConfig, Qwen3VLMoeConfig, ShieldGemma2Config, SmolVLMConfig, UdopConfig, VipLlavaConfig, VisionEncoderDecoderConfig.
/usr/local/lib/python3.12/dist-packages/modelopt/torch/utils/import_utils.py:32: UserWarning: Failed to import transformer engine plugin due to: AttributeError(βmodule βtransformer_engineβ has no attribute βpytorchββ). You may ignore this warning if you do not need this plugin.
warnings.warn(
/usr/local/lib/python3.12/dist-packages/modelopt/torch/utils/import_utils.py:32: UserWarning: Failed to import transformer_engine plugin due to: ImportError(β/usr/local/lib/python3.12/dist-packages/transformer_engine/transformer_engine_torch.cpython-312-aarch64-linux-gnu.so: undefined symbol: _ZN3c104cuda20CUDACachingAllocator9allocatorEβ). You may ignore this warning if you do not need this plugin.
warnings.warn(
ModelOpt save/restore enabled for transformers library.
ModelOpt save/restore enabled for peft library.
ModelOpt save/restore enabled for transformers library.
ModelOpt save/restore enabled for peft library.
Exporting standard model to ONNX format
Loading standard model from Qwen3-0.6B/quantized
Error during LLM model export: Qwen3-0.6B/quantized is not a local folder and is not a valid model identifier listed on βModels β Hugging Faceβ
If this is a private repository, make sure to pass a token having permission to this repo either by logging in with hf auth login or by passing token=<your_token>
Traceback:
Traceback (most recent call last):
File β/usr/local/lib/python3.12/dist-packages/huggingface_hub/utils/_http.pyβ, line 403, in hf_raise_for_status
response.raise_for_status()
File β/usr/local/lib/python3.12/dist-packages/requests/models.pyβ, line 1026, in raise_for_status
raise HTTPError(http_error_msg, response=self)
requests.exceptions.HTTPError: 401 Client Error: Unauthorized for url: https://huggingface.co/Qwen3-0.6B/quantized/resolve/main/tokenizer_config.json
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File β/usr/local/lib/python3.12/dist-packages/transformers/utils/hub.pyβ, line 479, in cached_files
hf_hub_download(
File β/usr/local/lib/python3.12/dist-packages/huggingface_hub/utils/_validators.pyβ, line 114, in _inner_fn
return fn(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/huggingface_hub/file_download.pyβ, line 1014, in hf_hub_download
return _hf_hub_download_to_cache_dir(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/huggingface_hub/file_download.pyβ, line 1121, in _hf_hub_download_to_cache_dir
_raise_on_head_call_error(head_call_error, force_download, local_files_only)
File β/usr/local/lib/python3.12/dist-packages/huggingface_hub/file_download.pyβ, line 1662, in _raise_on_head_call_error
raise head_call_error
File β/usr/local/lib/python3.12/dist-packages/huggingface_hub/file_download.pyβ, line 1550, in _get_metadata_or_catch_error
metadata = get_hf_file_metadata(
^^^^^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/huggingface_hub/utils/_validators.pyβ, line 114, in _inner_fn
return fn(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/huggingface_hub/file_download.pyβ, line 1467, in get_hf_file_metadata
r = _request_wrapper(
^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/huggingface_hub/file_download.pyβ, line 283, in _request_wrapper
response = _request_wrapper(
^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/huggingface_hub/file_download.pyβ, line 307, in _request_wrapper
hf_raise_for_status(response)
File β/usr/local/lib/python3.12/dist-packages/huggingface_hub/utils/_http.pyβ, line 453, in hf_raise_for_status
raise _format(RepositoryNotFoundError, message, response) from e
huggingface_hub.errors.RepositoryNotFoundError: 401 Client Error. (Request ID: Root=1-69f0b6bf-6f41776a2722131852237357;19222581-c3b7-4a62-b6d1-9f7968a99633)
Repository Not Found for url: https://huggingface.co/Qwen3-0.6B/quantized/resolve/main/tokenizer_config.json.
Please make sure you specified the correct repo_id and repo_type.
If you are trying to access a private or gated repo, make sure you are authenticated. For more details, see Quickstart Β· Hugging Face
Invalid username or password.
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File β/usr/local/lib/python3.12/dist-packages/tensorrt_edgellm/scripts/export_llm.pyβ, line 108, in main
export_llm_model(model_dir=args.model_dir,
File β/usr/local/lib/python3.12/dist-packages/tensorrt_edgellm/onnx_export/llm_export.pyβ, line 905, in export_llm_model
models_or_dict, tokenizer, processor = load_llm_model(
^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/tensorrt_edgellm/llm_models/model_utils.pyβ, line 515, in load_llm_model
model, tokenizer, processor = load_hf_model(model_dir, dtype, device)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/tensorrt_edgellm/llm_models/model_utils.pyβ, line 379, in load_hf_model
tokenizer = AutoTokenizer.from_pretrained(model_dir,
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/transformers/models/auto/tokenization_auto.pyβ, line 1089, in from_pretrained
tokenizer_config = get_tokenizer_config(pretrained_model_name_or_path, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/transformers/models/auto/tokenization_auto.pyβ, line 921, in get_tokenizer_config
resolved_config_file = cached_file(
^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/transformers/utils/hub.pyβ, line 322, in cached_file
file = cached_files(path_or_repo_id=path_or_repo_id, filenames=[filename], **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File β/usr/local/lib/python3.12/dist-packages/transformers/utils/hub.pyβ, line 511, in cached_files
raise OSError(
OSError: Qwen3-0.6B/quantized is not a local folder and is not a valid model identifier listed on βModels β Hugging Faceβ
If this is a private repository, make sure to pass a token having permission to this repo either by logging in with hf auth login or by passing token=<your_token>
root@855edf0a6149:~/tensorrt-edgellm-workspace#