Alpamayo-R1-10B TensorRT engine OOM on DRIVE AGX Thor

I am deploying nvidia/Alpamayo-R1-10B with TensorRT-Edge-LLM on DRIVE AGX Thor running DriveOS 7.2.5.0.

CUDA: 13.3
TensorRT reported by llm_build: 11.0.1
SDK container: driveos-sdk … 7.2.5.0-0004
Target: auto-thor

The ONNX model parses successfully and AttentionPlugin loads successfully. Engine serialization fails with:

Requested amount of GPU memory (15167621888 bytes) could not be allocated
OutOfMemory

This remains approximately 15.17 GB with maxBatchSize=1, maxInputLen=2048, and maxKVCacheCapacity=2048.

Is TensorRT 11.0.1 supported for Alpamayo-R1-10B on DRIVE AGX Thor with DriveOS 7.2.5? Is the fixed 15.17 GB allocation expected? What exact TensorRT-Edge-LLM commit, SDK container, CMake options, and FP16/INT8 workflow are validated for Thor? Is the intended deployment path a quantized INT8/QDQ Expert engine rather than direct FP16 LLM engine construction?

Is the intended DRIVE AGX Thor deployment path for Alpamayo an INT8/QDQ Expert engine rather than direct FP16 TensorRT-Edge-LLM conversion of the full 10B LLM? The direct FP16 LLM build requires a fixed 15.17 GB allocation and fails during serialization on Thor, even with batch size 1 and KV cache 2048. Please provide the official Thor workflow for SmoothQuant/INT8 Expert conversion, including the supported TensorRT-Edge-LLM release, ONNX export command, engine-build command, and whether the full LLM engine is expected to run on Thor.

Following confirmation i did

Available RAM before build: approximately 40–50 GB
virtual memory: unlimited
max memory size: unlimited
memlock: unlimited

can you please share the details

I recently installed alpamayo1_5, on Thor Dev kit with Jetpack 7.2.1; in a venv using following method. You could test and if not helpful to your use case, delete the venv.

This is amount of memory used when it is running.

free -h
               total        used        free      shared  buff/cache   available
Mem:           122Gi        21Gi        88Gi        51Mi        14Gi       101Gi

# According to jtop alpamayo1_5 used 14.5gb ram

To build alpamayo1_5 on Thor:

1. Install uv:
curl -LsSf https://astral.sh/uv/install.sh | sh

# Create and activate venv
uv venv a1_5_venv -p 3.12.3 --seed
source a1_5_venv/bin/activate

# Install huggingface_hub
uv pip install hf

hf auth login --token TOKEN  # Token generated from https://huggingface.co/settings/tokens

# Download the model:
hf download nvidia/Alpamayo-1.5-10B

# On hf go to the cosmos model and the physical dataset, and click checkboxes to be granted access to the gated repos.
https://huggingface.co/
nvidia/Cosmos-Reason2-8B
nvidia/PhysicalAI-Autonomous-Vehicles
  1. Clone Alpamayo recipes and model repo:
git clone https://github.com/NVlabs/alpamayo-recipes.git
git clone https://github.com/NVlabs/alpamayo1.5.git

# Install (pyproject.toml installs cpu torch/torchvision/torchaudio; we replace them with cuda versions in next block.
cd alpamayo1.5
uv sync --active --no-install-package flash-attn

# Next line substitute your desired recipe.
cd ../alpamayo-recipes/recipes/alpamayo1_5_quant
uv sync --active --no-install-package flash-attn
  1. Install cuda wheels.
uv pip install torch torchvision --index-url https://download.pytorch.org/whl/cu132

git clone -b release/2.11 --depth 1 https://github.com/pytorch/audio.git torchaudio
cd torchaudio
export CUDA_HOME=/usr/local/cuda
export USE_CUDA=1
export TORCH_CUDA_ARCH_LIST="11.0"
export BUILD_VERSION=2.11.0

uv build --wheel --no-build-isolation -v 

uv pip install dist/torchaudio-2.11.0-cp312-cp312-linux_aarch64.whl

  1. Skip flash-attn — use SDPA. torch includes built-in SDPA; it has a flash kernel that works on SM 110.

Uninstall the broken wheel and apply the patch at the bottom:

uv pip uninstall flash-attn

uv pip install triton

Apply the sdpa patch USE diff below

5. Run a recipe:
cd alpamayo-recipes
uv run --active --no-sync recipes/alpamayo1_5_quant/quantize.py \
  --quant_format=auto --auto_quantize_bits=6.5 \
  --num_of_calib_clips=100 --save_model_dir=./outputs


Git diff make alpamayo quantize.py use sdpa. Save as alpamayo15-thor-sdpa.patch and apply to alpamayo-recipes with git apply.

diff --git a/recipes/alpamayo1_5_quant/quantize.py b/recipes/alpamayo1_5_quant/quantize.py
index 8a8f3c7..78dc93b 100644
--- a/recipes/alpamayo1_5_quant/quantize.py
+++ b/recipes/alpamayo1_5_quant/quantize.py
@@ -267,9 +267,9 @@ def main():

     device = "cuda"
     mto.enable_huggingface_checkpointing()
-    model = Alpamayo1_5.from_pretrained(args.ckpt, dtype=torch.float16).to(
-        device=device, dtype=torch.float16
-    )
+    model = Alpamayo1_5.from_pretrained(
+        args.ckpt, dtype=torch.float16, attn_implementation="sdpa"
+    ).to(device=device, dtype=torch.float16)
     model.eval()

     processor = helper.get_processor(model.tokenizer)
---