Drive AGX thor with Alpamayo

DRIVE OS Version: Provide DRIVE OS version. Example: 7.2.5

I am deploying nvidia/Alpamayo-R1-10B with TensorRT-Edge-LLM on DRIVE AGX Thor running DriveOS 7.2.5.0.

CUDA: 13.3
TensorRT reported by llm_build: 11.0.1
SDK container: driveos-sdk … 7.2.5.0-0004
Target: auto-thor

The ONNX model parses successfully and AttentionPlugin loads successfully. Engine serialization fails with:

Requested amount of GPU memory (15167621888 bytes) could not be allocated
OutOfMemory

This remains approximately 15.17 GB with maxBatchSize=1, maxInputLen=2048, and maxKVCacheCapacity=2048.

  1. Is TensorRT 11.0.1 supported for Alpamayo-R1-10B on DRIVE AGX Thor with DriveOS 7.2.5, or is TensorRT 10 required?

  2. Is a fixed ~15.17 GB allocation during LLM engine serialization expected for this model on Thor?

  3. What exact combination is validated for Thor:

    • TensorRT-Edge-LLM release/commit
    • DriveOS SDK container
    • CMake options
    • TensorRT version
  4. Is the intended DRIVE AGX Thor deployment path a quantized INT8/QDQ Expert engine rather than direct FP16 TensorRT-Edge-LLM conversion of the full 10B LLM?

  5. If INT8/QDQ is the supported path, please provide the official Thor workflow, including:

    • the ONNX export command
    • the engine-build command
    • the required TensorRT-Edge-LLM release
  6. Is the full FP16 LLM engine expected to build and run on Thor at all, or only on larger discrete GPUs?

  7. Which smaller model can I build with the same TensorRT-Edge-LLM runtime to independently validate my Thor TensorRT setup?

Available RAM before build: approximately 40–50 GB
virtual memory: unlimited
max memory size: unlimited
memlock: unlimited

can you please clarify why the deployment is showing memory allocation issue irrespective of space is there ?

below are the tegrastats

@SivaRamaKrishnaNV : i am following the steps as mentioned Alpamayo-R1-10B (Vision-Language-Action) — TensorRT Edge-LLM

but still, it is failing can you please support, where it is a bug or how to build ?

Dear @nandhini.ravikumar1 ,
Did you test with TensorRT Edge LLM 0.9.0 as mentioned in TensorRT Edge-LLM — NVIDIA DriveOS 7.2.5 Linux SDK Early Access Developer Guide ?

Hello @SivaRamaKrishnaNV ,

Thank you. I reviewed the TensorRT Edge-LLM 0.9.0 installation guide. My current x86 host does not have an NVIDIA GPU, while the guide specifies an x86 Linux host with an NVIDIA GPU for tensorrt_edgellm.

The Alpamayo FP16 export completed successfully on my CPU-only host and produced onnx/llm, onnx/visual, and onnx/action, but I understand that this is outside the documented configuration.

Could you please confirm whether an NVIDIA GPU is mandatory for the non-quantized FP16 Alpamayo export, or whether CPU-only export is supported?

If a GPU is mandatory, is there a supported workflow to perform the export directly on DRIVE AGX Thor or through another NVIDIA-provided environment?

I currently used latest TensorRT Edge-LLM(0.10) but i will I will repeat the ONNX export, DRIVE OS SDK cross-build, and Thor engine build using the supported 0.9.0 release but can you please whether my HOST PC without GPU will impact?

I expect it to work. GPU is needed to quantize the model or run inference on the host.

Hello @SivaRamaKrishnaNV

Thanks ,

I have now retest with TensorRT Edge LLM 0.9.0 as mentioned in TensorRT Edge-LLM — NVIDIA DriveOS 7.2.5 Linux SDK Early Access Developer Guide

[12:14:54.790] [ERROR] [TensorRT] [resizingAllocator.cpp::allocateImpl::100] Error Code 1: Cuda Runtime (In allocateImpl at /_src/optimizer/builder/resizingAllocator.cpp:100)
[12:14:54.790] [WARNING] [TensorRT] Requested amount of GPU memory (15167679364 bytes) could not be allocated. There may not be enough free memory for allocation to succeed.

[12:14:55.610] [ERROR] [TensorRT] [myelinResourceManager.cpp::allocate::209] Error Code 2: OutOfMemory (Requested size was 15167679364 bytes.)
[12:14:55.610] [ERROR] [builderUtils.cpp:307:buildAndSerializeEngine] Failed to build serialized engine
[12:14:56.731] [ERROR] [llm_build.cpp:243:main] Failed to build LLM engine.

can you please provide any pointers ?

Thank you — confirmed CPU-only export works. I retested the full pipeline on 0.9.0 and the LLM engine build still fails with the same ~14.13 GB allocation (myelinResourceManager … OutOfMemory, 15167679364 bytes). The size is essentially unchanged across maxBatchSize/maxInputLen/maxKVCacheCapacity, so it appears to be a fixed Myelin working buffer, not KV/batch-scaled.

  1. Is a ~14.13 GB allocation expected for Alpamayo-R1-10B FP16 LLM engine build on Thor?
  2. Is the INT8/QDQ (quantized) action-expert / LLM path the intended Thor deployment rather than full FP16? If so, please share the exact export + build commands.
  3. Does increasing nr_hugepages (per the 0.6.x DriveOS large-model note) apply here?
  4. Is TensorRT 11.0.1 (reported by llm_build) the expected version for 0.9.0 on DRIVE OS 7.2.5, or is TensorRT 10 required?

@SivaRamaKrishnaNV :still there is issue with v0.9 as per installation guide . can you support

I am checking on this internally and update you.

Hello @SivaRamaKrishnaNV , I looked into few more details please find my observation below

Hardware: DRIVE AGX Thor developer kit
DriveOS: 7.2.5 (early access)
TensorRT-Edge-LLM: [fill in your version, e.g. 0.9.0]
CUDA: 13.3

Problem

Building the LLM engine for nvidia/Alpamayo-R1-10B (FP16, official TensorRT-Edge-LLM VLA workflow) consistently fails with an out-of-memory error during llm_build, even though the board has 58 GB unified memory with ~49 GB free.

./build/examples/llm/llm_build
–onnxDir …/Alpamayo-R1-10B/onnx/llm
–engineDir …/Alpamayo-R1-10B/engines/llm
–maxInputLen 3424 --maxKVCache

Failure (tail of log)

[TensorRT] Total Activation Memory: 100860416 bytes
[ERROR] [TensorRT] [resizingAllocator.cpp::allocateImpl::100] Error Code 1: Cuda Runtime
[WARNING] [TensorRT] Requested amount of GPU memory (15167679360 bytes) could not be allocated.
There may not be enough free memory for allocation to succeed.
[ERROR] [TensorRT] [myelinResourceManager.cpp::allocate::209] Error Code 2: OutOfMemory
(Requested size was 15167679360 bytes.)
[ERROR] [builderUtils.cpp:307:buildAndSerializeEngine] Failed to build serialized engine
[ERROR] [llm_build.cpp:243:main] Failed to build LLM engine.Capacity 4096 --maxBatchSize 6

The CUDA device only exposes 6 GiB

cudaMemGetInfo on the board:

cudaMemGetInfo ret = 0
GPU free = 5.84 GiB
GPU total = 6.0 GiB

Meanwhile the host has plenty of memory (free -h):

Mem: total 58Gi used 9.3Gi free 33Gi available 49Gi
Swap: 0B

So TensorRT requests a single 15167679360-byte (~14.13 GiB) allocation, but CUDA only exposes 6.0 GiB total — more than 2× the entire GPU-visible memory, so it can never succeed.
Additional checks

  • /proc/device-tree/reserved-memory/ has no gpu node (grep for gpu returns nothing), so the 6 GiB cap appears to be driver-enforced (nvgpu / DriveOS GPU memory manager), not a readable device-tree carveout.
  • tegrastats shows RAM 9068/59651MB and a largest-free-block of lfb 397x4MB (~1.5 GB), i.e. the GPU is idle and memory is fragmented into ≤1.5 GB contiguous blocks.
  • Root disk: was briefly 100% full, now 47% used — no change.
  • System RAM: 49 GB free — not the limit.
  • Locked memory: ulimit -l = unlimited.
  • Build flags: lowering maxBatchSize / maxInputLen / maxKVCacheCapacity does not meaningfully change the ~14.13 GiB allocation, consistent with it being a fixed Myelin working buffer for the 10B model.

Questions

  1. Is cudaMemGetInfo reporting only 6.0 GiB GPU total on a 58 GB Thor expected, or is the GPU memory allocation limit on this DriveOS 7.2.5 image misconfigured?
  2. How do we increase the GPU-visible/carveout memory so a 10B FP16 TensorRT engine (needs ~14+ GiB) can build? Is this an nvgpu / device-tree / boot-time / SDK setting, and does it require reflashing?
  3. If the limit cannot be raised, is there a validated INT8/QDQ Alpamayo-R1-10B LLM build profile for Thor that fits within the available GPU memory?

@SivaRamaKrishnaNV : Please for your review.

llm_build.log (32.4 KB)

Thanks for details. Try echo 8192 | sudo tee /sys/devices/system/node/node0/hugepages/hugepages-2048kB/nr_hugepages to see if it fixes memory issue. check increasing value to see if it fixes the issue.

Please see Carveout Customization and Profiling — NVIDIA DriveOS 7.0.3 Linux SDK Developer Guide helps.

Hello @SivaRamaKrishnaNV : It worked for building the engines , but there is issue while i run inferences .

thanks for your sharing detail, i will check and confirm whether it is working fine

Hello @SivaRamaKrishnaNV : when i tried to run inference, entry includes only output trajectory not the output text ?

Any idea , why this issue

Could you share the command, input files and complete log?

nvidia_submission_report.log (136.5 KB)

i am herewith attaching the log , added i used the below command

./build/examples/multimodal/action_inference
–engineDir “$ROOT/engines/llm”
–multimodalEngineDir “$ROOT/engines”
–inputFile “$WORKSPACE_DIR/input_action.json”
–outputFile “$WORKSPACE_DIR/output_action_fixed.json”
–batchSize 1
–maxGenerateLength 256
–debug
–dumpOutput
2>&1 | tee “$HOME/nvidia_submission_report.log”

thoruser@tegra-ubuntu:~/TensorRT-Edge-LLM$ jq '.responses[0] | {text_output_status: (if .output_text == "" then "EMPTY (Bug: String is completely unpopulated)" else .output_text end), trajectory_status: "SUCCESSFUL (\(.output_trajectory | length) spatial waypoints generated)"}' "$WORKSPACE_DIR/output_action_fixed.json" 
{
  "text_output_status": "EMPTY (Bug: String is completely unpopulated)",
  "trajectory_status": "SUCCESSFUL (64 spatial waypoints generated)"
}
thoruser@tegra-ubuntu:~/TensorRT-Edge-LLM$ 

thoruser@tegra-ubuntu:~/TensorRT-Edge-LLM$ jq '.responses[0] | {text_output_status: (if .output_text == "" then "EMPTY (Bug: String is completely unpopulated)" else .output_text end), trajectory_status: "SUCCESSFUL (\(.output_trajectory | length) spatial waypoints generated)"}' "$WORKSPACE_DIR/output_action_fixed.json" 
{
  "text_output_status": "EMPTY (Bug: String is completely unpopulated)",
  "trajectory_status": "SUCCESSFUL (64 spatial waypoints generated)"
}
thoruser@tegra-ubuntu:~/TensorRT-Edge-LLM$ 

Please find below my input_action.json

{
“batch_size”: 1,
“temperature”: 0.6,
“top_p”: 0.98,
“top_k”: 50,
“max_generate_length”: 256,
“requests”: [
{
“messages”: [
{
“role”: “system”,
“content”: [
{
“type”: “text”,
“text”: “You are a driving assistant that outputs actions based on video frames.”
},
{
“type”: “text”,
“text”: “You are a driving assistant that outputs actions based on video frames.”
}
]
},
{
“role”: “user”,
“content”: [
{
“type”: “text”,
“text”: “<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\n<|vision_start|><|image_pad|><|vision_end|>\noutput the driving action metrics.”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_00.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_01.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_02.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_03.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_04.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_05.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_06.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_07.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_08.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_09.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_10.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_11.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_12.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_13.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_14.png”
},
{
“type”: “image”,
“image”: “/home/thoruser/tensorrt-edgellm-workspace/frames/frame_15.png”
},
{
“type”: “trajectory”,
“trajectory”: [
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
],
[
0.0,
0.0,
0.0
]
]
}
]
}
]
}
],
“enable_thinking”: true
}

@SivaRamaKrishnaNV : Please let me know if any additional information are required

@SivaRamaKrishnaNV : is there any update on this topic

Thanks for sharing details. I will repro and get back to you. Does this block your development?

Thanks @SivaRamaKrishnaNV

Yes, it is blocking my development,

@SivaRamaKrishnaNV : Is there any update ?