I am deploying nvidia/Alpamayo-R1-10B with TensorRT-Edge-LLM on DRIVE AGX Thor running DriveOS 7.2.5.0.
CUDA: 13.3
TensorRT reported by llm_build: 11.0.1
SDK container: driveos-sdk … 7.2.5.0-0004
Target: auto-thor
The ONNX model parses successfully and AttentionPlugin loads successfully. Engine serialization fails with:
Requested amount of GPU memory (15167621888 bytes) could not be allocated
OutOfMemory
This remains approximately 15.17 GB with maxBatchSize=1, maxInputLen=2048, and maxKVCacheCapacity=2048.
Is TensorRT 11.0.1 supported for Alpamayo-R1-10B on DRIVE AGX Thor with DriveOS 7.2.5? Is the fixed 15.17 GB allocation expected? What exact TensorRT-Edge-LLM commit, SDK container, CMake options, and FP16/INT8 workflow are validated for Thor? Is the intended deployment path a quantized INT8/QDQ Expert engine rather than direct FP16 LLM engine construction?
Is the intended DRIVE AGX Thor deployment path for Alpamayo an INT8/QDQ Expert engine rather than direct FP16 TensorRT-Edge-LLM conversion of the full 10B LLM? The direct FP16 LLM build requires a fixed 15.17 GB allocation and fails during serialization on Thor, even with batch size 1 and KV cache 2048. Please provide the official Thor workflow for SmoothQuant/INT8 Expert conversion, including the supported TensorRT-Edge-LLM release, ONNX export command, engine-build command, and whether the full LLM engine is expected to run on Thor.
Following confirmation i did
Available RAM before build: approximately 40–50 GB
virtual memory: unlimited
max memory size: unlimited
memlock: unlimited
can you please share the details