# TRT LLM for Inference - two Sparks example is VERY slow

**URL:** <https://forums.developer.nvidia.com/t/trt-llm-for-inference-two-sparks-example-is-very-slow/348517>\
**Category:** DGX Spark / GB10\
**Created:** [October 21, 2025, 1:38pm UTC](https://forums.developer.nvidia.com/t/trt-llm-for-inference-two-sparks-example-is-very-slow/348517 "2025-10-21T13:38:59Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![vgoklani](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/vgoklani/32/11912_2.png) [@vgoklani](https://forums.developer.nvidia.com/u/vgoklani)\
**Post date:** [October 21, 2025, 1:38pm UTC](https://forums.developer.nvidia.com/t/trt-llm-for-inference-two-sparks-example-is-very-slow/348517/1 "2025-10-21T13:38:59Z")

</div>

I ran this example:

docker exec  
-e MODEL=“nvidia/Qwen3-235B-A22B-FP4”  
-e HF\_TOKEN=$HF\_TOKEN  
-it $TRTLLM\_MN\_CONTAINER bash -c ’  
mpirun -x HF\_TOKEN trtllm-llmapi-launch trtllm-serve $MODEL  
–tp\_size 2  
–backend pytorch  
–max\_num\_tokens 32768  
–max\_batch\_size 4  
–extra\_llm\_api\_options /tmp/extra-llm-api-config.yml  
–port 8355’

from this page: “[Try NVIDIA NIM APIs](https://build.nvidia.com/spark/trt-llm/stacked-sparks%E2%80%9D)

and the inference speed was very-very slow.

The example from the single spark page: “[Try NVIDIA NIM APIs](https://build.nvidia.com/spark/trt-llm/instructions%E2%80%9D)

export MODEL\_HANDLE=“openai/gpt-oss-120b”

docker run  
-e MODEL\_HANDLE=$MODEL\_HANDLE  
-e HF\_TOKEN=$HF\_TOKEN  
-v $HOME/.cache/huggingface/:/root/.cache/huggingface/  
–rm -it --ulimit memlock=-1 --ulimit stack=67108864  
–gpus=all --ipc=host --network host  
[nvcr.io/nvidia/tensorrt-llm/release:spark-single-gpu-dev](http://nvcr.io/nvidia/tensorrt-llm/release:spark-single-gpu-dev)  
bash -c ’  
export TIKTOKEN\_ENCODINGS\_BASE=“/tmp/harmony-reqs” &&  
mkdir -p $TIKTOKEN\_ENCODINGS\_BASE &&  
wget -P $TIKTOKEN\_ENCODINGS\_BASE [https://openaipublic.blob.core.windows.net/encodings/o200k\_base.tiktoken](https://openaipublic.blob.core.windows.net/encodings/o200k_base.tiktoken) &&  
wget -P $TIKTOKEN\_ENCODINGS\_BASE [https://openaipublic.blob.core.windows.net/encodings/cl100k\_base.tiktoken](https://openaipublic.blob.core.windows.net/encodings/cl100k_base.tiktoken) &&  
hf download $MODEL\_HANDLE &&  
python examples/llm-api/quickstart\_advanced.py  
–model\_dir $MODEL\_HANDLE  
–prompt “Paris is great because”  
–max\_tokens 64  
’  
was much-much faster, and I suspect that was because it used an optimized branch of tensor-rt:

”[nvcr.io/nvidia/tensorrt-llm/release:spark-single-gpu-dev”](http://nvcr.io/nvidia/tensorrt-llm/release:spark-single-gpu-dev%E2%80%9D)

Is there an optimized version for two sparks? Please respond with an actual answer and don’t just say “read the documentation”.thanks

[nvcr.io/nvidia/tensorrt-llm/release:spark-single-gpu-dev](http://nvcr.io/nvidia/tensorrt-llm/release:spark-single-gpu-dev)

---

<div class="post-metadata">

**Author:** ![cosinus](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/cosinus/32/454111_2.png) [@cosinus](https://forums.developer.nvidia.com/u/cosinus)\
**Post date:** [October 21, 2025, 1:55pm UTC](https://forums.developer.nvidia.com/t/trt-llm-for-inference-two-sparks-example-is-very-slow/348517/2 "2025-10-21T13:55:09Z")

</div>

You can see all availabe tags in the NGC catalog. There is no explicite dual-node image, but there is one newer image than spark-single-gpu-dev.

> **[GPU-optimized AI, Machine Learning, & HPC Software | NVIDIA NGC](https://catalog.ngc.nvidia.com/orgs/nvidia/teams/tensorrt-llm/containers/release/tags?version=1.2.0rc0.post1)**
>
> GPU-optimized AI, Machine Learning, & HPC Software | NVIDIA NGC

You could try that.

rc0.post1 was published on 10/14/2025 10:42 AM - bleeding edge  
single-spark was published on 10/06/2025 10:20 PM

Still waiting for a newer build, because there were improvements for RTX5090/sm120 for FP4 a few days ago in the main tree. That’s why I stumbled upon it.

---

<div class="post-metadata">

**Author:** ![vgoklani](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/vgoklani/32/11912_2.png) [@vgoklani](https://forums.developer.nvidia.com/u/vgoklani)\
**Post date:** [October 22, 2025, 3:06pm UTC](https://forums.developer.nvidia.com/t/trt-llm-for-inference-two-sparks-example-is-very-slow/348517/4 "2025-10-22T15:06:36Z")

</div>

We posted a detailed issue on github:

> <https://github.com/NVIDIA/dgx-spark-playbooks/issues/3>
>
> I followed the instructions from here:
> 
> https://github.com/NVIDIA/dgx-spark-play…books/tree/main/nvidia/trt-llm#step-11-serve-the-model
> 
> and ran the model across two sparks:
> 
> \`\`\`bash
> docker exec \\
> -e MODEL="nvidia/Qwen3-235B-A22B-FP4" \\
> -e HF\_TOKEN=$HF\_TOKEN \\
> -it $TRTLLM\_MN\_CONTAINER bash -c '
> mpirun -x HF\_TOKEN trtllm-llmapi-launch trtllm-serve $MODEL \\
> --tp\_size 2 \\
> --backend pytorch \\
> --max\_num\_tokens 32768 \\
> --max\_batch\_size 4 \\
> --extra\_llm\_api\_options /tmp/extra-llm-api-config.yml \\
> --port 8355'
> \`\`\`
> 
> but both the \`pre-fill\` and \`decoding\` speeds were incredibly slow. I expected this to be a lot faster since we were using tensor-parallel, have an nv-link connection between the two boxes, and ran an FP4 version of the model.
> 
> I also ran gpt-oss-120B on a single spark:
> 
> \`\`\`bash
> export MODEL\_HANDLE="openai/gpt-oss-120b"
> 
> docker run --name trtllm\_llm\_server --rm -it --gpus all --ipc host --network host \\
> -e HF\_TOKEN=$HF\_TOKEN \\
> -e MODEL\_HANDLE="$MODEL\_HANDLE" \\
> -v $HOME/.cache/huggingface/:/root/.cache/huggingface/ \\
> nvcr.io/nvidia/tensorrt-llm/release:spark-single-gpu-dev \\
> bash -c '
> export TIKTOKEN\_ENCODINGS\_BASE="/tmp/harmony-reqs" && \\
> mkdir -p $TIKTOKEN\_ENCODINGS\_BASE && \\
> wget -P $TIKTOKEN\_ENCODINGS\_BASE https://openaipublic.blob.core.windows.net/encodings/o200k\_base.tiktoken && \\
> wget -P $TIKTOKEN\_ENCODINGS\_BASE https://openaipublic.blob.core.windows.net/encodings/cl100k\_base.tiktoken && \\
> hf download $MODEL\_HANDLE && \\
> cat \> /tmp/extra-llm-api-config.yml \<\<EOF
> print\_iter\_log: false
> kv\_cache\_config:
> dtype: "auto"
> free\_gpu\_memory\_fraction: 0.9
> cuda\_graph\_config:
> enable\_padding: true
> disable\_overlap\_scheduler: true
> EOF
> trtllm-serve "$MODEL\_HANDLE" \\
> --max\_batch\_size 64 \\
> --trust\_remote\_code \\
> --port 8355 \\
> --extra\_llm\_api\_options /tmp/extra-llm-api-config.yml
> '
> \`\`\`
> 
> and both the \`pre-fill\` and \`decoding\` speeds were acceptable (i.e. similar to the token-generation speed that a user experiences on the chatgpt website). I suspect the improved performance was due to this container: \`nvcr.io/nvidia/tensorrt-llm/release:spark-single-gpu-dev\` which has optimized cutlass kernels for the spark. Why is the dual-spark tutorial using \`nvcr.io/nvidia/tensorrt-llm/release:1.0.0rc3\`?
> 
> Thanks!

If we are not able to resolve this, then we will just return the two sparks back to nvidia.

---

<div class="post-metadata">

**Author:** ![aniculescu](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/aniculescu/32/382649_2.png) [@aniculescu](https://forums.developer.nvidia.com/u/aniculescu)\
**Post date:** [October 22, 2025, 3:20pm UTC](https://forums.developer.nvidia.com/t/trt-llm-for-inference-two-sparks-example-is-very-slow/348517/6 "2025-10-22T15:20:52Z")

</div>

What performance numbers are you seeing for single vs stacked setups?

Edit: Please also use `trtllm-bench` to collect performance numbers

---

<div class="post-metadata">

**Author:** ![vgoklani](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/vgoklani/32/11912_2.png) [@vgoklani](https://forums.developer.nvidia.com/u/vgoklani)\
**Post date:** [October 22, 2025, 9:15pm UTC](https://forums.developer.nvidia.com/t/trt-llm-for-inference-two-sparks-example-is-very-slow/348517/8 "2025-10-22T21:15:22Z")

</div>

I’ve had much better luck with the `nvcr.io/nvidia/tensorrt-llm/release:1.2.0rc1` image that was released today. Could you please share a docker-command to run the `trtllm-bench` benchmark for both single and dual sparks. It feels like the single model instance is faster than the sharded model, so I need to properly benchmark them. Thanks!

---

<div class="post-metadata">

**Author:** ![eugr](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/eugr/32/449615_2.png) [@eugr](https://forums.developer.nvidia.com/u/eugr)\
**Post date:** [October 23, 2025, 5:07pm UTC](https://forums.developer.nvidia.com/t/trt-llm-for-inference-two-sparks-example-is-very-slow/348517/9 "2025-10-23T17:07:05Z")

</div>

You are comparing a model with 22B active parameters with a model with 3B active parameters. Of course gpt-oss-120b will be faster…

---

<div class="post-metadata">

**Author:** ![system](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/system/32/68080_2.png) [@system](https://forums.developer.nvidia.com/u/system)\
**Post date:** [December 10, 2025, 8:05pm UTC](https://forums.developer.nvidia.com/t/trt-llm-for-inference-two-sparks-example-is-very-slow/348517/10 "2025-12-10T20:05:46Z")

</div>

This topic was automatically closed 14 days after the last reply. New replies are no longer allowed.
