Distributed Inference - 200gb/s with bottleneck, am I missing something?

Hello, I’ve been up all night to figure this out.

It seems vLLM is perfect for DGX Spark, the only problem is that a 14b models take up to 100gb of ram using vLLM and the performance are not great as advertised, I’ve followed the tutorial from the official website.

I’m currently using llama.cpp RPC and I use half of the ram, even more.

So am I getting something wrong? or this is it. Around 10k for 2 spark and getting 5token/s for a 32b param is a bit.. I don’t know, just not worth.

I’ve checked, the GPU has been used so is not running on CPU.

So I start to question, this hardware is designed for fp4, it costs like a macbook but with half the ram, is it just me that maybe I don’t know how to fully use the machine or this is it?

Please check out our playbooks on how to use different workloads on the Spark: Spark Playbooks
You can also look over community projects on the forums: DGX Spark / GB10 Projects - NVIDIA Developer Forums
There are also a lot of threads here about how to get better performance if you specifically want to use vLLM: Run VLLM in Spark

A few things:

  1. llama.cpp RPC doesn’t support RDMA - you’ll get a lot of performance loss due to latency.
  2. llama.cpp doesn’t do tensor parallel, so you will not gain any performance in the cluster, no matter how low latency the link is.
  3. You need to use vLLM to gain from the cluster setup. On dense models you can get close to 2x performance gain. You can use our community Docker build to get started - it is optimized for cluster setup. Make sure you configure the networking first! GitHub - eugr/spark-vllm-docker: Docker configuration for running VLLM on dual DGX Sparks
  4. vLLM will use 90% of available memory by default, even for smaller models. It is optimized for concurrency, so it will use as much VRAM for KV cache as possible. You can limit memory usage by specifying --gpu-memory-utilization 0.5 to limit to 50%, for example.
  5. Please, be mindful that Spark memory bandwidth is only 273GB/s, so choose models accordingly. Use lower quants and/or sparse MoE models vs dense ones.

Here is an example of what you could expect performance wise from Spark. Please note that NVFP4 support is still in its infancy for this platform and AWQ quants perform better for now.

Model name Cluster (t/s) Single (t/s) Comment
Qwen/Qwen3-VL-32B-Instruct-FP8 12.00 7.00
cpatonn/Qwen3-VL-32B-Instruct-AWQ-4bit 21.00 12.00
GPT-OSS-120B 55.00 36.00 SGLang gives 75/53, new experimental vLLM version from our community member here has even better performance. Soon to be incorporated into Docker build.
RedHatAI/Qwen3-VL-235B-A22B-Instruct-NVFP4 21.00 N/A
QuantTrio/Qwen3-VL-235B-A22B-Instruct-AWQ 26.00 N/A
Qwen/Qwen3-VL-30B-A3B-Instruct-FP8 65.00 52.00
QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ 97.00 82.00
RedHatAI/Qwen3-30B-A3B-NVFP4 75.00 64.00
QuantTrio/MiniMax-M2-AWQ 41.00 N/A
QuantTrio/GLM-4.6-AWQ 17.00 N/A
zai-org/GLM-4.6V-FP8 24.00 N/A

@eugr - quick random question. Do you see any perceivable difference in coherence between the NVFP4 and AWQ versions of Qwen3 235b? Presumably NVFP4 offers less loss?

NVFP4 if it’s W4A4 (as I believe that quant is - I deleted it as I prefer AWQ for now) has more loss as it quantizes activation weights to 4-bit too, while AWQ keeps activation weights in BF16.

Now, NVFP4 W4A16 theoretically would be superior to AWQ quants as FP4 carries more precision than INT4, but in practice it depends on whether/how it was calibrated using a calibration dataset.

As for these specific quants, I didn’t notice any significant difference in practice, but again, I haven’t used the NVFP4 much.

Thanks – swinging by Microcenter to pick up that second Spark today and thinking about what big model I’m going to load up first. :-)