It seems vLLM is perfect for DGX Spark, the only problem is that a 14b models take up to 100gb of ram using vLLM and the performance are not great as advertised, I’ve followed the tutorial from the official website.
I’m currently using llama.cpp RPC and I use half of the ram, even more.
So am I getting something wrong? or this is it. Around 10k for 2 spark and getting 5token/s for a 32b param is a bit.. I don’t know, just not worth.
I’ve checked, the GPU has been used so is not running on CPU.
So I start to question, this hardware is designed for fp4, it costs like a macbook but with half the ram, is it just me that maybe I don’t know how to fully use the machine or this is it?
vLLM will use 90% of available memory by default, even for smaller models. It is optimized for concurrency, so it will use as much VRAM for KV cache as possible. You can limit memory usage by specifying --gpu-memory-utilization 0.5 to limit to 50%, for example.
Please, be mindful that Spark memory bandwidth is only 273GB/s, so choose models accordingly. Use lower quants and/or sparse MoE models vs dense ones.
Here is an example of what you could expect performance wise from Spark. Please note that NVFP4 support is still in its infancy for this platform and AWQ quants perform better for now.
Model name
Cluster (t/s)
Single (t/s)
Comment
Qwen/Qwen3-VL-32B-Instruct-FP8
12.00
7.00
cpatonn/Qwen3-VL-32B-Instruct-AWQ-4bit
21.00
12.00
GPT-OSS-120B
55.00
36.00
SGLang gives 75/53, new experimental vLLM version from our community member here has even better performance. Soon to be incorporated into Docker build.
@eugr - quick random question. Do you see any perceivable difference in coherence between the NVFP4 and AWQ versions of Qwen3 235b? Presumably NVFP4 offers less loss?
NVFP4 if it’s W4A4 (as I believe that quant is - I deleted it as I prefer AWQ for now) has more loss as it quantizes activation weights to 4-bit too, while AWQ keeps activation weights in BF16.
Now, NVFP4 W4A16 theoretically would be superior to AWQ quants as FP4 carries more precision than INT4, but in practice it depends on whether/how it was calibrated using a calibration dataset.
As for these specific quants, I didn’t notice any significant difference in practice, but again, I haven’t used the NVFP4 much.