Three models serving TP=2 on the same pair of DGX Sparks, same benchmark scripts. Prefill and TTFT first, one full-length request, 5 needles, token counts verified via /tokenize:
context Qwen3.8-FN FP8 GLM-5.3 NVFP4 DeepSeek-V4 NVFP4
ttft prefill ttft prefill ttft prefill
4,096 3.40s 1206 2.54s 1615 3.88s 1057
32,768 15.93s 2057 20.79s 1576 27.28s 1201
~131,072 50.60s 2591 66.71s 1965 114.65s 1134
Prefill in tok/s. Qwen’s prefill rate rises with context, 1206 to 2591. GLM and DeepSeek stay flat. At 131K that is Qwen at 50s against DeepSeek at 115s for the same prompt.
Needles 5 of 5 for all three at every depth. Max verified: GLM 131K, Qwen 245K, DeepSeek 265,781.
Single-stream decode, each at its own serving config:
Qwen3.8-FN NVFP4 29.6 tok/s 262K, no spec, seqs 64
Qwen3.8-FN FP8 23.0 tok/s 262K, no spec, seqs 64
GLM-5.3 NVFP4 21.5 tok/s MTP-4, seqs 6
DeepSeek-V4 NVFP4 41.0 tok/s 1M, DSpark spec, seqs 48
SWE-bench Pro, 40-instance enterprise subset, same harness:
Qwen3.8-Flash-Next 36/40 FP8 and NVFP4, 262K
Qwen3.8-Flash-Next 38/40 NVFP4, NVIDIA checkpoint, 1M YaRN, MTP-1
GLM-5.3-Flash 36/40 NVFP4, 262K
DeepSeek-V4-Flash I have not run on SWE-bench Pro.
Full configs and raw json per model: