NVIDIA Developer Forums
Qwen 3.8 27B + DFlash2
Accelerated Computing
DGX Spark / GB10 User Forum
DGX Spark / GB10
performance
danilo.luvizotto
August 19, 2026, 6:22pm
4
yes, I’m getting 40-42 tok/s
show post in topic
Related topics
Topic
Replies
Views
Activity
Qwen3.8-27B (NVFP4) on Single/Dual DGX Spark — SGLang + DFlash2, fully OpenAI-compatible
DGX Spark / GB10
15
5848
September 6, 2026
Qwen3.8-27B on dual Sparks
DGX Spark / GB10
agentic-ai
17
6534
August 22, 2026
Qwen3.8-27B at 34–38 tok/s on DGX Spark — open-source one-command setup (SGLang + NVFP4 + DSpark)
DGX Spark / GB10 Projects
llm
,
llama
,
agentic-ai
,
deepseek
106
19213
September 16, 2026
Qwen3.8-27B benchmarking on one DGX Spark: DFlash2 beat vLLM+MTP, and greedy beat the "thinking" sampler
DGX Spark / GB10 Projects
benchmarks
,
llm
,
agentic-ai
24
5465
September 7, 2026
Comprehensive Qwen3.8-27B Study on DGX Sparks: Quantization, Speculative Decoding, and TP/DP Scaling
DGX Spark / GB10 Projects
10
4779
August 26, 2026
DFlash LLM for DGX Spark - too good to be true?
DGX Spark / GB10
37
4398
April 17, 2026
Qwen3.8-Flash-Next on 1, 2 and 4 DGX Sparks with NVIDIA's official NVFP4 quant: 64 tok/s peak single stream
DGX Spark / GB10
27
5072
September 10, 2026
Single DGX-Spark - Qwen 3.8-Flash-Next at ~43tok/sec in Coding
DGX Spark / GB10
63
11922
September 14, 2026
1x Spark new workhorse: DFlash2 + SGLang is fairly fast for Qwen3.8 27B NVFP4
DGX Spark / GB10
3
1045
August 26, 2026
Qwen3.8-Flash-Next 180B, Single Solo DGX Spark With HashK-PLE NVFP4
DGX Spark / GB10 Projects
42
6375
September 5, 2026