Comprehensive Qwen3.8-27B Study on DGX Sparks: Quantization, Speculative Decoding, and TP/DP Scaling

We benchmarked Qwen3.8-27B across 11 configurations on one to four DGX Sparks to see what actually moves performance: precision, speculative decoding, runtime/image choice, and tensor parallelism.

Across the sequence, C1 generation increased from 4.5 to 77.3 tok/s. Below are the full concurrency sweeps and the exact recipes used at each step.

Stack

  • 4Γ— NVIDIA DGX Spark (GB10 Grace Blackwell, SM 12.1, 128 GB UMA, 273 GB/s LPDDR5X)

  • CX-7 (100 Gbps) interconnect between nodes

  • sparkrun v0.3.4

  • Benchmark: CordatusAI LLM Benchmark Tool β€” 128 tokens in, 128 tokens out, 10 rounds per concurrency level. Pass criteria at a concurrency: TTFT < 1000ms AND TPS >= 15 tok/s

  • Images: vLLM official (vllm/vllm-openai:qwen38), eugr vLLM (ghcr.io/spark-arena/dgx-vllm-eugr-nightly:latest), SGLang (lmsysorg/sglang, pinned by SHA in recipes)

  • Models: Qwen3.8-27B BF16 (Qwen/Qwen3.8-27B, 55.6 GB), Qwen3.8-27B FP8 (Qwen/Qwen3.8-27B-FP8, 30.9 GB), Qwen3.8-27B NVFP4 β€” InferAct (Inferact/Qwen3.8-27B-NVFP4, ~18 GB), Qwen3.8-27B NVFP4 β€” RadixArk (RadixArk/Qwen3.8-27B-NVFP4, ~18 GB), DSpark draft model (RadixArk/Qwen3.8-27B-DSpark, 1.36B)

Progression Summary

Step Configuration What changed TPS @ C1 TTFT @ C1 Max C (both pass)
1 Qwen3.8-27B BF16, TP=1, vLLM official β€” 4.5 335ms Never
2 Qwen3.8-27B BF16 + MTP, TP=1, vLLM official +MTP n=3 9.9 611ms Never
3 Qwen3.8-27B FP8, TP=1, vLLM official BF16β†’FP8 7.9 172ms Never
4 Qwen3.8-27B FP8 + MTP, TP=1, vLLM official +MTP n=3 17.1 392ms C4
5 Qwen3.8-27B NVFP4 (InferAct), TP=1, vLLM official FP8β†’NVFP4 9.7 155ms Never
6 Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=1, vLLM official +MTP n=3 18.5 341ms C8
7 Qwen3.8-27B NVFP4 (RadixArk) + DSpark, TP=1, SGLang engine + spec decode switch 36.6 246ms C4
8 Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=2, vLLM official +1 Spark 23.4 254ms C8
9 Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=2, eugr vLLM official→eugr 33.0 311ms C16
10 Qwen3.8-27B NVFP4 (RadixArk) + DSpark, TP=2, SGLang +1 Spark 51.8 219ms C16
11 Qwen3.8-27B NVFP4 (RadixArk) + DSpark, TP=4, SGLang +2 Sparks 77.3 227ms C16

Because Qwen3.8-27B is a dense model, every concurrent request multiplies against the same weight matrices. The GPU loads each weight tile once and serves all requests in a single batched GEMM. Since weight loading β€” the bottleneck in the memory-bandwidth-bound regime β€” doesn’t change with batch size, per-user TPS stays relatively flat at low concurrency, then gradually decreases as the batch grows large enough to saturate the SMs and enter the compute-bound regime. We observe this pattern across almost every run in this study. That is why most runs either never pass the concurrency threshold or pass it with a higher concurrency than C1 β€” once the model can serve enough tokens at C1, it holds up well under increasing concurrency. This is also why C1 TPS values β€” where decode is purely memory-bound β€” sit very close to the memory bandwidth theoretical ceiling.


BF16 (Qwen/Qwen3.8-27B), TP=1, vLLM official

The model is 55.6 GB. A simple memory-bandwidth-only upper bound is 273 / 55.6 = 4.9 tok/s; we measured 4.5 tok/s at C1. TPS stays close to that level as concurrency rises, but the configuration never reaches our 15 tok/s threshold. TTFT crosses 1 second at C8.

C TTFT (ms) TPS Pass?
1 335 4.5 βœ— (TPS)
2 610 4.3 βœ— (TPS)
4 797 4.2 βœ— (TPS)
8 1064 4.1 βœ— (both)
16 1670 3.8 βœ— (both)
recipe_version: "2"
model: Qwen/Qwen3.8-27B
runtime: vllm
container: vllm/vllm-openai:qwen38
​
executor_config:
  entrypoint: ""
​
max_nodes: 1
​
metadata:
  kv_dtype: fp8
  description: Qwen3.8-27B BF16, TP=1, no MTP
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.8
​
command: |
  vllm serve Qwen/Qwen3.8-27B \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder
sparkrun run ~/qwen38-27b-bf16.yaml

BF16 (Qwen/Qwen3.8-27B) + MTP, TP=1, vLLM official

Adding MTP with three speculative tokens raises C1 TPS from 4.5 to 9.9, a 2.2Γ— increase. The gain holds through the lower concurrency levels, but it is still not enough to reach 15 tok/s.

C TTFT (ms) TPS Pass?
1 611 9.9 βœ— (TPS)
2 867 10.0 βœ— (TPS)
4 916 10.0 βœ— (TPS)
8 1067 9.0 βœ— (both)
16 1250 7.9 βœ— (both)
recipe_version: "2"
model: Qwen/Qwen3.8-27B
runtime: vllm
container: vllm/vllm-openai:qwen38
​
executor_config:
  entrypoint: ""
​
max_nodes: 1
​
metadata:
  kv_dtype: fp8
  description: Qwen3.8-27B BF16 + MTP, TP=1
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.8
​
command: |
  vllm serve Qwen/Qwen3.8-27B \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --speculative-config '{"method":"mtp","num_speculative_tokens":3}'
sparkrun run ~/qwen38-27b-bf16-mtp.yaml

FP8 (Qwen/Qwen3.8-27B-FP8), TP=1, vLLM official

FP8 halves the model to 30.9 GB. Theoretical ceiling becomes 273 / 30.9 = 8.8 tok/s. We hit 7.9 β€” 90% of theoretical max. TTFT performance almost doubles as expected (335 β†’ 172ms) β€” half the weight loading during prefill. 7.9 TPS is still below 15.

C TTFT (ms) TPS Pass?
1 172 7.9 βœ— (TPS)
2 322 7.8 βœ— (TPS)
4 435 7.6 βœ— (TPS)
8 769 7.2 βœ— (TPS)
16 1401 6.4 βœ— (both)
32 3332 5.0 βœ— (both)
recipe_version: "2"
model: Qwen/Qwen3.8-27B-FP8
runtime: vllm
container: vllm/vllm-openai:qwen38
​
executor_config:
  entrypoint: ""
​
max_nodes: 1
​
metadata:
  kv_dtype: fp8
  description: Qwen3.8-27B FP8, TP=1, no MTP
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.8
​
command: |
  vllm serve Qwen/Qwen3.8-27B-FP8 \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder
sparkrun run ~/qwen38-27b-fp8.yaml

FP8 (Qwen/Qwen3.8-27B-FP8) + MTP, TP=1, vLLM official

This is the first configuration to pass our criteria. MTP raises C1 throughput from 7.9 to 17.1 tok/s, a 2.16Γ— increase, very close to the 2.2Γ— gain seen with BF16. C1 through C4 pass; at C8, TTFT is still comfortably below 1 second but TPS falls just below 15.

C TTFT (ms) TPS Pass?
1 392 17.1 βœ“
2 557 17.1 βœ“
4 565 16.1 βœ“
8 672 14.4 βœ— (TPS)
16 999 11.7 βœ— (TPS)
32 1504 8.6 βœ— (both)
recipe_version: "2"
model: Qwen/Qwen3.8-27B-FP8
runtime: vllm
container: vllm/vllm-openai:qwen38
​
executor_config:
  entrypoint: ""
​
max_nodes: 1
​
metadata:
  kv_dtype: fp8
  description: Qwen3.8-27B FP8 + MTP, TP=1
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.8
​
command: |
  vllm serve Qwen/Qwen3.8-27B-FP8 \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --speculative-config '{"method":"mtp","num_speculative_tokens":3}'
sparkrun run ~/qwen38-27b-fp8-mtp.yaml

NVFP4 (Inferact/Qwen3.8-27B-NVFP4), TP=1, vLLM official

NVFP4 reduces the model to roughly 18 GB, which gives a simple bandwidth-only estimate of about 15.2 tok/s. The measured C1 result is 9.7 tok/s, 64% of theoretical max. Thus, unlike BF16 and FP8, model footprint alone is no longer a good predictor of realized decode throughput on this stack. The new bottleneck might be the vLLM’s NVFP4 kernel efficiency on SM 12.

C TTFT (ms) TPS Pass?
1 155 9.7 βœ— (TPS)
2 270 9.6 βœ— (TPS)
4 333 9.3 βœ— (TPS)
8 520 8.8 βœ— (TPS)
16 899 7.7 βœ— (TPS)
32 1502 6.2 βœ— (both)
recipe_version: "2"
model: Inferact/Qwen3.8-27B-NVFP4
runtime: vllm
container: vllm/vllm-openai:qwen38

executor_config:
  entrypoint: ""

max_nodes: 1

metadata:
  kv_dtype: fp8
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (InferAct), TP=1, no MTP

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.8

command: |
  vllm serve Inferact/Qwen3.8-27B-NVFP4 \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder
sparkrun run ~/qwen38-27b-nvfp4-inferact.yaml

NVFP4 (Inferact/Qwen3.8-27B-NVFP4) + MTP, TP=1, vLLM official

MTP raises C1 TPS from 9.7 to 18.5, a 1.91Γ— increase. It passes through C8, with TTFT still at 603 ms there. At C16, TTFT remains below the threshold but TPS falls to 12.5.

C TTFT (ms) TPS Pass?
1 341 18.5 βœ“
2 465 18.8 βœ“
4 517 17.3 βœ“
8 603 15.3 βœ“
16 767 12.5 βœ— (TPS)
32 1095 9.2 βœ— (both)
recipe_version: "2"
model: Inferact/Qwen3.8-27B-NVFP4
runtime: vllm
container: vllm/vllm-openai:qwen38

executor_config:
  entrypoint: ""

max_nodes: 1

metadata:
  kv_dtype: fp8
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=1

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.8

command: |
  vllm serve Inferact/Qwen3.8-27B-NVFP4 \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --speculative-config '{"method":"mtp","num_speculative_tokens":3}'
sparkrun run ~/qwen38-27b-nvfp4-inferact-mtp.yaml

NVFP4 (RadixArk/Qwen3.8-27B-NVFP4) + DSpark, TP=1, SGLang

This is the strongest single-Spark configuration in the study. This single Spark configuration will beat the best 2-Spark vLLM run at C1–C2: 36.6 TPS vs 33.0 TPS, with half the node count. This success is the combination of the SGLang image NVFP4 optimization and DSpark (proposes 7 tokens per speculation) high acceptance rate.

The main weakness appears at higher concurrency. TTFT jumps from 394 ms at C4 to 1368 ms at C8. The recipe uses --cuda-graph-max-bs 4, so C8 and above are outside that configured CUDA-graph batch-size range. That coincides with the observed jump, although this benchmark alone does not isolate the exact causation.

C TTFT (ms) TPS Pass?
1 246 36.6 βœ“
2 349 31.3 βœ“
4 394 25.6 βœ“
8 1368 17.8 βœ— (TTFT)
16 8554 9.2 βœ— (both)
32 23045 4.7 βœ— (both)
recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4
runtime: sglang
container: lmsysorg/sglang@sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1

max_nodes: 1

metadata:
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (RadixArk) + DSpark β€” SGLang on DGX Spark

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.50
  served_model_name: qwen3.8-27b
  attention_backend: flashinfer
  speculative_algorithm: DSPARK
  speculative_draft_model_path: RadixArk/Qwen3.8-27B-DSpark
  speculative_dspark_block_size: 7
  speculative_draft_model_quantization: unquant

env:
  TORCHINDUCTOR_CACHE_DIR: /cache/inductor
  HF_HOME: /cache/huggingface
  HF_HUB_CACHE: /cache/huggingface/hub

command: |
  python3 -m sglang.launch_server \
    --trust-remote-code \
    --model-path {model} \
    --tp-size {tensor_parallel} \
    --served-model-name {served_model_name} \
    --mem-fraction-static {gpu_memory_utilization} \
    --attention-backend {attention_backend} \
    --chunked-prefill-size 8192 \
    --disable-prefill-cuda-graph \
    --cuda-graph-max-bs 4 \
    --speculative-algorithm {speculative_algorithm} \
    --speculative-draft-model-path {speculative_draft_model_path} \
    --speculative-dspark-block-size {speculative_dspark_block_size} \
    --speculative-draft-model-quantization {speculative_draft_model_quantization} \
    --mamba-scheduler-strategy extra_buffer \
    --enable-torch-compile --torch-compile-max-bs 4 \
    --num-continuous-decode-steps 2 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --host {host} \
    --port {port}
sparkrun run ~/qwen38-27b-sglang-dspark.yaml

NVFP4 (Inferact/Qwen3.8-27B-NVFP4) + MTP, TP=2, vLLM official

Moving the same InferAct + MTP setup from TP=1 to TP=2 raises C1 TPS from 18.5 to 23.4, about 26%. The scaling is useful but far from 2Γ—. Tensor parallelism reduces the amount of sharded model work handled by each GPU, while also introducing communication on the decode path, so linear scaling should not be expected. TTFT improves from 341 ms to 254 ms at C1. The configuration still passes through C8 and falls below the TPS threshold at C16.

C TTFT (ms) TPS Pass?
1 254 23.4 βœ“
2 423 22.4 βœ“
4 454 21.3 βœ“
8 570 17.0 βœ“
16 699 12.8 βœ— (TPS)
32 859 10.7 βœ— (TPS)
recipe_version: "2"
model: Inferact/Qwen3.8-27B-NVFP4
runtime: vllm-distributed
container: vllm/vllm-openai:qwen38

executor_config:
  entrypoint: ""

metadata:
  kv_dtype: fp8
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=2, official vLLM image

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 2
  gpu_memory_utilization: 0.75

command: |
  vllm serve Inferact/Qwen3.8-27B-NVFP4 \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --speculative-config '{"method":"mtp","num_speculative_tokens":3}'
sparkrun run ~/qwen38-27b-inferact-nvfp4-tp2.yaml --hosts <head>,<worker1>

NVFP4 (Inferact/Qwen3.8-27B-NVFP4) + MTP, TP=2, eugr vLLM

The eugr recipe reaches 33.0 tok/s at C1 versus 23.4 with the official vLLM configuration and remains ahead throughout the concurrency sweep. It is also the first vLLM setup here to satisfy both criteria at C16.

The native SM 12.1a support in the eugr stack is a plausible contributor to the improvement. In addition, the eugr recipe changes several serving parameters, including memory utilization, maximum model length, attention backend, load format, chunked prefill, and asynchronous scheduling.

C TTFT (ms) TPS Pass?
1 311 33.0 βœ“
2 467 31.8 βœ“
4 540 29.5 βœ“
8 599 24.0 βœ“
16 831 17.7 βœ“
32 1345 13.4 βœ— (both)
recipe_version: '2'
name: Qwen3.8-27B-NVFP4-MTP-TP2
model: Inferact/Qwen3.8-27B-NVFP4
container: ghcr.io/spark-arena/dgx-vllm-eugr-nightly:latest
runtime: vllm-distributed

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 2
  gpu_memory_utilization: 0.55
  max_model_len: 131072
  max_num_batched_tokens: 8192
  served_model_name: Qwen3.8-27B
  attention_backend: flashinfer
  tool_call_parser: qwen3_coder
  pipeline_parallel: 1
  load_format: fastsafetensors
  kv_cache_dtype: fp8
  reasoning_parser: qwen3
  speculative_config: '{"method":"mtp","num_speculative_tokens":3}'
  override_generation_config: '{"temperature":1.0,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'

env:
  VLLM_MARLIN_USE_ATOMIC_ADD: '1'
  TORCH_MATMUL_PRECISION: high
  FLASHINFER_DISABLE_VERSION_CHECK: '1'
  NVIDIA_FORWARD_COMPAT: '1'
  VLLM_HTTP_TIMEOUT_KEEP_ALIVE: '600'

command: |
  vllm serve {model} \
    --served-model-name {served_model_name} \
    --host {host} \
    --port {port} \
    -tp {tensor_parallel} \
    -pp {pipeline_parallel} \
    --trust-remote-code \
    --language-model-only \
    --kv-cache-dtype {kv_cache_dtype} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len {max_model_len} \
    --max-num-batched-tokens {max_num_batched_tokens} \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --async-scheduling \
    --load-format {load_format} \
    --attention-backend {attention_backend} \
    --enable-auto-tool-choice \
    --tool-call-parser {tool_call_parser} \
    --reasoning-parser {reasoning_parser} \
    --speculative-config '{speculative_config}' \
    --override-generation-config '{override_generation_config}'

metadata:
  description: Qwen3.8-27B NVFP4 (Inferact) + MTP β€” TP=2 dual Spark, eugr image
  model_params: 27B
  model_dtype: nvfp4
  kv_dtype: fp8
  quantization: nvfp4
  quant_bits: 4
sparkrun run ~/qwen38-27b-inferact-nvfp4-tp2-eugr.yaml --hosts <head>,<worker1>

NVFP4 (RadixArk/Qwen3.8-27B-NVFP4) + DSpark, TP=2, SGLang

This is a cleaner scaling comparison with the previous SGLang run: moving from TP=1 to TP=2 raises C1 TPS from 36.6 to 51.8, about 42%. The larger gain is not limited to C1. The configuration remains above both thresholds through C16, whereas the single-Spark run fails at C8 because of TTFT.

C TTFT (ms) TPS Pass?
1 219 51.8 βœ“
2 286 46.4 βœ“
4 324 36.0 βœ“
8 375 24.9 βœ“
16 460 16.7 βœ“
32 5682 9.7 βœ— (both)
recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4
runtime: sglang
container: lmsysorg/sglang@sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1

metadata:
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (RadixArk) + DSpark β€” SGLang TP=2 on 2x DGX Spark

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 2
  gpu_memory_utilization: 0.50
  served_model_name: qwen3.8-27b
  attention_backend: flashinfer
  speculative_algorithm: DSPARK
  speculative_draft_model_path: RadixArk/Qwen3.8-27B-DSpark
  speculative_dspark_block_size: 7
  speculative_draft_model_quantization: unquant

env:
  TORCHINDUCTOR_CACHE_DIR: /cache/huggingface/sglang-inductor
  HF_HOME: /cache/huggingface
  HF_HUB_CACHE: /cache/huggingface/hub

command: |
  python3 -m sglang.launch_server \
    --trust-remote-code \
    --model-path {model} \
    --tp-size {tensor_parallel} \
    --served-model-name {served_model_name} \
    --mem-fraction-static {gpu_memory_utilization} \
    --attention-backend {attention_backend} \
    --chunked-prefill-size 8192 \
    --disable-prefill-cuda-graph \
    --cuda-graph-max-bs 4 \
    --speculative-algorithm {speculative_algorithm} \
    --speculative-draft-model-path {speculative_draft_model_path} \
    --speculative-dspark-block-size {speculative_dspark_block_size} \
    --speculative-draft-model-quantization {speculative_draft_model_quantization} \
    --mamba-scheduler-strategy extra_buffer \
    --enable-torch-compile --torch-compile-max-bs 4 \
    --num-continuous-decode-steps 2 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --host {host} \
    --port {port}
sparkrun run ~/qwen38-27b-sglang-dspark-tp2.yaml --hosts <head>,<worker1>

NVFP4 (RadixArk/Qwen3.8-27B-NVFP4) + DSpark, TP=4, SGLang

Moving the same SGLang + DSpark setup from TP=2 to TP=4 raises C1 throughput from 51.8 to 77.3 tok/s, another 49%. This is the fastest low-concurrency configuration in the study: 77.3 tok/s with 227 ms TTFT at C1. The additional nodes also keep TTFT below 1 second through C32.

C TTFT (ms) TPS Pass?
1 227 77.3 βœ“
2 247 61.9 βœ“
4 264 48.1 βœ“
8 379 27.0 βœ“
16 411 18.6 βœ“
32 579 11.0 βœ— (TPS)
recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4
runtime: sglang
container: lmsysorg/sglang@sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1
​
metadata:
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (RadixArk) + DSpark β€” SGLang TP=4 on 4x DGX Spark
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 4
  gpu_memory_utilization: 0.50
  served_model_name: qwen3.8-27b
  attention_backend: flashinfer
  speculative_algorithm: DSPARK
  speculative_draft_model_path: RadixArk/Qwen3.8-27B-DSpark
  speculative_dspark_block_size: 7
  speculative_draft_model_quantization: unquant
​
env:
  TORCHINDUCTOR_CACHE_DIR: /cache/huggingface/sglang-inductor
  HF_HOME: /cache/huggingface
  HF_HUB_CACHE: /cache/huggingface/hub
​
command: |
  python3 -m sglang.launch_server \
    --trust-remote-code \
    --model-path {model} \
    --tp-size {tensor_parallel} \
    --served-model-name {served_model_name} \
    --mem-fraction-static {gpu_memory_utilization} \
    --attention-backend {attention_backend} \
    --chunked-prefill-size 8192 \
    --disable-prefill-cuda-graph \
    --cuda-graph-max-bs 4 \
    --speculative-algorithm {speculative_algorithm} \
    --speculative-draft-model-path {speculative_draft_model_path} \
    --speculative-dspark-block-size {speculative_dspark_block_size} \
    --speculative-draft-model-quantization {speculative_draft_model_quantization} \
    --mamba-scheduler-strategy extra_buffer \
    --enable-torch-compile --torch-compile-max-bs 4 \
    --num-continuous-decode-steps 2 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --host {host} \
    --port {port}
sparkrun run ~/qwen38-27b-sglang-dspark-tp4.yaml --hosts <head>,<worker1>,<worker2>,<worker3>

What this suggests

The progression seems to make a few things fairly clear for this particular.

Reducing precision helps, but precision alone is not enough. BF16 β†’ FP8 raises C1 throughput from 4.5 to 7.9 tok/s, while NVFP4 on the official vLLM stack reaches only 9.7 despite the much smaller model footprint. In contrast, speculative decoding changes the practical result much more: MTP takes both FP8 and NVFP4 from configurations that fail at every concurrency level to configurations that meet our threshold.

The runtime and serving stack matter at least as much as adding hardware. A single-Spark SGLang + DSpark configuration reaches 36.6 tok/s at C1, ahead of the two-Spark official-vLLM result and even slightly ahead of the two-Spark eugr result at C1. Going from two to four Sparks increases C1 throughput again, from 51.8 to 77.3 tok/s, but it does not increase higher concurrency performance that much: TP=2 and TP=4 both top out at C16 in the end. Under this benchmark, the four-Spark setup primarily a choice for maximum single-user or low-concurrency speed rather than additional multi-user capacity.

So, with the criteria used here, the two-Spark SGLang + DSpark setup looks like the most balanced point in the progression: it retains much of the low-concurrency speedup, serves through C16, and avoids using two additional Sparks without increasing the maximum passing concurrency. TP=4 is reasonable when 50–77 tok/s low-concurrency generation is worth the extra nodes. On a single Spark, SGLang + DSpark is the fastest option we tested for light concurrency, while NVFP4 + MTP on vLLM is slower but behaves more predictably through C8.

Those conclusions are specific to this benchmark and our 15 tok/s / 1 s thresholds; longer contexts or different serving objectives could shift the preferred configuration.


Thanks to RadixArk for the Qwen3.8-27B NVFP4 + DSpark models, InferAct for the NVFP4 variant, @basbunarhasan for the SGLang + DSpark DGX Spark recipe, @eugr_nv for the native SM 12.1a nightly images, and the sparkrun crew for the orchestration tool.

For the same 4Γ— Spark hardware, we also tested a different scaling strategy: instead of one TP=4 group, we split the four Sparks into two independent TP=2 groups behind a round-robin router β€” effectively TP=2 + DP=2.

The trade-off is lower single-user TPS, but substantially better scaling under concurrent load. This is the only configuration in the study that passes our criteria at C32.

Benchmark β†’ SMG Router (<head>:30000, round_robin)
                 β”œβ”€β”€ TP=2 Group 1: <head>:8000 ←→ <worker1>:8000
                 └── TP=2 Group 2: <worker2>:8000 ←→ <worker3>:8000
C TTFT (ms) TPS Pass?
1 207 54.6 βœ“
2 222 52.9 βœ“
4 263 46.8 βœ“
8 321 36.0 βœ“
16 354 24.2 βœ“
32 465 16.8 βœ“
64 5972 9.7 βœ— (both)

Each replica is the same TP=2 SGLang + DSpark configuration from the earlier test. The router simply distributes requests between the two groups, so at total concurrency 2C, each TP=2 group sees roughly concurrency C.

The results line up almost exactly:

TP=2 at C TPS TP=2+DP=2 at 2C TPS
C1 51.8 C2 52.9
C2 46.4 C4 46.8
C4 36.0 C8 36.0
C8 24.9 C16 24.2
C16 16.7 C32 16.8

That is a particularly clean result. TP=2 + DP=2 at total concurrency 2C behaves almost exactly like a single TP=2 group at concurrency C.

This also explains the C32 result: with the requests split across two replicas, each group is effectively handling about C16. A single TP=2 group at C16 measured 16.7 tok/s and 460 ms TTFT; TP=2 + DP=2 at C32 measures 16.8 tok/s and 465 ms. In this benchmark, the round-robin layer therefore adds negligible observable overhead, while the second replica almost directly translates into additional concurrent capacity.

Compared with TP=4 on the same four Sparks:

C TP=4 TPS TP=2+DP=2 TPS Winner Delta
1 77.3 54.6 TP=4 +42%
2 61.9 52.9 TP=4 +17%
4 48.1 46.8 TP=4 +3%
8 27.0 36.0 DP=2 +33%
16 18.6 24.2 DP=2 +30%
32 11.0 16.8 DP=2 +53%

So the choice between TP and DP changes with concurrency. TP=4 gives 77.3 tok/s at C1, while TP=2 + DP=2 gives 54.6 β€” about 29% lower. By C4, however, the difference has almost disappeared, and somewhere between C4 and C8 the advantage crosses over. From C8 onward, the replicated setup is clearly faster.

C32 is where the difference becomes most practical: TP=4 drops to 11.0 tok/s and fails our TPS threshold, while TP=2 + DP=2 still delivers 16.8 tok/s with 465 ms TTFT and passes both criteria.

For this workload, that makes TP=4 the better use of four Sparks when the priority is maximum single-user or low-concurrency speed. If the same hardware is intended to serve more simultaneous users, two TP=2 replicas are the more effective configuration: they give up some C1 performance, but almost double the useful concurrency range from C16 to C32.

The router here is SGLang’s Model Gateway (sgl-model-gateway), running as a separate Docker container rather than through sparkrun. We use its round_robin policy; it also provides worker health checking, circuit breakers, and Prometheus metrics.

Both groups use the same SGLang recipe β€” only the description differs:

recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4
runtime: sglang
container: lmsysorg/sglang@sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1
​
metadata:
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (RadixArk) + DSpark β€” SGLang TP=2 Group 1
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 2
  gpu_memory_utilization: 0.50
  served_model_name: qwen3.8-27b
  attention_backend: flashinfer
  speculative_algorithm: DSPARK
  speculative_draft_model_path: RadixArk/Qwen3.8-27B-DSpark
  speculative_dspark_block_size: 7
  speculative_draft_model_quantization: unquant
​
env:
  TORCHINDUCTOR_CACHE_DIR: /cache/huggingface/sglang-inductor
  HF_HOME: /cache/huggingface
  HF_HUB_CACHE: /cache/huggingface/hub
​
command: |
  python3 -m sglang.launch_server \
    --trust-remote-code \
    --model-path {model} \
    --tp-size {tensor_parallel} \
    --served-model-name {served_model_name} \
    --mem-fraction-static {gpu_memory_utilization} \
    --attention-backend {attention_backend} \
    --chunked-prefill-size 8192 \
    --disable-prefill-cuda-graph \
    --cuda-graph-max-bs 4 \
    --speculative-algorithm {speculative_algorithm} \
    --speculative-draft-model-path {speculative_draft_model_path} \
    --speculative-dspark-block-size {speculative_dspark_block_size} \
    --speculative-draft-model-quantization {speculative_draft_model_quantization} \
    --mamba-scheduler-strategy extra_buffer \
    --enable-torch-compile --torch-compile-max-bs 4 \
    --num-continuous-decode-steps 2 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --host {host} \
    --port {port}

Launch sequence (all from head node):

sparkrun run ~/qwen38-27b-sglang-dspark-tp2-group1.yaml --hosts <head>,<worker1>
sparkrun run ~/qwen38-27b-sglang-dspark-tp2-group2.yaml --hosts <worker2>,<worker3>
docker run -d --name sglang-router \
  --network host \
  lmsysorg/sgl-model-gateway:latest \
  --worker-urls http://<head>:8000 http://<worker2>:8000 \
  --policy round_robin \
  --host 0.0.0.0 --port 30000 \
  --prometheus-port 10001

Can you check with DFlash2?
My results with tuned for c1 SGLang

NVFP4 (RadixArk/Qwen3.8-27B-NVFP4) + DFlash2, TP=1, SGLang

Conc 1: TTFT=217ms \[OK\] TPS=56.6 \[OK\]
Conc 2: TTFT=330ms \[OK\] TPS=54.6 \[OK\]
Conc 4: TTFT=3938ms \[X\] TPS=24.2 \[OK\]
Conc 8: TTFT=10598ms \[X\] TPS=11.2 \[X\]
Conc 16: TTFT=14102ms \[X\] TPS=13.3 \[X\]

Amazing results!

One point to consider: with our benchmarking tool, the result table has two modes β€” p90 and mean. We usually use mean, but it might be the case that your table had p90 selected, which would give a slightly larger TPS β€” the MiaAI-Lab repo’s own results show 50.9 TPS.

As for the DFlash2, can you share which model variant you used (RadixArk/Qwen3.8-27B-NVFP4 or RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead) and your exact tunings and flags? We want to make sure we’re comparing the same setup.

Retest with mean

Conc 1: TTFT=217ms [OK] TPS=48.0 [OK]
Conc 2: TTFT=306ms [OK] TPS=42.3 [OK]
Conc 4: TTFT=3251ms [X] TPS=22.1 [OK]

I use GitHub - hasso5703/dgx-spark-qwen38: Fastest measured Qwen3.8-27B config for DGX Spark (GB10): SGLang + NVFP4 + DFlash2, deterministic boots, 50 tok/s greedy median, 148 tok/s at 8 streams, 258 at 32. One command, everything pinned, lossless. Β· GitHub as base but run it with llama-swap
My llama-swap model command

    cmd: |
      /usr/bin/docker run --rm
      --name qwen38-sglang
      --gpus all
      --memory 100g
      --memory-swap 100g
      --shm-size 16g
      --network host
      --ipc=host
      -e TORCHINDUCTOR_CACHE_DIR=/cache/inductor
      -v /home/pont/.config/qwen38/sglang-cache:/cache
      -v /home/pont/.cache/huggingface:/root/.cache/huggingface
      -v /home/pont/.config/qwen38:/out
      qwen38-dflash2:v1.2.2
      python3 -m sglang.launch_server
      --trust-remote-code
      --model-path RadixArk/Qwen3.8-27B-NVFP4
      --tp-size 1
      --served-model-name qwen3.8-27b
      --mem-fraction-static 0.45
      --attention-backend flashinfer
      --chunked-prefill-size 2048
      --cuda-graph-max-bs-decode 2
      --sleep-on-idle
      --speculative-algorithm DFLASH
      --speculative-draft-model-path z-lab/Qwen3.8-27B-DFlash2
      --speculative-draft-model-revision 50307d4c4cde6860d4eee73e2547cd786fe8e8a4
      --speculative-num-draft-tokens 8
      --speculative-draft-model-quantization unquant
      --mamba-radix-cache-strategy extra_buffer
      --mamba-ssm-dtype bfloat16
      --max-mamba-cache-size 32
      --max-running-requests 2
      --torch-compile-max-bs 2
      --num-continuous-decode-steps 1
      --reasoning-parser qwen3
      --tool-call-parser qwen3_coder
      --chat-template /out/chat-template-sglang.jinja
      --host 0.0.0.0
      --port ${PORT}
      --enable-metrics
      --enable-mfu-metrics
      --enable-cache-report
      --enable-request-time-stats-logging
      --log-requests
      --log-requests-level 1
      --log-requests-format json
      --show-time-cost

Thanks to @pontostroy for the suggestions. We experimented with the DFlash2 draft model and MiaAI-Lab’s DFlash2 setup workflow a bit and matched your C1 result, then spent more time on concurrency behavior and reproducibility. Here is the final setup we came up with:

Results β€” 47.9 TPS at C1

Same benchmark methodology as the study above (CordatusAI, 128 in / 128 out, 10 rounds, pass: TTFT < 1000ms AND TPS >= 15 tok/s).

Model: RadixArk/Qwen3.8-27B-NVFP4 (packed FP4, 21.9 GB)
Draft: z-lab/Qwen3.8-27B-DFlash2 (2B, BF16)
Image: lmsysorg/sglang:qwen38-27b-dflash2 (–full, commit c14312a66 + dflash2_nvfp4_head.patch)
--mem-fraction-static 0.50, --max-running-requests 16

C TTFT (ms) TPS Pass?
1 232 47.9 βœ“
2 320 39.8 βœ“
4 343 35.6 βœ“
8 389 27.6 βœ“
16 2357 18.6 TPS βœ“, TTFT βœ—

In addition to the 47.9 TPS at C1, TTFT holds at 343ms and TPS is 27.6 at C8 β€” the setup can be used with 8 concurrency in mind comfortably.

Running with sparkrun

The full configuration is captured in a single sparkrun recipe:

1. Build the DFlash2 image (one-time, on each Spark):

git clone https://github.com/MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark.git
cd Qwen3.8-27B-SGLang-DGX-Spark
bash patch/build-dflash2-image.sh --full

2. Save the recipe and launch:

recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4
runtime: sglang
container: lmsysorg/sglang:qwen38-27b-dflash2
​
max_nodes: 1
​
metadata:
  model_dtype: nvfp4
  kv_dtype: fp8
  description: Qwen3.8-27B NVFP4 (packed FP4) + DFlash2 β€” single Spark
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.50
  served_model_name: qwen3.8-27b
  attention_backend: flashinfer
  speculative_algorithm: DFLASH
  speculative_draft_model_path: z-lab/Qwen3.8-27B-DFlash2
  speculative_num_draft_tokens: 8
​
env:
  HF_HUB_OFFLINE: "0"
  TORCHINDUCTOR_CACHE_DIR: /cache/inductor
  HF_HOME: /cache/huggingface
  HF_HUB_CACHE: /cache/huggingface/hub
​
command: |
  python3 -m sglang.launch_server \
    --trust-remote-code \
    --model-path {model} \
    --tp-size {tensor_parallel} \
    --served-model-name {served_model_name} \
    --mem-fraction-static {gpu_memory_utilization} \
    --attention-backend {attention_backend} \
    --chunked-prefill-size 8192 \
    --disable-prefill-cuda-graph \
    --kv-cache-dtype fp8_e4m3 \
    --mamba-ssm-dtype bfloat16 \
    --mamba-full-memory-ratio 4.21 \
    --mamba-radix-cache-strategy extra_buffer \
    --max-mamba-cache-size 64 \
    --max-running-requests 16 \
    --context-length 262144 \
    --speculative-algorithm {speculative_algorithm} \
    --speculative-draft-model-path {speculative_draft_model_path} \
    --speculative-draft-model-revision 50307d4c4cde6860d4eee73e2547cd786fe8e8a4 \
    --speculative-num-draft-tokens {speculative_num_draft_tokens} \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --sampling-defaults model \
    --enable-metrics \
    --enable-cache-report \
    --stream-interval 1 \
    --host {host} \
    --port {port}
sparkrun run ~/qwen38-27b-dflash2.yaml --rootful --no-follow --no-rm

Optional β€” CPU pinning:

docker update --cpuset-cpus 5-9,15-19 $(docker ps --format "{{.Names}}" --filter "ancestor=lmsysorg/sglang:qwen38-27b-dflash2")

This pins the container to GB10’s 10 Cortex-X5 performance cores, keeping the scheduler and tokenizer off the 2.8 GHz efficiency cores. We measured less than 2 TPS difference at C1 and no meaningful change at higher concurrencies, so it can be considered optional.

MiaAI lab did not update her receipe. There is an official SGlang docker image with DFlash2 merged now
lmsysorg/sglang:dev-qwen38-27b-dflash2

Also you should try RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead, with the dlash2 above, better quality

Tested prefill size with very big context
2048 always better

tool-eval-bench --backend sglang --context-size 524288 --context-pressure 0.6

Chunked Prefill Size TTFT Long-Context Prefill Peak/Stable TFLOPS Result
2048 415 s ~390–480 tok/s ~180–182 TFLOPS πŸ₯‡ Best
4096 424.1 s ~384–481 tok/s ~176–179 TFLOPS πŸ₯ˆ Good
8192 441.1 s ~383–464 tok/s ~169–172 TFLOPS πŸ₯‰ Slowest

Thanks for the heads-up!

We also tested RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead with DFlash2, and the inference performance was very similar in our benchmarks.

When you say β€œbetter quality,” I believe that you mean output quality rather than inference speed. If you have comparative evaluation results or experiences between the two models, we’d be interested to see them.

We’ll also take a look at the new lmsysorg/sglang:dev-qwen38-27b-dflash2 image.

Something I messed around with today which hasn’t been mentioned yet: B12x + DFlash2 on Unsloth’s NVFP4. I don’t run benches, and just got this up, but for C1 with TP2 and batch of 8192 am seeing prose from 30s β†’ mid 40s and code 50s β†’ low 80s. Haven’t tried against FP8 quant yet.

We used the RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead with our existing setup for some time. There seems to be a quality increase as expected, while keeping the high performance.

As for the lmsysorg/sglang:dev-qwen38-27b-dflash2, our benchmarks showed ~3-5 TPS decrease at C1. We used the official cookbook with these selections:

  • Hardware: DGX Spark
  • Model: RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead (NVFP4 quantization, BF16 lm_head)
  • Draft model: incoai/Qwen3.8-27B-DFlash2
  • Speculative decoding: DFlash2 (8 draft tokens)
  • Nodes: single

The cookbook generated the serving cell with these parameters:

  • mem-fraction-static: 0.80
  • chunked-prefill-size: 2048
  • kv-cache-dtype: fp8_e4m3
  • attention-backend: flashinfer
  • reasoning-parser: qwen3
  • tool-call-parser: qwen3_coder

We launched it as-is via sparkrun. Here is the sparkrun recipe and run command:


Recipe (~/qwen38-27b-dflash2-official.yaml):

recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead
runtime: sglang
container: lmsysorg/sglang:dev-qwen38-27b-dflash2

max_nodes: 1

metadata:
  model_dtype: nvfp4
  kv_dtype: fp8
  description: Qwen3.8-27B NVFP4 (BF16-LMHead) + DFlash2 β€” official SGLang cookbook image

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.80
  served_model_name: qwen3.8-27b

env:
  HF_HUB_OFFLINE: "0"

command: |
  python3 -m sglang.launch_server \
    --trust-remote-code \
    --model-path {model} \
    --tp-size {tensor_parallel} \
    --served-model-name {served_model_name} \
    --mem-fraction-static {gpu_memory_utilization} \
    --attention-backend flashinfer \
    --chunked-prefill-size 2048 \
    --kv-cache-dtype fp8_e4m3 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --speculative-algorithm DFLASH \
    --speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2 \
    --speculative-num-draft-tokens 8 \
    --host {host} \
    --port {port}

Run command:

sparkrun run ~/qwen38-27b-dflash2-official.yaml --rootful --no-follow --no-rm