We benchmarked Qwen3.8-27B across 11 configurations on one to four DGX Sparks to see what actually moves performance: precision, speculative decoding, runtime/image choice, and tensor parallelism.
Across the sequence, C1 generation increased from 4.5 to 77.3 tok/s. Below are the full concurrency sweeps and the exact recipes used at each step.
Stack
-
4Γ NVIDIA DGX Spark (GB10 Grace Blackwell, SM 12.1, 128 GB UMA, 273 GB/s LPDDR5X)
-
CX-7 (100 Gbps) interconnect between nodes
-
sparkrun v0.3.4
-
Benchmark: CordatusAI LLM Benchmark Tool β 128 tokens in, 128 tokens out, 10 rounds per concurrency level. Pass criteria at a concurrency: TTFT < 1000ms AND TPS >= 15 tok/s
-
Images: vLLM official (
vllm/vllm-openai:qwen38), eugr vLLM (ghcr.io/spark-arena/dgx-vllm-eugr-nightly:latest), SGLang (lmsysorg/sglang, pinned by SHA in recipes) -
Models: Qwen3.8-27B BF16 (
Qwen/Qwen3.8-27B, 55.6 GB), Qwen3.8-27B FP8 (Qwen/Qwen3.8-27B-FP8, 30.9 GB), Qwen3.8-27B NVFP4 β InferAct (Inferact/Qwen3.8-27B-NVFP4, ~18 GB), Qwen3.8-27B NVFP4 β RadixArk (RadixArk/Qwen3.8-27B-NVFP4, ~18 GB), DSpark draft model (RadixArk/Qwen3.8-27B-DSpark, 1.36B)
Progression Summary
| Step | Configuration | What changed | TPS @ C1 | TTFT @ C1 | Max C (both pass) |
|---|---|---|---|---|---|
| 1 | Qwen3.8-27B BF16, TP=1, vLLM official | β | 4.5 | 335ms | Never |
| 2 | Qwen3.8-27B BF16 + MTP, TP=1, vLLM official | +MTP n=3 | 9.9 | 611ms | Never |
| 3 | Qwen3.8-27B FP8, TP=1, vLLM official | BF16βFP8 | 7.9 | 172ms | Never |
| 4 | Qwen3.8-27B FP8 + MTP, TP=1, vLLM official | +MTP n=3 | 17.1 | 392ms | C4 |
| 5 | Qwen3.8-27B NVFP4 (InferAct), TP=1, vLLM official | FP8βNVFP4 | 9.7 | 155ms | Never |
| 6 | Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=1, vLLM official | +MTP n=3 | 18.5 | 341ms | C8 |
| 7 | Qwen3.8-27B NVFP4 (RadixArk) + DSpark, TP=1, SGLang | engine + spec decode switch | 36.6 | 246ms | C4 |
| 8 | Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=2, vLLM official | +1 Spark | 23.4 | 254ms | C8 |
| 9 | Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=2, eugr vLLM | officialβeugr | 33.0 | 311ms | C16 |
| 10 | Qwen3.8-27B NVFP4 (RadixArk) + DSpark, TP=2, SGLang | +1 Spark | 51.8 | 219ms | C16 |
| 11 | Qwen3.8-27B NVFP4 (RadixArk) + DSpark, TP=4, SGLang | +2 Sparks | 77.3 | 227ms | C16 |
Because Qwen3.8-27B is a dense model, every concurrent request multiplies against the same weight matrices. The GPU loads each weight tile once and serves all requests in a single batched GEMM. Since weight loading β the bottleneck in the memory-bandwidth-bound regime β doesnβt change with batch size, per-user TPS stays relatively flat at low concurrency, then gradually decreases as the batch grows large enough to saturate the SMs and enter the compute-bound regime. We observe this pattern across almost every run in this study. That is why most runs either never pass the concurrency threshold or pass it with a higher concurrency than C1 β once the model can serve enough tokens at C1, it holds up well under increasing concurrency. This is also why C1 TPS values β where decode is purely memory-bound β sit very close to the memory bandwidth theoretical ceiling.
BF16 (Qwen/Qwen3.8-27B), TP=1, vLLM official
The model is 55.6 GB. A simple memory-bandwidth-only upper bound is 273 / 55.6 = 4.9 tok/s; we measured 4.5 tok/s at C1. TPS stays close to that level as concurrency rises, but the configuration never reaches our 15 tok/s threshold. TTFT crosses 1 second at C8.
| C | TTFT (ms) | TPS | Pass? |
|---|---|---|---|
| 1 | 335 | 4.5 | β (TPS) |
| 2 | 610 | 4.3 | β (TPS) |
| 4 | 797 | 4.2 | β (TPS) |
| 8 | 1064 | 4.1 | β (both) |
| 16 | 1670 | 3.8 | β (both) |
recipe_version: "2"
model: Qwen/Qwen3.8-27B
runtime: vllm
container: vllm/vllm-openai:qwen38
β
executor_config:
entrypoint: ""
β
max_nodes: 1
β
metadata:
kv_dtype: fp8
description: Qwen3.8-27B BF16, TP=1, no MTP
β
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
gpu_memory_utilization: 0.8
β
command: |
vllm serve Qwen/Qwen3.8-27B \
--host {host} \
--port {port} \
--tensor-parallel-size {tensor_parallel} \
--gpu-memory-utilization {gpu_memory_utilization} \
--max-model-len 262144 \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
sparkrun run ~/qwen38-27b-bf16.yaml
BF16 (Qwen/Qwen3.8-27B) + MTP, TP=1, vLLM official
Adding MTP with three speculative tokens raises C1 TPS from 4.5 to 9.9, a 2.2Γ increase. The gain holds through the lower concurrency levels, but it is still not enough to reach 15 tok/s.
| C | TTFT (ms) | TPS | Pass? |
|---|---|---|---|
| 1 | 611 | 9.9 | β (TPS) |
| 2 | 867 | 10.0 | β (TPS) |
| 4 | 916 | 10.0 | β (TPS) |
| 8 | 1067 | 9.0 | β (both) |
| 16 | 1250 | 7.9 | β (both) |
recipe_version: "2"
model: Qwen/Qwen3.8-27B
runtime: vllm
container: vllm/vllm-openai:qwen38
β
executor_config:
entrypoint: ""
β
max_nodes: 1
β
metadata:
kv_dtype: fp8
description: Qwen3.8-27B BF16 + MTP, TP=1
β
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
gpu_memory_utilization: 0.8
β
command: |
vllm serve Qwen/Qwen3.8-27B \
--host {host} \
--port {port} \
--tensor-parallel-size {tensor_parallel} \
--gpu-memory-utilization {gpu_memory_utilization} \
--max-model-len 262144 \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
sparkrun run ~/qwen38-27b-bf16-mtp.yaml
FP8 (Qwen/Qwen3.8-27B-FP8), TP=1, vLLM official
FP8 halves the model to 30.9 GB. Theoretical ceiling becomes 273 / 30.9 = 8.8 tok/s. We hit 7.9 β 90% of theoretical max. TTFT performance almost doubles as expected (335 β 172ms) β half the weight loading during prefill. 7.9 TPS is still below 15.
| C | TTFT (ms) | TPS | Pass? |
|---|---|---|---|
| 1 | 172 | 7.9 | β (TPS) |
| 2 | 322 | 7.8 | β (TPS) |
| 4 | 435 | 7.6 | β (TPS) |
| 8 | 769 | 7.2 | β (TPS) |
| 16 | 1401 | 6.4 | β (both) |
| 32 | 3332 | 5.0 | β (both) |
recipe_version: "2"
model: Qwen/Qwen3.8-27B-FP8
runtime: vllm
container: vllm/vllm-openai:qwen38
β
executor_config:
entrypoint: ""
β
max_nodes: 1
β
metadata:
kv_dtype: fp8
description: Qwen3.8-27B FP8, TP=1, no MTP
β
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
gpu_memory_utilization: 0.8
β
command: |
vllm serve Qwen/Qwen3.8-27B-FP8 \
--host {host} \
--port {port} \
--tensor-parallel-size {tensor_parallel} \
--gpu-memory-utilization {gpu_memory_utilization} \
--max-model-len 262144 \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
sparkrun run ~/qwen38-27b-fp8.yaml
FP8 (Qwen/Qwen3.8-27B-FP8) + MTP, TP=1, vLLM official
This is the first configuration to pass our criteria. MTP raises C1 throughput from 7.9 to 17.1 tok/s, a 2.16Γ increase, very close to the 2.2Γ gain seen with BF16. C1 through C4 pass; at C8, TTFT is still comfortably below 1 second but TPS falls just below 15.
| C | TTFT (ms) | TPS | Pass? |
|---|---|---|---|
| 1 | 392 | 17.1 | β |
| 2 | 557 | 17.1 | β |
| 4 | 565 | 16.1 | β |
| 8 | 672 | 14.4 | β (TPS) |
| 16 | 999 | 11.7 | β (TPS) |
| 32 | 1504 | 8.6 | β (both) |
recipe_version: "2"
model: Qwen/Qwen3.8-27B-FP8
runtime: vllm
container: vllm/vllm-openai:qwen38
β
executor_config:
entrypoint: ""
β
max_nodes: 1
β
metadata:
kv_dtype: fp8
description: Qwen3.8-27B FP8 + MTP, TP=1
β
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
gpu_memory_utilization: 0.8
β
command: |
vllm serve Qwen/Qwen3.8-27B-FP8 \
--host {host} \
--port {port} \
--tensor-parallel-size {tensor_parallel} \
--gpu-memory-utilization {gpu_memory_utilization} \
--max-model-len 262144 \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
sparkrun run ~/qwen38-27b-fp8-mtp.yaml
NVFP4 (Inferact/Qwen3.8-27B-NVFP4), TP=1, vLLM official
NVFP4 reduces the model to roughly 18 GB, which gives a simple bandwidth-only estimate of about 15.2 tok/s. The measured C1 result is 9.7 tok/s, 64% of theoretical max. Thus, unlike BF16 and FP8, model footprint alone is no longer a good predictor of realized decode throughput on this stack. The new bottleneck might be the vLLMβs NVFP4 kernel efficiency on SM 12.
| C | TTFT (ms) | TPS | Pass? |
|---|---|---|---|
| 1 | 155 | 9.7 | β (TPS) |
| 2 | 270 | 9.6 | β (TPS) |
| 4 | 333 | 9.3 | β (TPS) |
| 8 | 520 | 8.8 | β (TPS) |
| 16 | 899 | 7.7 | β (TPS) |
| 32 | 1502 | 6.2 | β (both) |
recipe_version: "2"
model: Inferact/Qwen3.8-27B-NVFP4
runtime: vllm
container: vllm/vllm-openai:qwen38
executor_config:
entrypoint: ""
max_nodes: 1
metadata:
kv_dtype: fp8
model_dtype: nvfp4
description: Qwen3.8-27B NVFP4 (InferAct), TP=1, no MTP
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
gpu_memory_utilization: 0.8
command: |
vllm serve Inferact/Qwen3.8-27B-NVFP4 \
--host {host} \
--port {port} \
--tensor-parallel-size {tensor_parallel} \
--gpu-memory-utilization {gpu_memory_utilization} \
--max-model-len 262144 \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
sparkrun run ~/qwen38-27b-nvfp4-inferact.yaml
NVFP4 (Inferact/Qwen3.8-27B-NVFP4) + MTP, TP=1, vLLM official
MTP raises C1 TPS from 9.7 to 18.5, a 1.91Γ increase. It passes through C8, with TTFT still at 603 ms there. At C16, TTFT remains below the threshold but TPS falls to 12.5.
| C | TTFT (ms) | TPS | Pass? |
|---|---|---|---|
| 1 | 341 | 18.5 | β |
| 2 | 465 | 18.8 | β |
| 4 | 517 | 17.3 | β |
| 8 | 603 | 15.3 | β |
| 16 | 767 | 12.5 | β (TPS) |
| 32 | 1095 | 9.2 | β (both) |
recipe_version: "2"
model: Inferact/Qwen3.8-27B-NVFP4
runtime: vllm
container: vllm/vllm-openai:qwen38
executor_config:
entrypoint: ""
max_nodes: 1
metadata:
kv_dtype: fp8
model_dtype: nvfp4
description: Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=1
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
gpu_memory_utilization: 0.8
command: |
vllm serve Inferact/Qwen3.8-27B-NVFP4 \
--host {host} \
--port {port} \
--tensor-parallel-size {tensor_parallel} \
--gpu-memory-utilization {gpu_memory_utilization} \
--max-model-len 262144 \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
sparkrun run ~/qwen38-27b-nvfp4-inferact-mtp.yaml
NVFP4 (RadixArk/Qwen3.8-27B-NVFP4) + DSpark, TP=1, SGLang
This is the strongest single-Spark configuration in the study. This single Spark configuration will beat the best 2-Spark vLLM run at C1βC2: 36.6 TPS vs 33.0 TPS, with half the node count. This success is the combination of the SGLang image NVFP4 optimization and DSpark (proposes 7 tokens per speculation) high acceptance rate.
The main weakness appears at higher concurrency. TTFT jumps from 394 ms at C4 to 1368 ms at C8. The recipe uses --cuda-graph-max-bs 4, so C8 and above are outside that configured CUDA-graph batch-size range. That coincides with the observed jump, although this benchmark alone does not isolate the exact causation.
| C | TTFT (ms) | TPS | Pass? |
|---|---|---|---|
| 1 | 246 | 36.6 | β |
| 2 | 349 | 31.3 | β |
| 4 | 394 | 25.6 | β |
| 8 | 1368 | 17.8 | β (TTFT) |
| 16 | 8554 | 9.2 | β (both) |
| 32 | 23045 | 4.7 | β (both) |
recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4
runtime: sglang
container: lmsysorg/sglang@sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1
max_nodes: 1
metadata:
model_dtype: nvfp4
description: Qwen3.8-27B NVFP4 (RadixArk) + DSpark β SGLang on DGX Spark
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
gpu_memory_utilization: 0.50
served_model_name: qwen3.8-27b
attention_backend: flashinfer
speculative_algorithm: DSPARK
speculative_draft_model_path: RadixArk/Qwen3.8-27B-DSpark
speculative_dspark_block_size: 7
speculative_draft_model_quantization: unquant
env:
TORCHINDUCTOR_CACHE_DIR: /cache/inductor
HF_HOME: /cache/huggingface
HF_HUB_CACHE: /cache/huggingface/hub
command: |
python3 -m sglang.launch_server \
--trust-remote-code \
--model-path {model} \
--tp-size {tensor_parallel} \
--served-model-name {served_model_name} \
--mem-fraction-static {gpu_memory_utilization} \
--attention-backend {attention_backend} \
--chunked-prefill-size 8192 \
--disable-prefill-cuda-graph \
--cuda-graph-max-bs 4 \
--speculative-algorithm {speculative_algorithm} \
--speculative-draft-model-path {speculative_draft_model_path} \
--speculative-dspark-block-size {speculative_dspark_block_size} \
--speculative-draft-model-quantization {speculative_draft_model_quantization} \
--mamba-scheduler-strategy extra_buffer \
--enable-torch-compile --torch-compile-max-bs 4 \
--num-continuous-decode-steps 2 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--host {host} \
--port {port}
sparkrun run ~/qwen38-27b-sglang-dspark.yaml
NVFP4 (Inferact/Qwen3.8-27B-NVFP4) + MTP, TP=2, vLLM official
Moving the same InferAct + MTP setup from TP=1 to TP=2 raises C1 TPS from 18.5 to 23.4, about 26%. The scaling is useful but far from 2Γ. Tensor parallelism reduces the amount of sharded model work handled by each GPU, while also introducing communication on the decode path, so linear scaling should not be expected. TTFT improves from 341 ms to 254 ms at C1. The configuration still passes through C8 and falls below the TPS threshold at C16.
| C | TTFT (ms) | TPS | Pass? |
|---|---|---|---|
| 1 | 254 | 23.4 | β |
| 2 | 423 | 22.4 | β |
| 4 | 454 | 21.3 | β |
| 8 | 570 | 17.0 | β |
| 16 | 699 | 12.8 | β (TPS) |
| 32 | 859 | 10.7 | β (TPS) |
recipe_version: "2"
model: Inferact/Qwen3.8-27B-NVFP4
runtime: vllm-distributed
container: vllm/vllm-openai:qwen38
executor_config:
entrypoint: ""
metadata:
kv_dtype: fp8
model_dtype: nvfp4
description: Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=2, official vLLM image
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 2
gpu_memory_utilization: 0.75
command: |
vllm serve Inferact/Qwen3.8-27B-NVFP4 \
--host {host} \
--port {port} \
--tensor-parallel-size {tensor_parallel} \
--gpu-memory-utilization {gpu_memory_utilization} \
--max-model-len 262144 \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
sparkrun run ~/qwen38-27b-inferact-nvfp4-tp2.yaml --hosts <head>,<worker1>
NVFP4 (Inferact/Qwen3.8-27B-NVFP4) + MTP, TP=2, eugr vLLM
The eugr recipe reaches 33.0 tok/s at C1 versus 23.4 with the official vLLM configuration and remains ahead throughout the concurrency sweep. It is also the first vLLM setup here to satisfy both criteria at C16.
The native SM 12.1a support in the eugr stack is a plausible contributor to the improvement. In addition, the eugr recipe changes several serving parameters, including memory utilization, maximum model length, attention backend, load format, chunked prefill, and asynchronous scheduling.
| C | TTFT (ms) | TPS | Pass? |
|---|---|---|---|
| 1 | 311 | 33.0 | β |
| 2 | 467 | 31.8 | β |
| 4 | 540 | 29.5 | β |
| 8 | 599 | 24.0 | β |
| 16 | 831 | 17.7 | β |
| 32 | 1345 | 13.4 | β (both) |
recipe_version: '2'
name: Qwen3.8-27B-NVFP4-MTP-TP2
model: Inferact/Qwen3.8-27B-NVFP4
container: ghcr.io/spark-arena/dgx-vllm-eugr-nightly:latest
runtime: vllm-distributed
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 2
gpu_memory_utilization: 0.55
max_model_len: 131072
max_num_batched_tokens: 8192
served_model_name: Qwen3.8-27B
attention_backend: flashinfer
tool_call_parser: qwen3_coder
pipeline_parallel: 1
load_format: fastsafetensors
kv_cache_dtype: fp8
reasoning_parser: qwen3
speculative_config: '{"method":"mtp","num_speculative_tokens":3}'
override_generation_config: '{"temperature":1.0,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
env:
VLLM_MARLIN_USE_ATOMIC_ADD: '1'
TORCH_MATMUL_PRECISION: high
FLASHINFER_DISABLE_VERSION_CHECK: '1'
NVIDIA_FORWARD_COMPAT: '1'
VLLM_HTTP_TIMEOUT_KEEP_ALIVE: '600'
command: |
vllm serve {model} \
--served-model-name {served_model_name} \
--host {host} \
--port {port} \
-tp {tensor_parallel} \
-pp {pipeline_parallel} \
--trust-remote-code \
--language-model-only \
--kv-cache-dtype {kv_cache_dtype} \
--gpu-memory-utilization {gpu_memory_utilization} \
--max-model-len {max_model_len} \
--max-num-batched-tokens {max_num_batched_tokens} \
--enable-prefix-caching \
--enable-chunked-prefill \
--async-scheduling \
--load-format {load_format} \
--attention-backend {attention_backend} \
--enable-auto-tool-choice \
--tool-call-parser {tool_call_parser} \
--reasoning-parser {reasoning_parser} \
--speculative-config '{speculative_config}' \
--override-generation-config '{override_generation_config}'
metadata:
description: Qwen3.8-27B NVFP4 (Inferact) + MTP β TP=2 dual Spark, eugr image
model_params: 27B
model_dtype: nvfp4
kv_dtype: fp8
quantization: nvfp4
quant_bits: 4
sparkrun run ~/qwen38-27b-inferact-nvfp4-tp2-eugr.yaml --hosts <head>,<worker1>
NVFP4 (RadixArk/Qwen3.8-27B-NVFP4) + DSpark, TP=2, SGLang
This is a cleaner scaling comparison with the previous SGLang run: moving from TP=1 to TP=2 raises C1 TPS from 36.6 to 51.8, about 42%. The larger gain is not limited to C1. The configuration remains above both thresholds through C16, whereas the single-Spark run fails at C8 because of TTFT.
| C | TTFT (ms) | TPS | Pass? |
|---|---|---|---|
| 1 | 219 | 51.8 | β |
| 2 | 286 | 46.4 | β |
| 4 | 324 | 36.0 | β |
| 8 | 375 | 24.9 | β |
| 16 | 460 | 16.7 | β |
| 32 | 5682 | 9.7 | β (both) |
recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4
runtime: sglang
container: lmsysorg/sglang@sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1
metadata:
model_dtype: nvfp4
description: Qwen3.8-27B NVFP4 (RadixArk) + DSpark β SGLang TP=2 on 2x DGX Spark
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 2
gpu_memory_utilization: 0.50
served_model_name: qwen3.8-27b
attention_backend: flashinfer
speculative_algorithm: DSPARK
speculative_draft_model_path: RadixArk/Qwen3.8-27B-DSpark
speculative_dspark_block_size: 7
speculative_draft_model_quantization: unquant
env:
TORCHINDUCTOR_CACHE_DIR: /cache/huggingface/sglang-inductor
HF_HOME: /cache/huggingface
HF_HUB_CACHE: /cache/huggingface/hub
command: |
python3 -m sglang.launch_server \
--trust-remote-code \
--model-path {model} \
--tp-size {tensor_parallel} \
--served-model-name {served_model_name} \
--mem-fraction-static {gpu_memory_utilization} \
--attention-backend {attention_backend} \
--chunked-prefill-size 8192 \
--disable-prefill-cuda-graph \
--cuda-graph-max-bs 4 \
--speculative-algorithm {speculative_algorithm} \
--speculative-draft-model-path {speculative_draft_model_path} \
--speculative-dspark-block-size {speculative_dspark_block_size} \
--speculative-draft-model-quantization {speculative_draft_model_quantization} \
--mamba-scheduler-strategy extra_buffer \
--enable-torch-compile --torch-compile-max-bs 4 \
--num-continuous-decode-steps 2 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--host {host} \
--port {port}
sparkrun run ~/qwen38-27b-sglang-dspark-tp2.yaml --hosts <head>,<worker1>
NVFP4 (RadixArk/Qwen3.8-27B-NVFP4) + DSpark, TP=4, SGLang
Moving the same SGLang + DSpark setup from TP=2 to TP=4 raises C1 throughput from 51.8 to 77.3 tok/s, another 49%. This is the fastest low-concurrency configuration in the study: 77.3 tok/s with 227 ms TTFT at C1. The additional nodes also keep TTFT below 1 second through C32.
| C | TTFT (ms) | TPS | Pass? |
|---|---|---|---|
| 1 | 227 | 77.3 | β |
| 2 | 247 | 61.9 | β |
| 4 | 264 | 48.1 | β |
| 8 | 379 | 27.0 | β |
| 16 | 411 | 18.6 | β |
| 32 | 579 | 11.0 | β (TPS) |
recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4
runtime: sglang
container: lmsysorg/sglang@sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1
β
metadata:
model_dtype: nvfp4
description: Qwen3.8-27B NVFP4 (RadixArk) + DSpark β SGLang TP=4 on 4x DGX Spark
β
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 4
gpu_memory_utilization: 0.50
served_model_name: qwen3.8-27b
attention_backend: flashinfer
speculative_algorithm: DSPARK
speculative_draft_model_path: RadixArk/Qwen3.8-27B-DSpark
speculative_dspark_block_size: 7
speculative_draft_model_quantization: unquant
β
env:
TORCHINDUCTOR_CACHE_DIR: /cache/huggingface/sglang-inductor
HF_HOME: /cache/huggingface
HF_HUB_CACHE: /cache/huggingface/hub
β
command: |
python3 -m sglang.launch_server \
--trust-remote-code \
--model-path {model} \
--tp-size {tensor_parallel} \
--served-model-name {served_model_name} \
--mem-fraction-static {gpu_memory_utilization} \
--attention-backend {attention_backend} \
--chunked-prefill-size 8192 \
--disable-prefill-cuda-graph \
--cuda-graph-max-bs 4 \
--speculative-algorithm {speculative_algorithm} \
--speculative-draft-model-path {speculative_draft_model_path} \
--speculative-dspark-block-size {speculative_dspark_block_size} \
--speculative-draft-model-quantization {speculative_draft_model_quantization} \
--mamba-scheduler-strategy extra_buffer \
--enable-torch-compile --torch-compile-max-bs 4 \
--num-continuous-decode-steps 2 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--host {host} \
--port {port}
sparkrun run ~/qwen38-27b-sglang-dspark-tp4.yaml --hosts <head>,<worker1>,<worker2>,<worker3>
What this suggests
The progression seems to make a few things fairly clear for this particular.
Reducing precision helps, but precision alone is not enough. BF16 β FP8 raises C1 throughput from 4.5 to 7.9 tok/s, while NVFP4 on the official vLLM stack reaches only 9.7 despite the much smaller model footprint. In contrast, speculative decoding changes the practical result much more: MTP takes both FP8 and NVFP4 from configurations that fail at every concurrency level to configurations that meet our threshold.
The runtime and serving stack matter at least as much as adding hardware. A single-Spark SGLang + DSpark configuration reaches 36.6 tok/s at C1, ahead of the two-Spark official-vLLM result and even slightly ahead of the two-Spark eugr result at C1. Going from two to four Sparks increases C1 throughput again, from 51.8 to 77.3 tok/s, but it does not increase higher concurrency performance that much: TP=2 and TP=4 both top out at C16 in the end. Under this benchmark, the four-Spark setup primarily a choice for maximum single-user or low-concurrency speed rather than additional multi-user capacity.
So, with the criteria used here, the two-Spark SGLang + DSpark setup looks like the most balanced point in the progression: it retains much of the low-concurrency speedup, serves through C16, and avoids using two additional Sparks without increasing the maximum passing concurrency. TP=4 is reasonable when 50β77 tok/s low-concurrency generation is worth the extra nodes. On a single Spark, SGLang + DSpark is the fastest option we tested for light concurrency, while NVFP4 + MTP on vLLM is slower but behaves more predictably through C8.
Those conclusions are specific to this benchmark and our 15 tok/s / 1 s thresholds; longer contexts or different serving objectives could shift the preferred configuration.
Thanks to RadixArk for the Qwen3.8-27B NVFP4 + DSpark models, InferAct for the NVFP4 variant, @basbunarhasan for the SGLang + DSpark DGX Spark recipe, @eugr_nv for the native SM 12.1a nightly images, and the sparkrun crew for the orchestration tool.