# Comprehensive Qwen3.8-27B Study on DGX Sparks: Quantization, Speculative Decoding, and TP/DP Scaling

**URL:** <https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102>\
**Category:** DGX Spark / GB10 Projects\
**Created:** [August 24, 2026, 9:18am UTC](https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102 "2026-08-24T09:18:29Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![emretoktas\_openzeka](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/emretoktas_openzeka/32/532000_2.png) [@emretoktas\_openzeka](https://forums.developer.nvidia.com/u/emretoktas_openzeka)\
**Post date:** [August 24, 2026, 9:18am UTC](https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102/1 "2026-08-24T09:18:30Z")

</div>

We benchmarked Qwen3.8-27B across 11 configurations on one to four DGX Sparks to see what actually moves performance: precision, speculative decoding, runtime/image choice, and tensor parallelism.

Across the sequence, C1 generation increased from 4.5 to 77.3 tok/s. Below are the full concurrency sweeps and the exact recipes used at each step.

## **Stack**

- 4× NVIDIA DGX Spark (GB10 Grace Blackwell, SM 12.1, 128 GB UMA, 273 GB/s LPDDR5X)

- CX-7 (100 Gbps) interconnect between nodes

- [sparkrun](https://github.com/spark-arena/sparkrun) v0.3.4

- Benchmark: [CordatusAI LLM Benchmark Tool](https://github.com/CordatusAI/llm-benchmark) — 128 tokens in, 128 tokens out, 10 rounds per concurrency level. Pass criteria at a concurrency: TTFT \< 1000ms AND TPS \>= 15 tok/s

- Images: vLLM official (`vllm/vllm-openai:qwen38`), eugr vLLM (`ghcr.io/spark-arena/dgx-vllm-eugr-nightly:latest`), SGLang (`lmsysorg/sglang`, pinned by SHA in recipes)

- Models: Qwen3.8-27B BF16 (`Qwen/Qwen3.8-27B`, 55.6 GB), Qwen3.8-27B FP8 (`Qwen/Qwen3.8-27B-FP8`, 30.9 GB), Qwen3.8-27B NVFP4 — InferAct (`Inferact/Qwen3.8-27B-NVFP4`, ~18 GB), Qwen3.8-27B NVFP4 — RadixArk (`RadixArk/Qwen3.8-27B-NVFP4`, ~18 GB), DSpark draft model (`RadixArk/Qwen3.8-27B-DSpark`, 1.36B)

## **Progression Summary**

| **Step** | **Configuration** | **What changed** | **TPS @ C1** | **TTFT @ C1** | **Max C (both pass)** |
| --- | --- | --- | --- | --- | --- |
| 1 | Qwen3.8-27B BF16, TP=1, vLLM official | — | 4.5 | 335ms | Never |
| 2 | Qwen3.8-27B BF16 + MTP, TP=1, vLLM official | +MTP n=3 | 9.9 | 611ms | Never |
| 3 | Qwen3.8-27B FP8, TP=1, vLLM official | BF16→FP8 | 7.9 | 172ms | Never |
| 4 | Qwen3.8-27B FP8 + MTP, TP=1, vLLM official | +MTP n=3 | 17.1 | 392ms | C4 |
| 5 | Qwen3.8-27B NVFP4 (InferAct), TP=1, vLLM official | FP8→NVFP4 | 9.7 | 155ms | Never |
| 6 | Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=1, vLLM official | +MTP n=3 | 18.5 | 341ms | C8 |
| 7 | Qwen3.8-27B NVFP4 (RadixArk) + DSpark, TP=1, SGLang | engine + spec decode switch | 36.6 | 246ms | C4 |
| 8 | Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=2, vLLM official | +1 Spark | 23.4 | 254ms | C8 |
| 9 | Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=2, eugr vLLM | official→eugr | 33.0 | 311ms | C16 |
| 10 | Qwen3.8-27B NVFP4 (RadixArk) + DSpark, TP=2, SGLang | +1 Spark | 51.8 | 219ms | C16 |
| 11 | Qwen3.8-27B NVFP4 (RadixArk) + DSpark, TP=4, SGLang | +2 Sparks | 77.3 | 227ms | C16 |

Because Qwen3.8-27B is a dense model, every concurrent request multiplies against the same weight matrices. The GPU loads each weight tile once and serves all requests in a single batched GEMM. Since weight loading — the bottleneck in the memory-bandwidth-bound regime — doesn’t change with batch size, per-user TPS stays relatively flat at low concurrency, then gradually decreases as the batch grows large enough to saturate the SMs and enter the compute-bound regime. We observe this pattern across almost every run in this study. That is why most runs either never pass the concurrency threshold or pass it with a higher concurrency than C1 — once the model can serve enough tokens at C1, it holds up well under increasing concurrency. This is also why C1 TPS values — where decode is purely memory-bound — sit very close to the memory bandwidth theoretical ceiling.

* * *

## **BF16 (Qwen/Qwen3.8-27B), TP=1, vLLM official**

The model is 55.6 GB. A simple memory-bandwidth-only upper bound is 273 / 55.6 = 4.9 tok/s; we measured 4.5 tok/s at C1. TPS stays close to that level as concurrency rises, but the configuration never reaches our 15 tok/s threshold. TTFT crosses 1 second at C8.

| **C** | **TTFT (ms)** | **TPS** | **Pass?** |
| --- | --- | --- | --- |
| 1 | 335 | 4.5 | ✗ (TPS) |
| 2 | 610 | 4.3 | ✗ (TPS) |
| 4 | 797 | 4.2 | ✗ (TPS) |
| 8 | 1064 | 4.1 | ✗ (both) |
| 16 | 1670 | 3.8 | ✗ (both) |

```auto
recipe_version: "2"
model: Qwen/Qwen3.8-27B
runtime: vllm
container: vllm/vllm-openai:qwen38
​
executor_config:
  entrypoint: ""
​
max_nodes: 1
​
metadata:
  kv_dtype: fp8
  description: Qwen3.8-27B BF16, TP=1, no MTP
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.8
​
command: |
  vllm serve Qwen/Qwen3.8-27B \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder

```

```auto
sparkrun run ~/qwen38-27b-bf16.yaml

```

* * *

## **BF16 (Qwen/Qwen3.8-27B) + MTP, TP=1, vLLM official**

Adding MTP with three speculative tokens raises C1 TPS from 4.5 to 9.9, a 2.2× increase. The gain holds through the lower concurrency levels, but it is still not enough to reach 15 tok/s.

| **C** | **TTFT (ms)** | **TPS** | **Pass?** |
| --- | --- | --- | --- |
| 1 | 611 | 9.9 | ✗ (TPS) |
| 2 | 867 | 10.0 | ✗ (TPS) |
| 4 | 916 | 10.0 | ✗ (TPS) |
| 8 | 1067 | 9.0 | ✗ (both) |
| 16 | 1250 | 7.9 | ✗ (both) |

```auto
recipe_version: "2"
model: Qwen/Qwen3.8-27B
runtime: vllm
container: vllm/vllm-openai:qwen38
​
executor_config:
  entrypoint: ""
​
max_nodes: 1
​
metadata:
  kv_dtype: fp8
  description: Qwen3.8-27B BF16 + MTP, TP=1
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.8
​
command: |
  vllm serve Qwen/Qwen3.8-27B \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

```

```auto
sparkrun run ~/qwen38-27b-bf16-mtp.yaml

```

* * *

## **FP8 (Qwen/Qwen3.8-27B-FP8), TP=1, vLLM official**

FP8 halves the model to 30.9 GB. Theoretical ceiling becomes 273 / 30.9 = 8.8 tok/s. We hit 7.9 — 90% of theoretical max. TTFT performance almost doubles as expected (335 → 172ms) — half the weight loading during prefill. 7.9 TPS is still below 15.

| **C** | **TTFT (ms)** | **TPS** | **Pass?** |
| --- | --- | --- | --- |
| 1 | 172 | 7.9 | ✗ (TPS) |
| 2 | 322 | 7.8 | ✗ (TPS) |
| 4 | 435 | 7.6 | ✗ (TPS) |
| 8 | 769 | 7.2 | ✗ (TPS) |
| 16 | 1401 | 6.4 | ✗ (both) |
| 32 | 3332 | 5.0 | ✗ (both) |

```auto
recipe_version: "2"
model: Qwen/Qwen3.8-27B-FP8
runtime: vllm
container: vllm/vllm-openai:qwen38
​
executor_config:
  entrypoint: ""
​
max_nodes: 1
​
metadata:
  kv_dtype: fp8
  description: Qwen3.8-27B FP8, TP=1, no MTP
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.8
​
command: |
  vllm serve Qwen/Qwen3.8-27B-FP8 \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder

```

```auto
sparkrun run ~/qwen38-27b-fp8.yaml

```

* * *

## **FP8 (Qwen/Qwen3.8-27B-FP8) + MTP, TP=1, vLLM official**

This is the first configuration to pass our criteria. MTP raises C1 throughput from 7.9 to 17.1 tok/s, a 2.16× increase, very close to the 2.2× gain seen with BF16. C1 through C4 pass; at C8, TTFT is still comfortably below 1 second but TPS falls just below 15.

| **C** | **TTFT (ms)** | **TPS** | **Pass?** |
| --- | --- | --- | --- |
| 1 | 392 | 17.1 | ✓ |
| 2 | 557 | 17.1 | ✓ |
| 4 | 565 | 16.1 | ✓ |
| 8 | 672 | 14.4 | ✗ (TPS) |
| 16 | 999 | 11.7 | ✗ (TPS) |
| 32 | 1504 | 8.6 | ✗ (both) |

```auto
recipe_version: "2"
model: Qwen/Qwen3.8-27B-FP8
runtime: vllm
container: vllm/vllm-openai:qwen38
​
executor_config:
  entrypoint: ""
​
max_nodes: 1
​
metadata:
  kv_dtype: fp8
  description: Qwen3.8-27B FP8 + MTP, TP=1
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.8
​
command: |
  vllm serve Qwen/Qwen3.8-27B-FP8 \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

```

```auto
sparkrun run ~/qwen38-27b-fp8-mtp.yaml

```

* * *

## **NVFP4 (Inferact/Qwen3.8-27B-NVFP4), TP=1, vLLM official**

NVFP4 reduces the model to roughly 18 GB, which gives a simple bandwidth-only estimate of about 15.2 tok/s. The measured C1 result is 9.7 tok/s, 64% of theoretical max. Thus, unlike BF16 and FP8, model footprint alone is no longer a good predictor of realized decode throughput on this stack. The new bottleneck might be the vLLM’s NVFP4 kernel efficiency on SM 12.

| **C** | **TTFT (ms)** | **TPS** | **Pass?** |
| --- | --- | --- | --- |
| 1 | 155 | 9.7 | ✗ (TPS) |
| 2 | 270 | 9.6 | ✗ (TPS) |
| 4 | 333 | 9.3 | ✗ (TPS) |
| 8 | 520 | 8.8 | ✗ (TPS) |
| 16 | 899 | 7.7 | ✗ (TPS) |
| 32 | 1502 | 6.2 | ✗ (both) |

```auto
recipe_version: "2"
model: Inferact/Qwen3.8-27B-NVFP4
runtime: vllm
container: vllm/vllm-openai:qwen38

executor_config:
  entrypoint: ""

max_nodes: 1

metadata:
  kv_dtype: fp8
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (InferAct), TP=1, no MTP

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.8

command: |
  vllm serve Inferact/Qwen3.8-27B-NVFP4 \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder

```

```auto
sparkrun run ~/qwen38-27b-nvfp4-inferact.yaml

```

* * *

## **NVFP4 (Inferact/Qwen3.8-27B-NVFP4) + MTP, TP=1, vLLM official**

MTP raises C1 TPS from 9.7 to 18.5, a 1.91× increase. It passes through C8, with TTFT still at 603 ms there. At C16, TTFT remains below the threshold but TPS falls to 12.5.

| **C** | **TTFT (ms)** | **TPS** | **Pass?** |
| --- | --- | --- | --- |
| 1 | 341 | 18.5 | ✓ |
| 2 | 465 | 18.8 | ✓ |
| 4 | 517 | 17.3 | ✓ |
| 8 | 603 | 15.3 | ✓ |
| 16 | 767 | 12.5 | ✗ (TPS) |
| 32 | 1095 | 9.2 | ✗ (both) |

```auto
recipe_version: "2"
model: Inferact/Qwen3.8-27B-NVFP4
runtime: vllm
container: vllm/vllm-openai:qwen38

executor_config:
  entrypoint: ""

max_nodes: 1

metadata:
  kv_dtype: fp8
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=1

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.8

command: |
  vllm serve Inferact/Qwen3.8-27B-NVFP4 \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

```

```auto
sparkrun run ~/qwen38-27b-nvfp4-inferact-mtp.yaml

```

* * *

## **NVFP4 (RadixArk/Qwen3.8-27B-NVFP4) + DSpark, TP=1, SGLang**

This is the strongest single-Spark configuration in the study. This single Spark configuration will beat the best 2-Spark vLLM run at C1–C2: 36.6 TPS vs 33.0 TPS, with half the node count. This success is the combination of the SGLang image NVFP4 optimization and DSpark (proposes 7 tokens per speculation) high acceptance rate.

The main weakness appears at higher concurrency. TTFT jumps from 394 ms at C4 to 1368 ms at C8. The recipe uses `--cuda-graph-max-bs 4`, so C8 and above are outside that configured CUDA-graph batch-size range. That coincides with the observed jump, although this benchmark alone does not isolate the exact causation.

| **C** | **TTFT (ms)** | **TPS** | **Pass?** |
| --- | --- | --- | --- |
| 1 | 246 | 36.6 | ✓ |
| 2 | 349 | 31.3 | ✓ |
| 4 | 394 | 25.6 | ✓ |
| 8 | 1368 | 17.8 | ✗ (TTFT) |
| 16 | 8554 | 9.2 | ✗ (both) |
| 32 | 23045 | 4.7 | ✗ (both) |

```auto
recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4
runtime: sglang
container: lmsysorg/sglang@sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1

max_nodes: 1

metadata:
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (RadixArk) + DSpark — SGLang on DGX Spark

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.50
  served_model_name: qwen3.8-27b
  attention_backend: flashinfer
  speculative_algorithm: DSPARK
  speculative_draft_model_path: RadixArk/Qwen3.8-27B-DSpark
  speculative_dspark_block_size: 7
  speculative_draft_model_quantization: unquant

env:
  TORCHINDUCTOR_CACHE_DIR: /cache/inductor
  HF_HOME: /cache/huggingface
  HF_HUB_CACHE: /cache/huggingface/hub

command: |
  python3 -m sglang.launch_server \
    --trust-remote-code \
    --model-path {model} \
    --tp-size {tensor_parallel} \
    --served-model-name {served_model_name} \
    --mem-fraction-static {gpu_memory_utilization} \
    --attention-backend {attention_backend} \
    --chunked-prefill-size 8192 \
    --disable-prefill-cuda-graph \
    --cuda-graph-max-bs 4 \
    --speculative-algorithm {speculative_algorithm} \
    --speculative-draft-model-path {speculative_draft_model_path} \
    --speculative-dspark-block-size {speculative_dspark_block_size} \
    --speculative-draft-model-quantization {speculative_draft_model_quantization} \
    --mamba-scheduler-strategy extra_buffer \
    --enable-torch-compile --torch-compile-max-bs 4 \
    --num-continuous-decode-steps 2 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --host {host} \
    --port {port}

```

```auto
sparkrun run ~/qwen38-27b-sglang-dspark.yaml

```

* * *

## **NVFP4 (Inferact/Qwen3.8-27B-NVFP4) + MTP, TP=2, vLLM official**

Moving the same InferAct + MTP setup from TP=1 to TP=2 raises C1 TPS from 18.5 to 23.4, about 26%. The scaling is useful but far from 2×. Tensor parallelism reduces the amount of sharded model work handled by each GPU, while also introducing communication on the decode path, so linear scaling should not be expected. TTFT improves from 341 ms to 254 ms at C1. The configuration still passes through C8 and falls below the TPS threshold at C16.

| **C** | **TTFT (ms)** | **TPS** | **Pass?** |
| --- | --- | --- | --- |
| 1 | 254 | 23.4 | ✓ |
| 2 | 423 | 22.4 | ✓ |
| 4 | 454 | 21.3 | ✓ |
| 8 | 570 | 17.0 | ✓ |
| 16 | 699 | 12.8 | ✗ (TPS) |
| 32 | 859 | 10.7 | ✗ (TPS) |

```auto
recipe_version: "2"
model: Inferact/Qwen3.8-27B-NVFP4
runtime: vllm-distributed
container: vllm/vllm-openai:qwen38

executor_config:
  entrypoint: ""

metadata:
  kv_dtype: fp8
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (InferAct) + MTP, TP=2, official vLLM image

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 2
  gpu_memory_utilization: 0.75

command: |
  vllm serve Inferact/Qwen3.8-27B-NVFP4 \
    --host {host} \
    --port {port} \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

```

```auto
sparkrun run ~/qwen38-27b-inferact-nvfp4-tp2.yaml --hosts <head>,<worker1>

```

* * *

## **NVFP4 (Inferact/Qwen3.8-27B-NVFP4) + MTP, TP=2, eugr vLLM**

The eugr recipe reaches 33.0 tok/s at C1 versus 23.4 with the official vLLM configuration and remains ahead throughout the concurrency sweep. It is also the first vLLM setup here to satisfy both criteria at C16.

The native SM 12.1a support in the eugr stack is a plausible contributor to the improvement. In addition, the eugr recipe changes several serving parameters, including memory utilization, maximum model length, attention backend, load format, chunked prefill, and asynchronous scheduling.

| **C** | **TTFT (ms)** | **TPS** | **Pass?** |
| --- | --- | --- | --- |
| 1 | 311 | 33.0 | ✓ |
| 2 | 467 | 31.8 | ✓ |
| 4 | 540 | 29.5 | ✓ |
| 8 | 599 | 24.0 | ✓ |
| 16 | 831 | 17.7 | ✓ |
| 32 | 1345 | 13.4 | ✗ (both) |

```auto
recipe_version: '2'
name: Qwen3.8-27B-NVFP4-MTP-TP2
model: Inferact/Qwen3.8-27B-NVFP4
container: ghcr.io/spark-arena/dgx-vllm-eugr-nightly:latest
runtime: vllm-distributed

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 2
  gpu_memory_utilization: 0.55
  max_model_len: 131072
  max_num_batched_tokens: 8192
  served_model_name: Qwen3.8-27B
  attention_backend: flashinfer
  tool_call_parser: qwen3_coder
  pipeline_parallel: 1
  load_format: fastsafetensors
  kv_cache_dtype: fp8
  reasoning_parser: qwen3
  speculative_config: '{"method":"mtp","num_speculative_tokens":3}'
  override_generation_config: '{"temperature":1.0,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'

env:
  VLLM_MARLIN_USE_ATOMIC_ADD: '1'
  TORCH_MATMUL_PRECISION: high
  FLASHINFER_DISABLE_VERSION_CHECK: '1'
  NVIDIA_FORWARD_COMPAT: '1'
  VLLM_HTTP_TIMEOUT_KEEP_ALIVE: '600'

command: |
  vllm serve {model} \
    --served-model-name {served_model_name} \
    --host {host} \
    --port {port} \
    -tp {tensor_parallel} \
    -pp {pipeline_parallel} \
    --trust-remote-code \
    --language-model-only \
    --kv-cache-dtype {kv_cache_dtype} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len {max_model_len} \
    --max-num-batched-tokens {max_num_batched_tokens} \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --async-scheduling \
    --load-format {load_format} \
    --attention-backend {attention_backend} \
    --enable-auto-tool-choice \
    --tool-call-parser {tool_call_parser} \
    --reasoning-parser {reasoning_parser} \
    --speculative-config '{speculative_config}' \
    --override-generation-config '{override_generation_config}'

metadata:
  description: Qwen3.8-27B NVFP4 (Inferact) + MTP — TP=2 dual Spark, eugr image
  model_params: 27B
  model_dtype: nvfp4
  kv_dtype: fp8
  quantization: nvfp4
  quant_bits: 4

```

```auto
sparkrun run ~/qwen38-27b-inferact-nvfp4-tp2-eugr.yaml --hosts <head>,<worker1>

```

* * *

## **NVFP4 (RadixArk/Qwen3.8-27B-NVFP4) + DSpark, TP=2, SGLang**

This is a cleaner scaling comparison with the previous SGLang run: moving from TP=1 to TP=2 raises C1 TPS from 36.6 to 51.8, about 42%. The larger gain is not limited to C1. The configuration remains above both thresholds through C16, whereas the single-Spark run fails at C8 because of TTFT.

| **C** | **TTFT (ms)** | **TPS** | **Pass?** |
| --- | --- | --- | --- |
| 1 | 219 | 51.8 | ✓ |
| 2 | 286 | 46.4 | ✓ |
| 4 | 324 | 36.0 | ✓ |
| 8 | 375 | 24.9 | ✓ |
| 16 | 460 | 16.7 | ✓ |
| 32 | 5682 | 9.7 | ✗ (both) |

```auto
recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4
runtime: sglang
container: lmsysorg/sglang@sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1

metadata:
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (RadixArk) + DSpark — SGLang TP=2 on 2x DGX Spark

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 2
  gpu_memory_utilization: 0.50
  served_model_name: qwen3.8-27b
  attention_backend: flashinfer
  speculative_algorithm: DSPARK
  speculative_draft_model_path: RadixArk/Qwen3.8-27B-DSpark
  speculative_dspark_block_size: 7
  speculative_draft_model_quantization: unquant

env:
  TORCHINDUCTOR_CACHE_DIR: /cache/huggingface/sglang-inductor
  HF_HOME: /cache/huggingface
  HF_HUB_CACHE: /cache/huggingface/hub

command: |
  python3 -m sglang.launch_server \
    --trust-remote-code \
    --model-path {model} \
    --tp-size {tensor_parallel} \
    --served-model-name {served_model_name} \
    --mem-fraction-static {gpu_memory_utilization} \
    --attention-backend {attention_backend} \
    --chunked-prefill-size 8192 \
    --disable-prefill-cuda-graph \
    --cuda-graph-max-bs 4 \
    --speculative-algorithm {speculative_algorithm} \
    --speculative-draft-model-path {speculative_draft_model_path} \
    --speculative-dspark-block-size {speculative_dspark_block_size} \
    --speculative-draft-model-quantization {speculative_draft_model_quantization} \
    --mamba-scheduler-strategy extra_buffer \
    --enable-torch-compile --torch-compile-max-bs 4 \
    --num-continuous-decode-steps 2 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --host {host} \
    --port {port}

```

```auto
sparkrun run ~/qwen38-27b-sglang-dspark-tp2.yaml --hosts <head>,<worker1>

```

* * *

## **NVFP4 (RadixArk/Qwen3.8-27B-NVFP4) + DSpark, TP=4, SGLang**

Moving the same SGLang + DSpark setup from TP=2 to TP=4 raises C1 throughput from 51.8 to 77.3 tok/s, another 49%. This is the fastest low-concurrency configuration in the study: 77.3 tok/s with 227 ms TTFT at C1. The additional nodes also keep TTFT below 1 second through C32.

| **C** | **TTFT (ms)** | **TPS** | **Pass?** |
| --- | --- | --- | --- |
| 1 | 227 | 77.3 | ✓ |
| 2 | 247 | 61.9 | ✓ |
| 4 | 264 | 48.1 | ✓ |
| 8 | 379 | 27.0 | ✓ |
| 16 | 411 | 18.6 | ✓ |
| 32 | 579 | 11.0 | ✗ (TPS) |

```auto
recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4
runtime: sglang
container: lmsysorg/sglang@sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1
​
metadata:
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (RadixArk) + DSpark — SGLang TP=4 on 4x DGX Spark
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 4
  gpu_memory_utilization: 0.50
  served_model_name: qwen3.8-27b
  attention_backend: flashinfer
  speculative_algorithm: DSPARK
  speculative_draft_model_path: RadixArk/Qwen3.8-27B-DSpark
  speculative_dspark_block_size: 7
  speculative_draft_model_quantization: unquant
​
env:
  TORCHINDUCTOR_CACHE_DIR: /cache/huggingface/sglang-inductor
  HF_HOME: /cache/huggingface
  HF_HUB_CACHE: /cache/huggingface/hub
​
command: |
  python3 -m sglang.launch_server \
    --trust-remote-code \
    --model-path {model} \
    --tp-size {tensor_parallel} \
    --served-model-name {served_model_name} \
    --mem-fraction-static {gpu_memory_utilization} \
    --attention-backend {attention_backend} \
    --chunked-prefill-size 8192 \
    --disable-prefill-cuda-graph \
    --cuda-graph-max-bs 4 \
    --speculative-algorithm {speculative_algorithm} \
    --speculative-draft-model-path {speculative_draft_model_path} \
    --speculative-dspark-block-size {speculative_dspark_block_size} \
    --speculative-draft-model-quantization {speculative_draft_model_quantization} \
    --mamba-scheduler-strategy extra_buffer \
    --enable-torch-compile --torch-compile-max-bs 4 \
    --num-continuous-decode-steps 2 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --host {host} \
    --port {port}

```

```auto
sparkrun run ~/qwen38-27b-sglang-dspark-tp4.yaml --hosts <head>,<worker1>,<worker2>,<worker3>

```

* * *

## **What this suggests**

The progression seems to make a few things fairly clear for this particular.

Reducing precision helps, but precision alone is not enough. BF16 → FP8 raises C1 throughput from 4.5 to 7.9 tok/s, while NVFP4 on the official vLLM stack reaches only 9.7 despite the much smaller model footprint. In contrast, speculative decoding changes the practical result much more: MTP takes both FP8 and NVFP4 from configurations that fail at every concurrency level to configurations that meet our threshold.

The runtime and serving stack matter at least as much as adding hardware. A single-Spark SGLang + DSpark configuration reaches 36.6 tok/s at C1, ahead of the two-Spark official-vLLM result and even slightly ahead of the two-Spark eugr result at C1. Going from two to four Sparks increases C1 throughput again, from 51.8 to 77.3 tok/s, but it does not increase higher concurrency performance that much: TP=2 and TP=4 both top out at C16 in the end. Under this benchmark, the four-Spark setup primarily a choice for maximum single-user or low-concurrency speed rather than additional multi-user capacity.

So, with the criteria used here, the two-Spark SGLang + DSpark setup looks like the most balanced point in the progression: it retains much of the low-concurrency speedup, serves through C16, and avoids using two additional Sparks without increasing the maximum passing concurrency. TP=4 is reasonable when 50–77 tok/s low-concurrency generation is worth the extra nodes. On a single Spark, SGLang + DSpark is the fastest option we tested for light concurrency, while NVFP4 + MTP on vLLM is slower but behaves more predictably through C8.

Those conclusions are specific to this benchmark and our 15 tok/s / 1 s thresholds; longer contexts or different serving objectives could shift the preferred configuration.

* * *

Thanks to RadixArk for the Qwen3.8-27B NVFP4 + DSpark models, InferAct for the NVFP4 variant, @basbunarhasan for the SGLang + DSpark DGX Spark recipe, @eugr_nv for the native SM 12.1a nightly images, and the sparkrun crew for the orchestration tool.

---

<div class="post-metadata">

**Author:** ![emretoktas\_openzeka](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/emretoktas_openzeka/32/532000_2.png) [@emretoktas\_openzeka](https://forums.developer.nvidia.com/u/emretoktas_openzeka)\
**Post date:** [August 24, 2026, 9:39am UTC](https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102/2 "2026-08-24T09:39:32Z")

</div>

For the same 4× Spark hardware, we also tested a different scaling strategy: instead of one TP=4 group, we split the four Sparks into two independent TP=2 groups behind a round-robin router — effectively TP=2 + DP=2.

The trade-off is lower single-user TPS, but substantially better scaling under concurrent load. This is the only configuration in the study that passes our criteria at C32.

```auto
Benchmark → SMG Router (<head>:30000, round_robin)
                 ├── TP=2 Group 1: <head>:8000 ←→ <worker1>:8000
                 └── TP=2 Group 2: <worker2>:8000 ←→ <worker3>:8000

```

| **C** | **TTFT (ms)** | **TPS** | **Pass?** |
| --- | --- | --- | --- |
| 1 | 207 | 54.6 | ✓ |
| 2 | 222 | 52.9 | ✓ |
| 4 | 263 | 46.8 | ✓ |
| 8 | 321 | 36.0 | ✓ |
| 16 | 354 | 24.2 | ✓ |
| 32 | 465 | 16.8 | ✓ |
| 64 | 5972 | 9.7 | ✗ (both) |

Each replica is the same TP=2 SGLang + DSpark configuration from the earlier test. The router simply distributes requests between the two groups, so at total concurrency 2C, each TP=2 group sees roughly concurrency C.

The results line up almost exactly:

| **TP=2 at C** | **TPS** | **TP=2+DP=2 at 2C** | **TPS** |
| --- | --- | --- | --- |
| C1 | 51.8 | C2 | 52.9 |
| C2 | 46.4 | C4 | 46.8 |
| C4 | 36.0 | C8 | 36.0 |
| C8 | 24.9 | C16 | 24.2 |
| C16 | 16.7 | C32 | 16.8 |

That is a particularly clean result. TP=2 + DP=2 at total concurrency 2C behaves almost exactly like a single TP=2 group at concurrency C.

This also explains the C32 result: with the requests split across two replicas, each group is effectively handling about C16. A single TP=2 group at C16 measured 16.7 tok/s and 460 ms TTFT; TP=2 + DP=2 at C32 measures 16.8 tok/s and 465 ms. In this benchmark, the round-robin layer therefore adds negligible observable overhead, while the second replica almost directly translates into additional concurrent capacity.

Compared with TP=4 on the same four Sparks:

| **C** | **TP=4 TPS** | **TP=2+DP=2 TPS** | **Winner** | **Delta** |
| --- | --- | --- | --- | --- |
| 1 | 77.3 | 54.6 | TP=4 | +42% |
| 2 | 61.9 | 52.9 | TP=4 | +17% |
| 4 | 48.1 | 46.8 | TP=4 | +3% |
| 8 | 27.0 | 36.0 | DP=2 | +33% |
| 16 | 18.6 | 24.2 | DP=2 | +30% |
| 32 | 11.0 | 16.8 | DP=2 | +53% |

So the choice between TP and DP changes with concurrency. TP=4 gives 77.3 tok/s at C1, while TP=2 + DP=2 gives 54.6 — about 29% lower. By C4, however, the difference has almost disappeared, and somewhere between C4 and C8 the advantage crosses over. From C8 onward, the replicated setup is clearly faster.

C32 is where the difference becomes most practical: TP=4 drops to 11.0 tok/s and fails our TPS threshold, while TP=2 + DP=2 still delivers 16.8 tok/s with 465 ms TTFT and passes both criteria.

For this workload, that makes TP=4 the better use of four Sparks when the priority is maximum single-user or low-concurrency speed. If the same hardware is intended to serve more simultaneous users, two TP=2 replicas are the more effective configuration: they give up some C1 performance, but almost double the useful concurrency range from C16 to C32.

The router here is SGLang’s Model Gateway (`sgl-model-gateway`), running as a separate Docker container rather than through sparkrun. We use its `round_robin` policy; it also provides worker health checking, circuit breakers, and Prometheus metrics.

Both groups use the same SGLang recipe — only the description differs:

```auto
recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4
runtime: sglang
container: lmsysorg/sglang@sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1
​
metadata:
  model_dtype: nvfp4
  description: Qwen3.8-27B NVFP4 (RadixArk) + DSpark — SGLang TP=2 Group 1
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 2
  gpu_memory_utilization: 0.50
  served_model_name: qwen3.8-27b
  attention_backend: flashinfer
  speculative_algorithm: DSPARK
  speculative_draft_model_path: RadixArk/Qwen3.8-27B-DSpark
  speculative_dspark_block_size: 7
  speculative_draft_model_quantization: unquant
​
env:
  TORCHINDUCTOR_CACHE_DIR: /cache/huggingface/sglang-inductor
  HF_HOME: /cache/huggingface
  HF_HUB_CACHE: /cache/huggingface/hub
​
command: |
  python3 -m sglang.launch_server \
    --trust-remote-code \
    --model-path {model} \
    --tp-size {tensor_parallel} \
    --served-model-name {served_model_name} \
    --mem-fraction-static {gpu_memory_utilization} \
    --attention-backend {attention_backend} \
    --chunked-prefill-size 8192 \
    --disable-prefill-cuda-graph \
    --cuda-graph-max-bs 4 \
    --speculative-algorithm {speculative_algorithm} \
    --speculative-draft-model-path {speculative_draft_model_path} \
    --speculative-dspark-block-size {speculative_dspark_block_size} \
    --speculative-draft-model-quantization {speculative_draft_model_quantization} \
    --mamba-scheduler-strategy extra_buffer \
    --enable-torch-compile --torch-compile-max-bs 4 \
    --num-continuous-decode-steps 2 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --host {host} \
    --port {port}

```

Launch sequence (all from head node):

```auto
sparkrun run ~/qwen38-27b-sglang-dspark-tp2-group1.yaml --hosts <head>,<worker1>

```

```auto
sparkrun run ~/qwen38-27b-sglang-dspark-tp2-group2.yaml --hosts <worker2>,<worker3>

```

```auto
docker run -d --name sglang-router \
  --network host \
  lmsysorg/sgl-model-gateway:latest \
  --worker-urls http://<head>:8000 http://<worker2>:8000 \
  --policy round_robin \
  --host 0.0.0.0 --port 30000 \
  --prometheus-port 10001

```

---

<div class="post-metadata">

**Author:** ![pontostroy](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@pontostroy](https://forums.developer.nvidia.com/u/pontostroy)\
**Post date:** [August 24, 2026, 12:40pm UTC](https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102/3 "2026-08-24T12:40:43Z")

</div>

> [@emretoktas\_openzeka](#):
>
> ## **NVFP4 (RadixArk/Qwen3.8-27B-NVFP4) + DSpark, TP=1, SGLang**
> 
> | **C** | **TTFT (ms)** | **TPS** | **Pass?** |
> | --- | --- | --- | --- |
> | 1 | 246 | 36.6 | ✓ |
> | 2 | 349 | 31.3 | ✓ |
> | 4 | 394 | 25.6 | ✓ |
> | 8 | 1368 | 17.8 | ✗ (TTFT) |
> | 16 | 8554 | 9.2 | ✗ (both) |
> | 32 | 23045 | 4.7 | ✗ (both) |

Can you check with DFlash2?  
My results with tuned for c1 SGLang

## **NVFP4 (RadixArk/Qwen3.8-27B-NVFP4) + DFlash2, TP=1, SGLang**

```auto
Conc 1: TTFT=217ms \[OK\] TPS=56.6 \[OK\]
Conc 2: TTFT=330ms \[OK\] TPS=54.6 \[OK\]
Conc 4: TTFT=3938ms \[X\] TPS=24.2 \[OK\]
Conc 8: TTFT=10598ms \[X\] TPS=11.2 \[X\]
Conc 16: TTFT=14102ms \[X\] TPS=13.3 \[X\]

```

---

<div class="post-metadata">

**Author:** ![emretoktas\_openzeka](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/emretoktas_openzeka/32/532000_2.png) [@emretoktas\_openzeka](https://forums.developer.nvidia.com/u/emretoktas_openzeka)\
**Post date:** [August 25, 2026, 8:55am UTC](https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102/4 "2026-08-25T08:55:35Z")

</div>

Amazing results!

One point to consider: with our benchmarking tool, the result table has two modes — p90 and mean. We usually use mean, but it might be the case that your table had p90 selected, which would give a slightly larger TPS — the MiaAI-Lab repo’s own results show 50.9 TPS.

As for the DFlash2, can you share which model variant you used (`RadixArk/Qwen3.8-27B-NVFP4` or `RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead`) and your exact tunings and flags? We want to make sure we’re comparing the same setup.

---

<div class="post-metadata">

**Author:** ![pontostroy](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@pontostroy](https://forums.developer.nvidia.com/u/pontostroy)\
**Post date:** [August 25, 2026, 11:02am UTC](https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102/5 "2026-08-25T11:02:42Z")

</div>

Retest with mean

```auto
Conc 1: TTFT=217ms [OK] TPS=48.0 [OK]
Conc 2: TTFT=306ms [OK] TPS=42.3 [OK]
Conc 4: TTFT=3251ms [X] TPS=22.1 [OK]

```

I use [GitHub - hasso5703/dgx-spark-qwen38: Fastest measured Qwen3.8-27B config for DGX Spark (GB10): SGLang + NVFP4 + DFlash2, deterministic boots, 50 tok/s greedy median, 148 tok/s at 8 streams, 258 at 32. One command, everything pinned, lossless. · GitHub](https://github.com/hasso5703/dgx-spark-qwen38) as base but run it with llama-swap  
My llama-swap model command

```auto
    cmd: |
      /usr/bin/docker run --rm
      --name qwen38-sglang
      --gpus all
      --memory 100g
      --memory-swap 100g
      --shm-size 16g
      --network host
      --ipc=host
      -e TORCHINDUCTOR_CACHE_DIR=/cache/inductor
      -v /home/pont/.config/qwen38/sglang-cache:/cache
      -v /home/pont/.cache/huggingface:/root/.cache/huggingface
      -v /home/pont/.config/qwen38:/out
      qwen38-dflash2:v1.2.2
      python3 -m sglang.launch_server
      --trust-remote-code
      --model-path RadixArk/Qwen3.8-27B-NVFP4
      --tp-size 1
      --served-model-name qwen3.8-27b
      --mem-fraction-static 0.45
      --attention-backend flashinfer
      --chunked-prefill-size 2048
      --cuda-graph-max-bs-decode 2
      --sleep-on-idle
      --speculative-algorithm DFLASH
      --speculative-draft-model-path z-lab/Qwen3.8-27B-DFlash2
      --speculative-draft-model-revision 50307d4c4cde6860d4eee73e2547cd786fe8e8a4
      --speculative-num-draft-tokens 8
      --speculative-draft-model-quantization unquant
      --mamba-radix-cache-strategy extra_buffer
      --mamba-ssm-dtype bfloat16
      --max-mamba-cache-size 32
      --max-running-requests 2
      --torch-compile-max-bs 2
      --num-continuous-decode-steps 1
      --reasoning-parser qwen3
      --tool-call-parser qwen3_coder
      --chat-template /out/chat-template-sglang.jinja
      --host 0.0.0.0
      --port ${PORT}
      --enable-metrics
      --enable-mfu-metrics
      --enable-cache-report
      --enable-request-time-stats-logging
      --log-requests
      --log-requests-level 1
      --log-requests-format json
      --show-time-cost

```

---

<div class="post-metadata">

**Author:** ![emretoktas\_openzeka](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/emretoktas_openzeka/32/532000_2.png) [@emretoktas\_openzeka](https://forums.developer.nvidia.com/u/emretoktas_openzeka)\
**Post date:** [August 25, 2026, 1:59pm UTC](https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102/6 "2026-08-25T13:59:20Z")

</div>

Thanks to @pontostroy for the suggestions. We experimented with the [DFlash2 draft model](https://huggingface.co/z-lab/Qwen3.8-27B-DFlash2) and MiaAI-Lab’s [DFlash2 setup workflow](https://github.com/MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark) a bit and matched your C1 result, then spent more time on concurrency behavior and reproducibility. Here is the final setup we came up with:

## **Results — 47.9 TPS at C1**

Same benchmark methodology as the study above (CordatusAI, 128 in / 128 out, 10 rounds, pass: TTFT \< 1000ms AND TPS \>= 15 tok/s).

Model: `RadixArk/Qwen3.8-27B-NVFP4` (packed FP4, 21.9 GB)  
Draft: `z-lab/Qwen3.8-27B-DFlash2` (2B, BF16)  
Image: `lmsysorg/sglang:qwen38-27b-dflash2` (–full, commit `c14312a66` + `dflash2_nvfp4_head.patch`)  
`--mem-fraction-static 0.50`, `--max-running-requests 16`

| **C** | **TTFT (ms)** | **TPS** | **Pass?** |
| --- | --- | --- | --- |
| 1 | 232 | 47.9 | ✓ |
| 2 | 320 | 39.8 | ✓ |
| 4 | 343 | 35.6 | ✓ |
| 8 | 389 | 27.6 | ✓ |
| 16 | 2357 | 18.6 | TPS ✓, TTFT ✗ |

In addition to the 47.9 TPS at C1, TTFT holds at 343ms and TPS is 27.6 at C8 — the setup can be used with 8 concurrency in mind comfortably.

## **Running with sparkrun**

The full configuration is captured in a single sparkrun recipe:

**1. Build the DFlash2 image (one-time, on each Spark):**

```auto
git clone https://github.com/MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark.git
cd Qwen3.8-27B-SGLang-DGX-Spark
bash patch/build-dflash2-image.sh --full

```

**2. Save the recipe and launch:**

```auto
recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4
runtime: sglang
container: lmsysorg/sglang:qwen38-27b-dflash2
​
max_nodes: 1
​
metadata:
  model_dtype: nvfp4
  kv_dtype: fp8
  description: Qwen3.8-27B NVFP4 (packed FP4) + DFlash2 — single Spark
​
defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.50
  served_model_name: qwen3.8-27b
  attention_backend: flashinfer
  speculative_algorithm: DFLASH
  speculative_draft_model_path: z-lab/Qwen3.8-27B-DFlash2
  speculative_num_draft_tokens: 8
​
env:
  HF_HUB_OFFLINE: "0"
  TORCHINDUCTOR_CACHE_DIR: /cache/inductor
  HF_HOME: /cache/huggingface
  HF_HUB_CACHE: /cache/huggingface/hub
​
command: |
  python3 -m sglang.launch_server \
    --trust-remote-code \
    --model-path {model} \
    --tp-size {tensor_parallel} \
    --served-model-name {served_model_name} \
    --mem-fraction-static {gpu_memory_utilization} \
    --attention-backend {attention_backend} \
    --chunked-prefill-size 8192 \
    --disable-prefill-cuda-graph \
    --kv-cache-dtype fp8_e4m3 \
    --mamba-ssm-dtype bfloat16 \
    --mamba-full-memory-ratio 4.21 \
    --mamba-radix-cache-strategy extra_buffer \
    --max-mamba-cache-size 64 \
    --max-running-requests 16 \
    --context-length 262144 \
    --speculative-algorithm {speculative_algorithm} \
    --speculative-draft-model-path {speculative_draft_model_path} \
    --speculative-draft-model-revision 50307d4c4cde6860d4eee73e2547cd786fe8e8a4 \
    --speculative-num-draft-tokens {speculative_num_draft_tokens} \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --sampling-defaults model \
    --enable-metrics \
    --enable-cache-report \
    --stream-interval 1 \
    --host {host} \
    --port {port}

```

```auto
sparkrun run ~/qwen38-27b-dflash2.yaml --rootful --no-follow --no-rm

```

**Optional — CPU pinning:**

```auto
docker update --cpuset-cpus 5-9,15-19 $(docker ps --format "{{.Names}}" --filter "ancestor=lmsysorg/sglang:qwen38-27b-dflash2")

```

This pins the container to GB10’s 10 Cortex-X5 performance cores, keeping the scheduler and tokenizer off the 2.8 GHz efficiency cores. We measured less than 2 TPS difference at C1 and no meaningful change at higher concurrencies, so it can be considered optional.

---

<div class="post-metadata">

**Author:** ![xkm121](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@xkm121](https://forums.developer.nvidia.com/u/xkm121)\
**Post date:** [August 25, 2026, 2:16pm UTC](https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102/7 "2026-08-25T14:16:54Z")

</div>

MiaAI lab did not update her receipe. There is an official SGlang docker image with DFlash2 merged now  
lmsysorg/sglang:dev-qwen38-27b-dflash2

Also you should try RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead, with the dlash2 above, better quality

---

<div class="post-metadata">

**Author:** ![pontostroy](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@pontostroy](https://forums.developer.nvidia.com/u/pontostroy)\
**Post date:** [August 25, 2026, 4:59pm UTC](https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102/8 "2026-08-25T16:59:47Z")

</div>

Tested prefill size with very big context  
2048 always better

tool-eval-bench --backend sglang --context-size 524288 --context-pressure 0.6

| Chunked Prefill Size | TTFT | Long-Context Prefill | Peak/Stable TFLOPS | Result |
| --- | --- | --- | --- | --- |
| 2048 | 415 s | ~390–480 tok/s | ~180–182 TFLOPS | 🥇 Best |
| 4096 | 424.1 s | ~384–481 tok/s | ~176–179 TFLOPS | 🥈 Good |
| 8192 | 441.1 s | ~383–464 tok/s | ~169–172 TFLOPS | 🥉 Slowest |

---

<div class="post-metadata">

**Author:** ![emretoktas\_openzeka](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/emretoktas_openzeka/32/532000_2.png) [@emretoktas\_openzeka](https://forums.developer.nvidia.com/u/emretoktas_openzeka)\
**Post date:** [August 25, 2026, 6:17pm UTC](https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102/9 "2026-08-25T18:17:39Z")

</div>

Thanks for the heads-up!

We also tested `RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead` with DFlash2, and the inference performance was very similar in our benchmarks.

When you say “better quality,” I believe that you mean output quality rather than inference speed. If you have comparative evaluation results or experiences between the two models, we’d be interested to see them.

We’ll also take a look at the new `lmsysorg/sglang:dev-qwen38-27b-dflash2` image.

---

<div class="post-metadata">

**Author:** ![jrsphd](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/jrsphd/32/459183_2.png) [@jrsphd](https://forums.developer.nvidia.com/u/jrsphd)\
**Post date:** [August 25, 2026, 9:29pm UTC](https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102/10 "2026-08-25T21:29:44Z")

</div>

Something I messed around with today which hasn’t been mentioned yet: B12x + DFlash2 on Unsloth’s NVFP4. I don’t run benches, and just got this up, but for C1 with TP2 and batch of 8192 am seeing prose from 30s → mid 40s and code 50s → low 80s. Haven’t tried against FP8 quant yet.

---

<div class="post-metadata">

**Author:** ![emretoktas\_openzeka](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/emretoktas_openzeka/32/532000_2.png) [@emretoktas\_openzeka](https://forums.developer.nvidia.com/u/emretoktas_openzeka)\
**Post date:** [August 26, 2026, 11:22am UTC](https://forums.developer.nvidia.com/t/comprehensive-qwen3-8-27b-study-on-dgx-sparks-quantization-speculative-decoding-and-tp-dp-scaling/381102/11 "2026-08-26T11:22:37Z")

</div>

We used the `RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead` with our existing setup for some time. There seems to be a quality increase as expected, while keeping the high performance.

As for the `lmsysorg/sglang:dev-qwen38-27b-dflash2`, our benchmarks showed ~3-5 TPS decrease at C1. We used the [official cookbook](https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-27B#hw=dgx-spark&variant=default&quant=nvfp4-bf16-head&nodes=single&spec=dflash&tier=low-latency&ssmDtype=float32) with these selections:

- Hardware: DGX Spark
- Model: RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead (NVFP4 quantization, BF16 lm\_head)
- Draft model: incoai/Qwen3.8-27B-DFlash2
- Speculative decoding: DFlash2 (8 draft tokens)
- Nodes: single

The cookbook generated the serving cell with these parameters:

- mem-fraction-static: 0.80
- chunked-prefill-size: 2048
- kv-cache-dtype: fp8\_e4m3
- attention-backend: flashinfer
- reasoning-parser: qwen3
- tool-call-parser: qwen3\_coder

We launched it as-is via sparkrun. Here is the sparkrun recipe and run command:

* * *

Recipe (`~/qwen38-27b-dflash2-official.yaml`):

```yaml
recipe_version: "2"
model: RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead
runtime: sglang
container: lmsysorg/sglang:dev-qwen38-27b-dflash2

max_nodes: 1

metadata:
  model_dtype: nvfp4
  kv_dtype: fp8
  description: Qwen3.8-27B NVFP4 (BF16-LMHead) + DFlash2 — official SGLang cookbook image

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.80
  served_model_name: qwen3.8-27b

env:
  HF_HUB_OFFLINE: "0"

command: |
  python3 -m sglang.launch_server \
    --trust-remote-code \
    --model-path {model} \
    --tp-size {tensor_parallel} \
    --served-model-name {served_model_name} \
    --mem-fraction-static {gpu_memory_utilization} \
    --attention-backend flashinfer \
    --chunked-prefill-size 2048 \
    --kv-cache-dtype fp8_e4m3 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --speculative-algorithm DFLASH \
    --speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2 \
    --speculative-num-draft-tokens 8 \
    --host {host} \
    --port {port}

```

Run command:

```bash
sparkrun run ~/qwen38-27b-dflash2-official.yaml --rootful --no-follow --no-rm

```
