4-node DGX Spark Cluster with DeepSeek V4 flash-0731-Dspark benchmark (prefill=2,500 t/s, decode=90 t/s)

Most of the work was done by GPT. All I did was plug in the cables and press the power button. (lol)

Here are some quick notes:

1. C=1, 50K context length, cold: TTFT ~ 20s, prefill ~ 2,500 tok/s, decode ~ 90 tok/s

2. C=1, 150K context length, KV cache hit: TTFT ~ 0.776s, effective prefill ~ 193,280 tok/s, decode ~ 90 tok/s

3. C=6, 150K context length, KV cache hit: TTFT ~ 4.334s, decode ~ 40.4 tok/s per request (avg)

## 2. Hardware Configuration

Private 4-node Spark cluster, wired through a QRS812 switch fabric.

Component Configuration

| Compute | 4x ASUS Spark / DGX Spark-class nodes |

| Switch | QRS812 |

| Cabling | 2x MK-TW-QDD-20x QSFP56 400G breakout DAC |

| Fabric | Spark port2 / QRS812 fabric |

| Cluster mode | 4-node tensor parallel, TP=4 |

Total build cost was roughly **USD $20.5K-$21.1K**, with some parts bought before the recent price hikes.

## 3. QRS812 Switch / RDMA Latency Benchmark

Ran a full-mesh RDMA verbs latency matrix across the QRS812-facing Spark port2 fabric.

Test profile:

Item Value

| RDMA device | `roceP2p1s0f1` |

| RoCE GID index | `11` |

| Message size | `2B` |

| Iterations | `10000` |

| Tests | `ib_write_lat`, `ib_read_lat`, `ib_send_lat` |

Summary:

| Test | Avg Latency Range | p99 Range | Notes |

|—|—:|—:|—|

| RDMA write | `2.93-3.49 us` | `4.17-4.94 us` | Tight across the full mesh |

| RDMA read | `5.64-6.34 us` | `8.27-10.45 us` | Higher than write/send, but stable |

| RDMA send | `2.55-3.44 us` | `2.95-4.65 us` | Most consistent p99 |

Full matrix:

QRS812 Spark Fabric RDMA Latency Matrix

device=roceP2p1s0f1 gid=11 size=2B iters=10000 tests=write read send

write latency

FROM\TO spark1 spark2 spark3 spark4

spark1 self 3.49us p99=4.94 3.20us p99=4.44 3.22us p99=4.52

spark2 3.49us p99=4.82 self 3.22us p99=4.54 3.20us p99=4.47

spark3 3.19us p99=4.35 3.23us p99=4.62 self 2.93us p99=4.17

spark4 3.25us p99=4.62 3.22us p99=4.70 2.97us p99=4.36 self

read latency

FROM\TO spark1 spark2 spark3 spark4

spark1 self 6.28us p99=9.55 6.34us p99=9.50 6.33us p99=10.21

spark2 6.32us p99=10.00 self 6.23us p99=9.47 6.28us p99=10.45

spark3 5.68us p99=8.27 5.66us p99=8.78 self 5.71us p99=9.73

spark4 5.66us p99=8.27 5.64us p99=8.83 5.70us p99=8.66 self

send latency

FROM\TO spark1 spark2 spark3 spark4

spark1 self 3.44us p99=4.65 3.17us p99=4.40 3.17us p99=4.50

spark2 3.10us p99=3.50 self 2.82us p99=3.22 2.83us p99=3.23

spark3 2.85us p99=3.25 2.85us p99=3.26 self 2.55us p99=2.95

spark4 2.84us p99=3.26 2.84us p99=3.24 2.55us p99=2.95 self

```

Additional fabric checks:

| Check | Result |

|— |— |

| Full-mesh ICMP RTT over QRS812 | about `0.49-1.26 ms` observed |

| TCP iperf2, Spark1 → Spark2, 16 streams | about `105 Gbit/s` |

| RDMA `ib_write_bw`, Spark1 → Spark2 over QRS812 | `107.66 Gbit/s` |

| Direct QSFP112 baseline, Spark1 → Spark3 | `107.66 Gbit/s` |

| NIC PHY errors after transfer / benchmark | CRC, symbol, discard, link-down all `0` |

The QRS812 fabric looks clean for Spark-to-Spark RoCE traffic. Write/send latencies sit in the low single-digit microseconds, read around 6 us.

## 4. Model And Quantization Profile

The cluster serves **DeepSeek-V4-Flash-0731** with **DSpark speculative decoding**.

| Item | Value |

|— |— |

| Model | `deepseek-ai/DeepSeek-V4-Flash-0731` |

| Runtime | vLLM + DSpark |

| Parallelism | TP=4 across 4 Spark nodes |

| Max model length | `512000` |

| KV cache dtype | `nvfp4_ds_mla` |

| Prefix cache | enabled |

| Scheduling | async scheduling + chunked prefill |

| DSpark speculative tokens | `MTP_NUM_TOKENS=3` |

| vLLM fingerprint | `vllm-0.21.1rc1.dev339+g1967a5627bc3-tp4-d2211cb5` |

Note this isn’t a plain FP8 deployment. Other vLLM examples serve the official model with FP8 KV; this GR7a path uses the experimental **NVFP4 DS-MLA KV** path instead.

## 5. vLLM Image / GR7a Runtime Notes

Current profile label is **GR7a**: a 4-node QRS812 profile based on existing community DSpark / NVFP4 work, adapted to this hardware layout.

Community sources we built from or cross-checked:

Our main modifications:

| Area | Modification |

|—|—|

| Topology | Expanded the recipe to 4 nodes over QRS812, `NNODES=4`, `TP=4` |

| Fabric | Switched NCCL/Gloo to Spark port2 QRS812 fabric: `enP2p1s0f1np1`, `roceP2p1s0f1`, GID `11` |

| Model | Swapped to official `DeepSeek-V4-Flash-0731` weights |

| Runtime | Kept Stage-C NVFP4 DS-MLA KV path |

| Patch 4 | Bind-mounted patched `dspark.py` with `shared_experts.gate_up_proj` mapping |

| Concurrency | Used Keys-style request-stable DSpark slot mapping and ragged mixed prefill/decode handling |

| GR7a C12 profile | Raised `MAX_NUM_SEQS` to `12` for C8/C12 shared-prefix cache-hit tests |

Patch 4 was the big one. When we first brought the 4-node cluster up, the patch wasn’t applied and we didn’t notice — decode sat around `23-26 tok/s` in C=1 long-context tests, and DSpark acceptance dropped to roughly `~25%`, well below the ~60% we normally see on the 2-node setup. After we applied the `shared_experts.gate_up_proj` mapping fix, C=1 decode came back to roughly the `90 tok/s` class and acceptance recovered to the `60-80%` range in typical runs (varies with concurrency — C=1 cold reached ~95%).

## 6. C=1 Benchmark: Cold Start vs KV Cache Hit

Method: one fresh prompt primes the cache, then the exact same prompt is sent again. The cache-hit FFTS number is `prompt_tokens / TTFT`, so it’s effective cached first-token speed, not raw uncached prefill throughput.

| Context | Mode | TTFT | Prefill / Effective FFTS | Decode |

|---  |---          |---       |---             |---         |

| 50K | Cold prime | `19.622s` | `2,548 tok/s` | `93.15 tok/s` |

| 50K | KV cache hit | `0.413s` | `120,977 tok/s` | `84.80 tok/s` |

| 100K | Cold prime | `41.510s` | `2,409 tok/s` | `85.34 tok/s` |

| 100K | KV cache hit | `0.578s` | `173,082 tok/s` | `81.02 tok/s` |

| 150K | Cold prime | `64.433s` | `2,328 tok/s` | `90.15 tok/s` |

| 150K | KV cache hit | `0.776s` | `193,280 tok/s` | `90.30 tok/s` |

Health stayed clean after the run: `/v1/models` healthy, QRS812-facing NIC error counters at zero.

## 7. 150K Benchmark: C=1 / 2 / 4 / 6, Cold vs KV Cache Hit

Target context: `150K` prompt tokens.

Generation: `max_tokens=512`, streaming, `temperature=0`, `ignore_eos=true`.

### Cold Unique-Prompt Wave

| C | Success | TTFT p50 / p95 | Decode p50 | Aggregate TPS | DSpark Acceptance |

|---:|---:    |---:             |---:         |---:        |---:      |

| 1 | `1/1` | `64.153s / 64.153s` | `93.99 tok/s` | `7.21` | `95.71%` |

| 2 | `2/2` | `98.308s / 126.740s` | `41.36 tok/s` | `7.49` | `90.67%` |

| 4 | `3/4` | `134.139s / 193.381s` | `3.64 tok/s` | `1.28` | `77.45%` |

| 6 | `3/6` | `132.949s / 191.495s` | `1.85 tok/s` | `1.28` | `59.96%` |

C=4 and C=6 cold unique-prompt waves hit client-side timeouts. Kept them in anyway — they’re useful negative evidence: simultaneous uncached 150K requests aren’t this config’s sweet spot.

### KV Cache-Hit Average

Average of `cache_hit_1` and `cache_hit_2`.

| C | Success | Avg TTFT p50 | Decode p50 | Aggregate TPS | Spark Accelerator Power |

|---|---    |---       |---             |---     |---         |

| 1 | `2/2` | `0.657s` | `92.18 tok/s` | `82.41` | `145.4 W` |

| 2 | `4/4` | `1.055s` | `64.03 tok/s` | `102.35` | `166.3 W` |

| 4 | `8/8` | `3.265s` | `48.69 tok/s` | `123.08` | `136.7 W` |

| 6 | `12/12` | `4.334s` | `40.39 tok/s` | `165.54` | `169.3 W` |

Prefix/KV cache reuse completely changes the serving profile. C=6 cache-hit finished `12/12` at around `165 tok/s` aggregate output.

## 8. 150K Benchmark: C=8 / 12 KV Cache Hit

Method: one shared 150K C=1 request primes the cache first, then C8 and C12 waves reuse the exact same prompt.

Shared C=1 prime:

| Success | TTFT p50 | Decode p50 | Aggregate TPS |

|---    |---        |---           |---      |

| `1/1` | `65.086s` | `74.46 tok/s` | `6.98` |

Cache-hit waves:

| C | Wave | Success | TTFT p50 / p95 | Decode p50 | Aggregate TPS | DSpark Acceptance |

|---|---|---      |---    |---                |---            |---       |

| 8 | cache_hit_1 | `8/8` | `5.634s / 5.646s` | `29.97 tok/s` | `149.18` | `48.04%` |

| 8 | cache_hit_2 | `8/8` | `5.888s / 5.897s` | `30.59 tok/s` | `159.25` | `53.95%` |

| 12 | cache_hit_1 | `12/12` | `5.199s / 8.218s` | `26.25 tok/s` | `217.85` | `63.85%` |

| 12 | cache_hit_2 | `12/12` | `8.210s / 8.801s` | `25.38 tok/s` | `199.62` | `57.19%` |

Two-wave average:

| C | Success | Avg TTFT p50 / p95 | Avg Decode p50 | Avg Aggregate TPS | Avg Acceptance |

|---|---      |---                |---            |---       |---       |

| 8 | `16/16` | `5.761s / 5.771s` | `30.28 tok/s` | `154.22` | `51.00%` |

| 12 | `24/24` | `6.705s / 8.509s` | `25.82 tok/s` | `208.74` | `60.52%` |

C12 gives ~35% more aggregate throughput than C8, at the cost of higher p95 TTFT. Usable as a high-concurrency cache-hit lane, but it doesn’t replace cold 150K parallel prefill.

## 9. Throughput Comparison: 2x 2-node Spark vs 1x 4-node Spark

| Setup | Topology | Workload Shape | Concurrency | Aggregate Throughput |

|— | — |— |— |— |

| 2-node Spark cluster | TP=2, single replica | historical DSpark code-generation benchmark | C=12 | `230.10 tok/s` |

| 2x 2-node Spark clusters | two independent TP=2 replicas | estimated from two similar replicas | C=12 + C=12 | `~460 tok/s` theoretical aggregate |

| 4-node Spark cluster | TP=4 over QRS812 | 150K shared-prefix KV cache hit | C=8 | `154.22 tok/s` |

| 4-node Spark cluster | TP=4 over QRS812 | 150K shared-prefix KV cache hit | C=12 | `208.74 tok/s` |

| 4-node Spark cluster | TP=4 over QRS812 | 150K cold unique prompts | C=6 | `1.28 tok/s`, partial timeout |

Bottom line: if you can split requests across replicas, the old **2x 2-node layout still wins on raw aggregate decode throughput**. The **4-node TP=4 QRS812 setup** is mainly worth it for a single large shared-memory / long-context serving lane, especially when prefix/KV cache hits apply. Cold parallel 150K unique prompts stay inefficient even on 4 nodes.

## Notes

- Spark power numbers are `nvidia-smi power.draw` sums only, not wall-plug whole-system power.

- QRS812 switch power wasn’t included in the benchmark tables. A manual telemetry check showed about `34.4 W`, but it wasn’t wired into the benchmark sampler.

- All hardware benchmark numbers here are from our private setup; they weren’t public before this post.

Something must be wrong with your 4 node setup. At C12, 2-nodes with TP=2 is 230 while 4-nodes with TP=4 only reaches 209.
So with 50% more compute you got a 10% reduction in T/s.

Useful post — thanks especially for the RDMA latency matrix, that’s the part
I can’t produce from a 2-node direct-attach setup.

For comparison, 2x GB10 at TP=2, direct-attach, no switch, same 0731 + Patch 4 stack:

metric your 4-node TP=4 my 2-node TP=2
decode, C=1 90-94 tok/s 88.3 peak / 73.3 mean
prefill, cold 2,409 tok/s @ 100K 2,644 tok/s @ 78K
KV pool ~256 GiB ~50 GiB = 2.87M tok
build cost ~$20.5K ~$8K

So decode is a tie and prefill is slightly ahead on half the hardware. That’s
consistent with prefill being compute-bound and decode being all-reduce-latency
bound — TP=4 doesn’t help either, and the switch hop costs you roughly what the
4-way weight split gains.

Your actual win is the KV pool: ~5.1x larger, which is what buys ~14 concurrent
1M-context sessions instead of 2.74. That’s the case for going 4-node, and I’d
lead with it — the prefill and decode headlines undersell the build.

One caveat on the C=8/C=12 rows: they’re shared-prefix cache-hit. I measured that
inflating aggregate throughput ~17% at c6 versus distinct prompts on identical
hardware, so those aren’t comparable to anyone’s cold-prompt numbers. It also
explains part of what @mashie is seeing — your 2-node 230 tok/s row is a
short-prompt code-gen benchmark and your 4-node 209 row is a 150K cache-hit run,
so those two aren’t measuring the same thing and the comparison isn’t 50% more
compute for 10% less throughput.

Also curious what MTP depth you’re running — acceptance falling to ~50% at C=8
looks low. I’m seeing 4.86 accepted tokens per decode step here.

I’ve got almost the exact same setup (1xGB10, 3X Sparks, QRS812) and I’m seeing similar results. My only question is what are you using for vision in your stack? I’ve had to cut my memory allocation to Deepseek quite a bit to make room for another model for vision. I’ve been using Qwen3.6 35B-A3B for vision because if I’m giving up resources to run more vLLM containers, I may as well use it for a model that can do vision + some lighter tasks that don’t need Deepseek.

Thanks for posting the full matrix — and @jovan3, your 2-node comparison is what made this worth chasing. We have a 4-node GB10 box too, ran the same TP=2 vs TP=4 question on it, and got a partly opposite answer. I think both results are correct and the difference is instructive, so here’s our data.

Our setup

Item Value
Compute 4x DGX Spark (GB10, sm_121a, 121 GiB unified each)
Switch MikroTik CRS504-4XQ-IN
Fabric 100 Gb/s per port, QSFP28 DAC, RoCE v2 (100 Gb/sec (4X EDR))
Also present 200 Gb/s direct QSFP between each pair — used by the old TP=2 layout only
Model DeepSeek-V4-Flash-0731, fp8 KV, 1M context
Image aidendle94/sparkrun-vllm-ds4-gb10:production-3.75
Speculation DSpark k=3, draft_sample_method: probabilistic
Workload Russian-language contract/legal analysis, reasoning_effort=max — not code

That last row matters. On this model we measured roughly 2x swing in tok/s between code and prose workloads, so please don’t line our absolute numbers up against a code benchmark.

Note the handicap: our old layout ran each TP=2 pair over a 200G direct cable. Going TP=4 forced everything onto the 100G switch — half the per-link bandwidth plus a hop. We expected that to hurt. It didn’t.

Method

Two configs, same image, same flags, only --tensor-parallel-size, --nnodes and the HCA/NIC selection changed:

  • Fleet: 2 independent TP=2 replicas (8 slots each) behind nginx with consistent-hash sticky sessions, clients hit the balancer
  • TP=4: one instance on all four nodes, 16 slots, clients hit the engine directly

Same prompts both sides, fixed output length (min_tokens + ignore_eos — otherwise one long request drags the aggregate), warm-up discarded, 4-6 repeats per point. Our measured noise floor is ~4% on tok/s and ~2 pp on acceptance; anything smaller we don’t claim.

Single stream — TP=4 wins clearly

Metric Fleet TP=4
Decode, reasoning_effort=max 44.4 tok/s 71.2 +60%
Decode, llama-benchy tg256 @ d0 47.6 64.4 +35%
Decode, llama-benchy tg256 @ d8192 46.3 67.5 +46%
Prefill, pp2048 @ d0 2,053 2,615 +27%
Prefill, pp2048 @ d8192 2,160 2,625 +22%
TTFT, prefix-cache hit 0.40 s 0.29 s −27%
DSpark acceptance 57.3% 56.7% unchanged
Accepted tokens per step 2.72 2.70 unchanged

Those last two rows are the whole story. Speculation behaves identically — same acceptance, same accepted tokens per step. The steps themselves got faster. That points at memory bandwidth, not all-reduce latency: with four nodes each one reads a quarter of the weights per decode step instead of a half.

Concurrency — depends entirely on what dominates

Decode-dominated (short prompt, 600-token output):

C Fleet aggregate TP=4 aggregate
1 46.8 70.3
2 83.7 99.9
4 67.0 162.9
8 136.4 260.6
16 229.9 360.0

Prefill-dominated (llama-benchy, 2048-token prompts, 256-token outputs, c4):

Metric Fleet TP=4
Prefill @ d0 3,388 2,507 −26%
Prefill @ d8192 4,023 2,544 −37%
Decode @ d8192 75.7 55.1 −27%
TTFT @ d8192 7,580 ms 11,046 ms +46%

So the fleet is genuinely better at one thing: bursts of concurrent fresh long prompts. Two replicas prefill in parallel on physically separate machines; one TP=4 instance funnels them through a single scheduler. If your traffic is “many new documents at once, short answers”, two replicas remain the right call.

Caveat on that table: variance at c4 was large (TTFT 11,046 ± 4,451 ms), and llama-benchy moved 0.4.0 → 0.4.1.dev1 between our two runs with a changed output format, so treat it as directional.

The setting that decided it: --max-cudagraph-capture-size

This is the part I’d most want others to take away. Our first TP=4 run used our production value of 8, inherited from the TP=2 config:

C capture=8 capture=64
1 69.1 70.3 +2%
2 46.8 52.3 +12%
4 15.3 41.4 +170%
8 13.5 33.6 +148%
16 16.1 23.4 +45%

With speculative decoding the uniform decode batch is max_num_seqs x (1 + k). At k=3 and 16 slots that’s 64 tokens, and capture sizes get rounded to multiples of 1+k. A cap of 8 therefore covers only 1-2 concurrent requests — the moment a third arrives you fall out of CUDA graphs into eager and lose roughly two thirds of decode speed.

64 is exactly the ceiling here. Going higher (128, 512) captures shapes that can never be dispatched. Cost of 8 → 64: startup 8 → 9 min, KV pool unchanged.

Honest caveat: our fleet replicas also ran capture=8, so their multi-stream numbers above are depressed by the same limit. Raising it there would lift them. But at C=1 both configs are fully inside CUDA graphs, and that’s where the +60% lives — no graph tuning gives the fleet a 4-way weight split.

KV pool — the part we underestimated

Fleet TP=4
Pool 2.43M + 2.39M, separate 7.80M, single
Available to one conversation 2.4M 7.8M
Concurrent 1M-context sessions 2.36 7.44

Two pools don’t add. A session lives on one replica and can’t touch the other’s cache, so the practical ceiling for one long conversation was 2.4M. @jovan3 called this out as the real case for 4-node and we now agree — it’s 3.2x more for a single conversation, not 1.6x more in total.

It also killed a recurring annoyance: after restarting one replica its prefix cache was cold, so users hashed onto it saw 22% hit rate while the other sat at 69%. One pool, no split.

Why our conclusion differs from this thread’s

@jovan3’s explanation — decode is all-reduce-latency bound, the switch hop eats what the 4-way split gains — is almost certainly right for that setup. The difference we see is the KV dtype: that build runs nvfp4_ds_mla, roughly half the bytes per token of our fp8. Less KV traffic per step means less bandwidth pressure, so the interconnect becomes the limiter and TP=4 stops paying. We’re on fp8 at a 1M window, firmly bandwidth-bound, so splitting weights four ways helps a lot.

The diagnostic is cheap: watch accepted tokens per step. If it stays flat while tok/s moves, you changed step cost, not speculation — that’s a bandwidth story. Ours didn’t budge (2.72 → 2.70) while decode went up 60%.

Two gotchas worth flagging

tool-eval-bench: the v2.5.0 tag moved. Our August run at v2.5.0 reported 84 scenarios / 168 points. The same pinned tag today gives 69 scenarios / 138 points, nothing skipped. Scores across those two runs are not comparable even though the version string matches. If you track this benchmark over time, record the scenario count, not just the tag.

What is comparable is per-turn latency, and it corroborates everything above: median turn 3.4 s → 2.2 s, seconds per scenario 12.55 → 7.99 (−36%), which is what a +50% decode speedup should produce.

Don’t proxy /metrics through the balancer. If you scrape the nginx address, Prometheus gets whichever replica the hash picked and draws a phantom series. We return 404 for that path in the fleet config.

What we gave up, and why we took the trade

Not free:

  • No rolling restarts. Two replicas let us upgrade one while the other served. Now every restart is ~9 minutes of full downtime.
  • One blast radius — any node takes the whole cluster with it.
  • Cold start 9 min vs 6.

We took it because of this, from 14 days of our own telemetry:

 1 concurrent request : 96.6% of active time
 2                    :  1.0%
 4                    :  1.7%   <- our own benchmarks
16                    :  0.7%   <- our own benchmarks

In real use it’s a single request essentially always, so the fleet’s second replica was idle hardware, and the one workload where two replicas win — concurrent fresh prefill — is 2.4% of our time and even that was us testing. Different traffic shape, different answer.

Config that produced these numbers

--tensor-parallel-size 4 --nnodes 4 --node-rank N --master-port 29501
--kv-cache-dtype fp8 --block-size 256 --enable-prefix-caching
--max-model-len 1048576 --max-num-seqs 16 --max-num-batched-tokens 8192
--max-cudagraph-capture-size 64
--compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE","custom_ops":["all"]}'
--speculative-config '{"method":"dspark","num_speculative_tokens":3,"draft_sample_method":"probabilistic"}'
--attention-backend FLASHINFER_MLA_SPARSE_DSV4 --moe-backend b12x
--disable-custom-all-reduce --no-scheduler-reserve-full-isl --async-scheduling
--enable-chunked-prefill --gpu-memory-utilization 0.85

Fabric env on all four ranks, pointed at the CRS504-facing port:

NCCL_NET=IB  NCCL_IB_DISABLE=0  NCCL_IB_HCA=rocep1s0f0
NCCL_SOCKET_IFNAME=enp1s0f0np0  GLOO_SOCKET_IFNAME=enp1s0f0np0
NCCL_IB_GID_INDEX=<computed at runtime>  NCCL_CUMEM_ENABLE=0  NCCL_NVLS_ENABLE=0

Two operational notes:

  1. Compute the RoCE GID index at runtime. A hardcoded index broke our cluster twice after IP reassignment. Walk /sys/class/infiniband/<hca>/ports/1/gids/ and match your own address.
  2. Run the launch command locally on each node. The JSON arguments don’t survive multiple layers of SSH quoting.

Quality was verified at 16 slots with CUDA graphs on: 8 concurrent requests, no degenerate output, arithmetic correct. Worth checking explicitly — that failure mode leaves the speculation metrics looking perfectly healthy.

Happy to run anything specific if it would help nail down the bandwidth-vs-latency question. An nvfp4_ds_mla run on our box would settle it, but that KV path isn’t in our image.