Here’s my version, if anyone is interested.
It’s tuned to my workload and setup on a 3x switchless FE setup.
Here’s my version, if anyone is interested.
It’s tuned to my workload and setup on a 3x switchless FE setup.
DeepSeek V4.1 Flash on three DGX Sparks: weekend release (+16% single-stream decode, +14% prefill, 123 s cold start)
This is an update on my three-Spark DeepSeek V4.1 Flash setup:
Seven changes were promoted over the weekend, and all of them are on main:
| Measurement | Friday | Now | Change | Improved? |
|---|---|---|---|---|
| Decode, 1 stream, code (reasoning on) | 54.0 tok/s | 62.6 tok/s | +15.9% | yes |
| Decode, 1 stream, prose (reasoning on) | 43.9 tok/s | 49.7 tok/s | +13.2% | yes |
| First token, short prompt, 1 stream | 0.24–0.26 s | 0.20–0.22 s | −14% | yes |
| Prefill, 32K cold (filler text) | 4.2k tok/s | 4.8k tok/s | +14% | yes |
| 32K prefix replay, cold | 7.59 s | 6.65 s | −12% | yes |
| Four concurrent 64K contexts | 12.7 tok/s each | 13.6 tok/s each | +7% | yes |
| Cold start to API ready | 150 s | 123 s | −18% | yes |
| Lowest free memory under load (dgx1) | 6.31 GiB | 6.44 GiB | +0.13 GiB | yes, but |
| Quality gate (fixed task, 5 runs) | 5/5 | 5/5 | — | no worse |
Decode, aggregate tok/s across all streams:
| Workload | 1 stream | 2 | 4 | 8 |
|---|---|---|---|---|
| JSON, answer only | 78.5 | 106.8 | 164.2 | 241.1 |
| Code, answer only | 78.9 | 112.2 | 169.3 | 226.2 |
| Code, reasoning on | 62.6 | 89.4 | 137.0 | 188.4 |
| Prose, answer only | 55.9 | 84.0 | 126.4 | 172.2 |
| Prose, reasoning on | 49.7 | 74.9 | 110.3 | 164.0 |
Method: measured with bin/spark3 bench:
Speculative decoding
Prefill
Runtime
torch._library.utils.fill_defaults rereads the op schema once for every argument, so its cost grows with the square of the argument count. A one-read replacement cuts a 64-argument custom op from 873 µs to 48 µs per call, and this applies to every MoE launch outside a CUDA graph.Operations
Memory: all changes were qualified against the stack’s memory guards (5 GiB free at startup, 3 GiB under load). The tightest node ends the weekend with slightly more headroom than it started with.
fill_defaults fix is a good candidate for upstream PyTorch, and the vLLM patches are being prepared for upstream review.**A note on comparisons:** benchmarking methods vary, including the prompt set, whether reasoning is on, filler versus real text, prefix-cache state, and whether the best or the mean run is reported. As a result, these numbers will generally read lower than figures posted for other builds, even on identical hardware.
For example, running another published benchmark’s protocol on an earlier build of this stack gave 95 tok/s for single-stream code and 359 tok/s at eight streams on a counting task, against 63 and 241 tok/s with the method above.
Thanks to Local Inference Lab for the vLLM and B12X branches this builds on. Questions, reproductions and comparison runs are welcome.
DeepSeek V4.1 Flash on three DGX Sparks: 512K context, 4.9× KV cache, prefill flat to 249K tokens, and a 64 KiB memory saver
This is an update to Monday’s release of my three-Spark DeepSeek V4.1 Flash setup:
Five changes were promoted to main since Monday (r5k–r5o). They grow the context and KV cache, speed up long-prompt prefill, and fix several correctness problems. A driver patch, Memory Saver, recovers memory on a 64 KiB-page kernel, and that 64 KiB profile is now the production default. Together they take the context limit from 131K to 512K tokens. Decode is level to slightly faster.
Monday → now
| Measurement | Monday (r5j) | Now (r5o, 64 KiB profile) | Change | Improved? |
|---|---|---|---|---|
| Context limit per request | 131,072 tokens | 524,288 tokens | 4× | yes |
| KV cache capacity | 575,304 tokens (4.4 full windows) | 2,845,543 tokens (5.4 full 512K windows) | 4.9× | yes |
| Longest real-text prefill measured | 61K tokens at 3.78k tok/s | 249K tokens at 3.68k tok/s | 4× the length, −2% | yes |
| Real-text prefill, about 60K tokens | 3.78k tok/s | 3.80k tok/s | level | — |
| 32K prefix replay, cold | 6.65 s | 6.55 s | −1.4% | yes |
| Four concurrent long contexts | four 64K at 13.6 tok/s each | four 485K at 22.4 tok/s each, zero preemptions | new capability | yes |
| Decode, 1 stream, code (reasoning on) | 62.6 tok/s, 48.3 ms step | 64.8 tok/s, 45.4 ms step | +3.5% | slightly |
| Decode, 1 stream, prose (reasoning on) | 49.7 tok/s, 42.3 ms step | 52.8 tok/s, 40.8 ms step | +6.2% | slightly |
| Decode, 8 streams, code / prose | 188.4 / 164.0 tok/s | 195.8 / 171.2 tok/s | +4% | slightly |
| First token, short prompt, 1 stream (prose / code) | 203 / 221 ms | 197 / 206 ms | −3 to −7% | yes |
| Worst decode/prefill disagreement (logprob) | 8.23 nats | 2.54 nats | −69% | yes |
| Prefill argmax disagrees with decode | 3.62% of tokens | 2.84% of tokens | −22% | yes |
| Wrong results beside co-resident kernels (stress test) | up to 609 / 12,000 calls | 0 / 12,000 | fixed | yes |
| Indexer selections, same prompt twice | 78% of rows differ (one layer) | identical | fixed | yes |
| Lowest free memory under load (dgx1) | 6.44 GiB with four 64K contexts | 5.83 GiB with four 485K contexts | 7.6× the context resident | yes, but |
| Quality gate (fixed task, 5 runs) | 5/5 | 5/5 | — | no worse |
| Needle in a haystack | — | 3/3 at 519,142 tokens | new | — |
Notes on the table:
Long-context prefill (real text, cold)
| Prompt | Monday (r5j) | Before the indexer split (r5k) | After the split (r5l, 4 KiB) | First token (r5l) |
|---|---|---|---|---|
| 3.7K tokens | 3.83k tok/s | 3.87k | 3.90k | 0.94 s |
| 13.4K | 3.84k | 3.73k | 3.79k | 3.6 s |
| 29.4K | 3.81k¹ | 3.74k | 3.83k | 7.7 s |
| 57.9K | 3.78k¹ | 3.69k | 3.83k | 15.3 s |
| 116.7K | over the limit | 3.56k | 3.80k | 30.5 s |
| 183.1K | over the limit | 3.42k | 3.76k | 48.6 s |
¹ Monday’s real-text prompts were 36.3K and 61.4K tokens; the source text has grown since.
Prefill used to fall off with context depth. It now holds about 3.8k tok/s from 4K to 183K tokens. On the 64 KiB production profile it measured 3.77k tok/s at 28.8K tokens, 3.80k at 59.7K and 3.68k at 249.2K (66 s to first token).
Method
The method is the same as Monday’s: bin/spark3 bench, temperature 0, 256 output tokens, means of at least three samples, reasoning tokens counted as output, and prefill measured cold on both filler and real text.
What changed
Context and memory
Long-context prefill
Quality and correctness
Decode/prefill consistency (r5k). The ratio-2 compressor ring is now written during decode, so compressed entries for generated tokens match what prefill builds. The FP4 KV writer also rounds like DeepSeek’s reference quantizer (B12X #435). Together these cut the worst decode-versus-prefill disagreement from 8.2 to 2.5 nats.
Exact indexer selections (r5m). Indexer top-k now breaks exact score ties by lowest position, as DeepSeek’s reference does. Before, two identical runs chose different positions for 78% of one layer’s prefill rows.
Fenced pipeline stages (r5n, r5o). Several B12X kernels released a pipeline stage for its TMA refill while shared-memory reads from that stage could still be pending. When another kernel shared the SM, the refill could overwrite data before it was read. The effects were:
All are now fenced, at 0 / 12,000, and a SASS check rejects any stage release with shared loads still outstanding. The cost was within 0.3–0.7%.
Memory Saver (experimental; our default, opt-in for everyone else)
Testing
Next
fill_defaults change.On comparisons
The caveat from Monday’s post still applies. These figures use real prompts with reasoning on, count rejected drafts, and report means rather than the best run, so they will read lower than numbers measured other ways on the same hardware.
Thanks again to Local Inference Lab for the vLLM and B12X branches this builds on. Questions, reproductions and comparison runs are welcome.
Sounds nice! Thanks for the job! That makes me think about buying a third ;)
This DeepSeek V4.1 Flash deployment on DGX Spark has a new default. With r6, the model’s compute kernels are TileLang kernels written for SM121, with DeepSeek’s own TileKernels underneath and one-shot RoCE collectives from my sparknet library. They replace B12X’s kernels. On the same hardware, recipe and benchmark, r6 is never behind B12X, and it is ahead on single-stream step time, at high concurrency and in prefill up to 1M tokens.
Repository: GitHub - christopherowen/spark-ds41f: Reproducible Docker/vLLM deployment, tuning, and benchmarks for DeepSeek V4.1 Flash on a switchless three-node DGX Spark fabric. · GitHub (the project formerly known as spark3-vllm-ds41f)
Four Sparks (TP4, 1M context, 16 sequences)
Aggregate tok/s, mean ± 95% interval. Bold marks a point whose interval does not overlap the other family’s.
| Streams | 1 | 2 | 4 | 8 | 16 |
|---|---|---|---|---|---|
| Prose, B12X r5p | 62.1 ± 3.8% | 92.7 ± 8.0% | 141.6 ± 7.9% | 215.7 ± 7.5% | 297.3 ± 1.2% |
| Prose, TileLang r6 | 62.4 ± 1.2% | 100.2 ± 2.7% | 147.1 ± 10.2% | 222.1 ± 0.3% | 318.4 ± 1.6% |
| Code, B12X r5p | 72.5 ± 7.5% | 111.3 ± 5.4% | 168.3 ± 9.6% | 245.7 ± 1.3% | 323.5 ± 1.5% |
| Code, TileLang r6 | 76.0 ± 0.6% | 116.8 ± 18.9% | 174.9 ± 0.6% | 264.3 ± 0.6% | 349.8 ± 2.5% |
| Prefill, real source text (tok/s) | 32K | 256K | 500K | 1M |
|---|---|---|---|---|
| B12X r5p | 5,117 ± 1.7% | 4,873 ± 2.5% | 4,645 ± 1.7% | 4,026 ± 0.9% |
| TileLang r6 | 5,836 ± 4.2% | 5,471 ± 3.1% | 5,072 ± 1.3% | 4,370 ± 0.6% |
Single-stream steps are 31.79 ms (prose) and 34.75 ms (code), against B12X’s 34.00 and 37.72: 6.5% and 7.9% shorter.
Three Sparks (TP3, 8 sequences at 512K)
| Streams | 1 | 2 | 4 | 8 |
|---|---|---|---|---|
| Prose, B12X r5p | 50.4 ± 12.2% | 75.9 ± 10.6% | 118.8 ± 12.9% | 173.3 ± 8.4% |
| Prose, TileLang r6 | 52.2 ± 0.7% | 84.4 ± 0.6% | 125.7 ± 0.5% | 198.6 ± 0.4% |
| Code, B12X r5p | 60.1 ± 2.0% | 90.8 ± 5.0% | 136.3 ± 6.3% | 192.3 ± 3.2% |
| Code, TileLang r6 | 63.0 ± 0.5% | 97.3 ± 8.5% | 159.2 ± 44.2% | 240.3 ± 1.0% |
| Prefill, real source text (tok/s) | 32K | 256K | 500K |
|---|---|---|---|
| B12X r5p | 3,778 ± 14.9% | 3,632 ± 2.2% | 3,432 ± 0.6% |
| TileLang r6 | 4,306 ± 20.9% | 4,034 ± 0.3% | 3,743 ± 0.8% |
Single-stream steps are 38.21 ms and 42.43 ms, against 41.92 and 46.44: 8.9% and 8.6% shorter. Eight streams decode 15% more prose and 25% more code.
Both families run the same native FP8/FP4/BF16 weights, the same DSpark speculative decoding (five drafts), the same memory budget and the same benchmark: quality gate, prose and code with reasoning on, temperature 0, 256 output tokens, three samples per point, real source text for prefill with two repeats. Each family uses DSpark cost curves profiled for its own kernels. The r6 numbers were measured on 2026-10-05, TP4 on the four-node ring and TP3 after recabling the triangle the same afternoon. B12X’s are r5p’s acceptance benchmark from the same day, with its 500K point from a run that morning.
Reading the intervals
B12X’s intervals are wider because its output text changes from sample to sample, and DSpark acceptance moves with the text. TileLang produces the same text every sample. Its two wide points, code at two and four streams, come from streams interleaving differently between samples.
An earlier TP4 run this morning showed TileLang behind at two-stream code (105.0 against 111.3). That point is one fixed text, and acceptance on that one text was lower. Across 13 texts at TP3, acceptance was level (2.401 against 2.384 accepted drafts per step), with about ±5% spread per text. In r6 the same point reads 116.8 ± 18.9%, level within its interval. The intervals cover sample-to-sample spread on one text, not text-to-text acceptance. Single-stream step time doesn’t depend on acceptance, so it is the cleanest kernel comparison.
Why TileLang
TileLang (tile-ai/tilelang) is a Python DSL for writing GPU kernels at the tile level. You state the tiles, the shared-memory staging, the software pipeline and the tensor-core MMAs explicitly. TileLang lowers that through TVM to CUDA, JIT-compiles it and caches the result. You keep control of layout and data movement without a C++ template stack, and a kernel is a few hundred readable lines.
What made it worth trying on the Spark is that DeepSeek writes production kernels in it. Their TileKernels library describes itself as “dozens of highly optimized kernels” used in their internal training and inference. The v2.0.0 release (deepseek-ai/TileKernels#34, 2026-09-30, 229 files) covers:
TileKernels lists SM90 and SM100 as its targets. Everything I use from it runs unmodified on GB10 (SM121).
Two properties come with it:
How it works
The compiler. TileLang 0.1.15, built from source with three patches for the SM120 family:
sm_121a;mma.sync.kind::mxf8f6f4.block_scale for MXFP8 and FP8 × FP4;kind::mxf4nvf4.block_scale.scale_vec::2X for MXFP4.They are on GitHub - christopherowen/tilelang: Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels · GitHub (branch deepseek-v41-sm120), not upstream yet.
The DS4.1 kernels, as a vLLM kernel backend:
TileKernels supplies the MoE top-k gate, the SwiGLU and quantization casts, and mHC’s post and stream collapse. DeepSeek moved the mHC projection itself to DeepGEMM, which doesn’t run on SM12x. That projection is a TileLang kernel of mine, with the V4.1 lagged pre-mix and the stream update folded in registers.
The collectives come from sparknet (below).
The switch. The configuration’s kernel_backend names the family for the whole model, drafter included, and bin/spark doctor checks that the backend flags and environment match it. Everything else matches the B12X configuration field by field, so the comparison isolates the kernels.
Kernel by kernel against B12X
Rank-0 decode profile at TP4, one stream, median six-row verification step, both families on the same four Sparks. The whole step is 39.3 ms against 44.9 ms (GPU busy 30.0 against 36.8 ms).
| Kernel group | B12X | TileLang | Change |
|---|---|---|---|
| Routed experts | 21.05 ms | 16.50 ms | −22% |
| Shared expert (side stream, per call) | 52.7 + 46.7 µs | 10.0 + 19.1 µs | −71% |
| Collectives | 4.57 ms | 4.15 ms | −9% |
| mHC | 1.91 ms | 1.55 ms | −19% |
| Sparse MLA | 1.22 ms | 1.10 ms | −10% |
| Router | 0.59 ms | 0.45 ms | −24% |
| Quantization and norms | 0.44 ms | 0.32 ms | −28% |
| Fused Q-A/KV (per call) | 18.0 µs | 14.2 µs | −21% |
| DSpark main projection (per call) | 175.6 µs | 173.2 µs | −1.4% |
| Q-B (per call) | 20.8 µs | 21.9 µs | +5% |
| Indexer Q-B (per call) | 34.3 µs | 35.2 µs | +3% |
Q-B and the indexer’s Q-B still trail slightly in serving, although both win in isolation: warm in L2, cold from DRAM, and racing their own L2 prefetch. Three different Q-B tiles all measured 21.8–21.9 µs in serving, so the tile isn’t the limit. B12X stores weights tile-packed, so each tile is one contiguous read; TileLang reads row-strided weights. That layout is the next thing to try.
The shared Triton kernels (page mapping, chunk metadata, the attention output quantizer) read slightly slower in TileLang’s profile. They’re the same kernels: the L2 weight prefetch stream overlaps them more. With the prefetch off they run at or below B12X’s times, and the prefetch saves about 1.7 ms per step overall.
A race in TileLang’s warp-specialized pipeline
The new 16-row decode tiles failed a repeated bit check against the prefill kernel. Identical inputs gave different and wrong results, in up to every run for some tile shapes. The cause is TileLang’s warp-specialized pipeline on SM121, where producer warps stage the block scales with cp.async for consumer warps. TileLang’s unspecialized pipeline never failed in any configuration and is just as fast for decode. So decode tiles use it, while prefill keeps warp specialization, which is up to 17% faster there. Adding memory clobbers to TileLang’s mbarrier assembly did not fix it, so compiler hoisting isn’t the cause. The 64-row tiles in the earlier image showed no failure in 18,000 runs. I’ll report it upstream with the reproducer from the experiment.
Numerics and determinism
On TP3, the two families score the same 8,188 tokens equally:
| Measure | TileLang | B12X |
|---|---|---|
| Mean NLL | 1.5525 | 1.5527 |
| Top-1 accuracy | 67.84% | 68.04% |
| Decode against its own prefill: logprob gap | 0.0469 | 0.0565 |
| Decode against its own prefill: argmax differs | 2.59% | 3.32% |
Quality gates passed 5/5 in every window.
B12X is still in the image, with every non-determinism and race fix from the last weeks in the tree:
Switching back is a configuration change (r5p’s B12X configuration is kept in the repository). I’ll likely retire B12X in the next release, once the remaining kernels below are ported, so maybe grab the patches now for your recipe if you want them.
Collectives moved to dgx-spark-networking
The one-shot RoCE collectives (formerly RoCEnante inside B12X) now live in their own public repository, dgx-spark-networking (Python package sparknet): GitHub - christopherowen/dgx-spark-networking: Switchless RoCE collectives, NCCL profiles and fabric tooling for DGX Spark inference recipes · GitHub
It has the one-shot all-reduce and all-gather for direct-cabled Sparks, plus NCCL profiles for 2-, 3- and 4-node and switched fabrics. The TileLang family uses it at TP3 (direct triangle) and TP4 (ring), and nothing in it is specific to this model.
Using it
Both recipes are promoted configurations in config/:
config/cluster.json (64 KiB pages; config/cluster-4k.json for a 4 KiB kernel). docs/replicate.md walks through the hosts, the image build and the first start.config/cluster-tp4.json, selected with --cluster-config config/cluster-tp4.json.Four Sparks, briefly:
config/examples/nodes-ring4.json as config/nodes-ring4.local.json: ranks in ring order, management IPs, and each peer’s RoCE devices.config/cluster-tp4.json (the promoted TP4 profile) for your site: head address, home directory and interface names. Then run bin/spark --cluster-config config/cluster-tp4.json doctor.bin/spark build prepare && bin/spark build image --apply), copy the image to the other nodes, download the model on each, then bin/spark --cluster-config config/cluster-tp4.json cluster start --apply.The TP4 section of docs/replicate.md has the same steps.
What is still missing
bin/spark tuning still speaks B12X’s transport settings; teach it sparknet’s.On the experience
Compared with the FlashInfer and CUTLASS development I did for gpt-oss-120b back in the day, this was a pleasure. TileLang kernels are short Python I can read and change in an afternoon, and the generated CUDA is there to inspect when an MMA, or a pipeline, does something unexpected. Most of the time went into the model, not the toolchain.
Credits
The experiments, configurations and raw benchmark reports are in the repository:
experiments/2026-10-05-tilelang-r6 for r6 and both benchmarks;experiments/2026-10-05-tilelang-decode-kernels for the kernel-by-kernel work, the race census and the profiles;2026-10-03-tilelang-kernels through 2026-10-04-tilelang-1m for the module work.Setup: DGX Spark (GB10), driver 580.178.04, kernel 7.0.0-1019-nvidia-64k, DeepSeek V4.1 Flash native weights. Benchmarks are temperature 0, 256 output tokens, reasoning on, three samples per point, aggregate throughput, 95% intervals; prefill is real source text, two repeats.
Questions, and results from other Spark setups, are welcome.
Thanks for keeping the TP3 configuration supported! As soon as my third node arrives, this recipe is going to be the very first thing I try out. Appreciate the detailed write-up!
BTW, DSV4F TP3 config is supported by spark-vllm-docker too if you happen to use it.
Thank you for contributing to my new deepseek 4.1 flash recipe thread :).
Did you give it a shot?
Please clarify that your sparknet library’s one-shot collectives are a vendored derivative of b12x’s RoCEnante, with your switchless-routing extensions and subsequent backend work. RoCEnante hasn’t moved out of b12x. The repository documents this relationship, but the forum announcement currently reads as though the upstream project relocated.
Personally I have no issue with forks (or I wouldn’t have picked the license I did), but renaming the project, declaring our pun “dead” and then engaging in multiple rounds of self-promotion on social media is in astonishingly bad taste.
RoCEnante (and the rest of b12x) will be moving to FlashInfer officially today/tomorrow, and we’ll begin the full upstream integration process shortly thereafter.