# New Deepseek 4.1 flash recipe

**URL:** <https://forums.developer.nvidia.com/t/new-deepseek-4-1-flash-recipe/384438>\
**Category:** DGX Spark / GB10\
**Tags:** deepseek\
**Created:** [September 26, 2026, 5:19pm UTC](https://forums.developer.nvidia.com/t/new-deepseek-4-1-flash-recipe/384438 "2026-09-26T17:19:34Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![christopher\_owen](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/christopher_owen/32/469895_2.png) [@christopher\_owen](https://forums.developer.nvidia.com/u/christopher_owen)\
**Post date:** [September 26, 2026, 5:19pm UTC](https://forums.developer.nvidia.com/t/new-deepseek-4-1-flash-recipe/384438/1 "2026-09-26T17:19:34Z")

</div>

Here’s my version, if anyone is interested.

It’s tuned to my workload and setup on a 3x switchless FE setup.

> **[GitHub - christopherowen/spark3-vllm-ds41f: Reproducible Docker/vLLM deployment, tuning, and...](https://github.com/christopherowen/spark3-vllm-ds41f)**
>
> Reproducible Docker/vLLM deployment, tuning, and benchmarks for DeepSeek V4.1 Flash on a switchless three-node DGX Spark fabric.

---

<div class="post-metadata">

**Author:** ![christopher\_owen](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/christopher_owen/32/469895_2.png) [@christopher\_owen](https://forums.developer.nvidia.com/u/christopher_owen)\
**Post date:** [September 28, 2026, 6:02am UTC](https://forums.developer.nvidia.com/t/new-deepseek-4-1-flash-recipe/384438/2 "2026-09-28T06:02:53Z")

</div>

**DeepSeek V4.1 Flash on three DGX Sparks: weekend release (+16% single-stream decode, +14% prefill, 123 s cold start)**

This is an update on my three-Spark DeepSeek V4.1 Flash setup:

- native FP4/FP8 weights
- tensor parallel 3 over RoCE
- vLLM with B12X CuTe DSL kernels
- DSpark speculative decoding

Seven changes were promoted over the weekend, and all of them are on main:

> **[GitHub - christopherowen/spark3-vllm-ds41f: Reproducible Docker/vLLM deployment, tuning, and...](https://github.com/christopherowen/spark3-vllm-ds41f)**
>
> Reproducible Docker/vLLM deployment, tuning, and benchmarks for DeepSeek V4.1 Flash on a switchless three-node DGX Spark fabric.

### Friday → Monday

| Measurement | Friday | Now | Change | Improved? |
| --- | --- | --- | --- | --- |
| Decode, 1 stream, code (reasoning on) | 54.0 tok/s | 62.6 tok/s | +15.9% | yes |
| Decode, 1 stream, prose (reasoning on) | 43.9 tok/s | 49.7 tok/s | +13.2% | yes |
| First token, short prompt, 1 stream | 0.24–0.26 s | 0.20–0.22 s | −14% | yes |
| Prefill, 32K cold (filler text) | 4.2k tok/s | 4.8k tok/s | +14% | yes |
| 32K prefix replay, cold | 7.59 s | 6.65 s | −12% | yes |
| Four concurrent 64K contexts | 12.7 tok/s each | 13.6 tok/s each | +7% | yes |
| Cold start to API ready | 150 s | 123 s | −18% | yes |
| Lowest free memory under load (dgx1) | 6.31 GiB | 6.44 GiB | +0.13 GiB | yes, but |
| Quality gate (fixed task, 5 runs) | 5/5 | 5/5 | — | no worse |

### Current numbers

Decode, aggregate tok/s across all streams:

| Workload | 1 stream | 2 | 4 | 8 |
| --- | --- | --- | --- | --- |
| JSON, answer only | 78.5 | 106.8 | 164.2 | 241.1 |
| Code, answer only | 78.9 | 112.2 | 169.3 | 226.2 |
| Code, reasoning on | 62.6 | 89.4 | 137.0 | 188.4 |
| Prose, answer only | 55.9 | 84.0 | 126.4 | 172.2 |
| Prose, reasoning on | 49.7 | 74.9 | 110.3 | 164.0 |

- **Time to first token (short prompts):** 178–221 ms at one stream and 443–568 ms at eight.
- **Real-text prefill (Python source, cold):** a flat 3.8k tok/s from 4K to 61K tokens, so a 61K-token prompt reaches its first token in 15.8 s.

**Method:** measured with `bin/spark3 bench`:

- temperature 0 and 256 output tokens per request
- each figure is the mean of at least three samples
- reasoning tokens count as output
- prefill is measured cold, both on repeated filler text and on real text

### What changed

**Speculative decoding**

- **Newer vLLM and B12X, patched NCCL, 5 drafts:** rebased onto Local Inference Lab’s latest vLLM and B12X branches with a patched NCCL, and raised the draft length to 5 tokens. Code answers now accept about 3.1 drafts per verification step.
- **Dead verification rows:** draft rows whose estimated chance of surviving verification is below 0.2 skip the MoE expert weight reads.
- **Cheaper drafting:** greedy drafting stays sharded by vocabulary, and the drafter head and Markov table run in NVFP4. Together these save 3.7 ms per step.

**Prefill**

- **Sequence-parallel prefill:** each node handles a third of the rows for the per-row work in the encoder layers (hyper-connection mixes, norms, the Engram gate and the front of attention). Prefill is 9–14% faster.
- **When it switches on:** the threshold now comes from the RoCE all-reduce transport’s size limit, so sequence parallelism engages for prompts of 205 tokens or more at TP3.

**Runtime**

- **Custom-op call overhead:** PyTorch’s `torch._library.utils.fill_defaults` rereads the op schema once for every argument, so its cost grows with the square of the argument count. A one-read replacement cuts a 64-argument custom op from 873 µs to 48 µs per call, and this applies to every MoE launch outside a CUDA graph.
- **CuTe DSL 4.7.1:** every consumer now uses 4.7.1. B12X’s pin was moved forward rather than downgrading anything else. This also lets the SM121 L2 weight-prefetch path compile; it had been silently inactive before.
- **Driver JIT cache:** the cache now persists, so the cuBLASLt kernels for the vision tower’s startup profiling pass are no longer rebuilt at every boot.
- **Warmup:** the two Triton kernels used by speculative-decoding verification now compile during warmup. No kernel JIT-compiles after the service reports ready.

**Operations**

- **Launcher:** it reuses one SSH connection per node and checks readiness every second. Cold start fell from 150 s to 123 s.

**Memory:** all changes were qualified against the stack’s memory guards (5 GiB free at startup, 3 GiB under load). The tightest node ends the weekend with slightly more headroom than it started with.

### Known issues and next steps

- **First request after a restart:** it still takes about 0.37 s longer to its first token than later requests (about 580 ms against about 210 ms). JIT compilation is ruled out; the cause is under investigation.
- **Dense FP8 GEMMs:** they reach about 160 GB/s of an achievable ~236 GB/s. This is the largest remaining decode lever.
- **Single-stream decode:** it still trails four-node setups.
- **Upstreaming:** the `fill_defaults` fix is a good candidate for upstream PyTorch, and the vLLM patches are being prepared for upstream review.

\*\*A note on comparisons:\*\* benchmarking methods vary, including the prompt set, whether reasoning is on, filler versus real text, prefix-cache state, and whether the best or the mean run is reported. As a result, these numbers will generally read lower than figures posted for other builds, even on identical hardware.

For example, running another published benchmark’s protocol on an earlier build of this stack gave 95 tok/s for single-stream code and 359 tok/s at eight streams on a counting task, against 63 and 241 tok/s with the method above.

Thanks to Local Inference Lab for the vLLM and B12X branches this builds on. Questions, reproductions and comparison runs are welcome.

---

<div class="post-metadata">

**Author:** ![christopher\_owen](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/christopher_owen/32/469895_2.png) [@christopher\_owen](https://forums.developer.nvidia.com/u/christopher_owen)\
**Post date:** [October 2, 2026, 11:34am UTC](https://forums.developer.nvidia.com/t/new-deepseek-4-1-flash-recipe/384438/4 "2026-10-02T11:34:13Z")

</div>

**DeepSeek V4.1 Flash on three DGX Sparks: 512K context, 4.9× KV cache, prefill flat to 249K tokens, and a 64 KiB memory saver**

This is an update to Monday’s release of my three-Spark DeepSeek V4.1 Flash setup:

- native FP4/FP8 weights
- tensor parallel 3 over RoCE
- vLLM with B12X CuTe DSL kernels
- DSpark speculative decoding

Five changes were promoted to main since Monday (r5k–r5o). They grow the context and KV cache, speed up long-prompt prefill, and fix several correctness problems. A driver patch, Memory Saver, recovers memory on a 64 KiB-page kernel, and that 64 KiB profile is now the production default. Together they take the context limit from 131K to 512K tokens. Decode is level to slightly faster.

**Monday → now**

| Measurement | Monday (r5j) | Now (r5o, 64 KiB profile) | Change | Improved? |
| --- | --- | --- | --- | --- |
| Context limit per request | 131,072 tokens | 524,288 tokens | 4× | yes |
| KV cache capacity | 575,304 tokens (4.4 full windows) | 2,845,543 tokens (5.4 full 512K windows) | 4.9× | yes |
| Longest real-text prefill measured | 61K tokens at 3.78k tok/s | 249K tokens at 3.68k tok/s | 4× the length, −2% | yes |
| Real-text prefill, about 60K tokens | 3.78k tok/s | 3.80k tok/s | level | — |
| 32K prefix replay, cold | 6.65 s | 6.55 s | −1.4% | yes |
| Four concurrent long contexts | four 64K at 13.6 tok/s each | four 485K at 22.4 tok/s each, zero preemptions | new capability | yes |
| Decode, 1 stream, code (reasoning on) | 62.6 tok/s, 48.3 ms step | 64.8 tok/s, 45.4 ms step | +3.5% | slightly |
| Decode, 1 stream, prose (reasoning on) | 49.7 tok/s, 42.3 ms step | 52.8 tok/s, 40.8 ms step | +6.2% | slightly |
| Decode, 8 streams, code / prose | 188.4 / 164.0 tok/s | 195.8 / 171.2 tok/s | +4% | slightly |
| First token, short prompt, 1 stream (prose / code) | 203 / 221 ms | 197 / 206 ms | −3 to −7% | yes |
| Worst decode/prefill disagreement (logprob) | 8.23 nats | 2.54 nats | −69% | yes |
| Prefill argmax disagrees with decode | 3.62% of tokens | 2.84% of tokens | −22% | yes |
| Wrong results beside co-resident kernels (stress test) | up to 609 / 12,000 calls | 0 / 12,000 | fixed | yes |
| Indexer selections, same prompt twice | 78% of rows differ (one layer) | identical | fixed | yes |
| Lowest free memory under load (dgx1) | 6.44 GiB with four 64K contexts | 5.83 GiB with four 485K contexts | 7.6× the context resident | yes, but |
| Quality gate (fixed task, 5 runs) | 5/5 | 5/5 | — | no worse |
| Needle in a haystack | — | 3/3 at 519,142 tokens | new | — |

Notes on the table:

- “Now” is the current production profile: r5o on the 64 KiB kernel with Memory Saver. It was measured with the same harness and prompts as Monday’s reference, three samples per point. That run covered prose and code with reasoning on and real-text prefill; the answer-only cases and filler-text prefill were not rerun.
- The consistency, correctness and repeatability rows come from the r5k–r5o promotion tests. The 64 KiB profile changes memory, not arithmetic.
- Decode tok/s depends on draft acceptance, and the quality fixes changed the generated text. Step time is the cleaner comparison: 6% lower for code and 3% lower for prose.
- The long-context row uses 4,096-token replies; Monday’s 64K test used 256.

**Long-context prefill (real text, cold)**

| Prompt | Monday (r5j) | Before the indexer split (r5k) | After the split (r5l, 4 KiB) | First token (r5l) |
| --- | --- | --- | --- | --- |
| 3.7K tokens | 3.83k tok/s | 3.87k | 3.90k | 0.94 s |
| 13.4K | 3.84k | 3.73k | 3.79k | 3.6 s |
| 29.4K | 3.81k¹ | 3.74k | 3.83k | 7.7 s |
| 57.9K | 3.78k¹ | 3.69k | 3.83k | 15.3 s |
| 116.7K | over the limit | 3.56k | 3.80k | 30.5 s |
| 183.1K | over the limit | 3.42k | 3.76k | 48.6 s |

¹ Monday’s real-text prompts were 36.3K and 61.4K tokens; the source text has grown since.

Prefill used to fall off with context depth. It now holds about 3.8k tok/s from 4K to 183K tokens. On the 64 KiB production profile it measured 3.77k tok/s at 28.8K tokens, 3.80k at 59.7K and 3.68k at 249.2K (66 s to first token).

**Method**

The method is the same as Monday’s: `bin/spark3 bench`, temperature 0, 256 output tokens, means of at least three samples, reasoning tokens counted as output, and prefill measured cold on both filler and real text.

**What changed**

_Context and memory_

- **Display carve-out (r5k).** The GB10 firmware reserves about 2.1 GiB at the top of memory for display scanout, and ordinary allocations never use it. The embedding and output-head weights (842.5 MiB per rank) now live there, imported through a DRM dumb buffer and dma-buf. That frees ordinary memory for the KV cache, which grows from 1.40 to 2.20 GiB per rank: the 256K limit and 2.3× capacity of the 4 KiB profile. The text console keeps its framebuffer, and dgx3’s HDMI login prompt still works.
- **Decode page metadata (r5k).** This metadata is now shared across CUDA graphs.
- **64 KiB profile with Memory Saver, now the default.** The KV cache grows from 2.2 to 3.5 GiB per rank and the request limit from 262,144 to 524,288 tokens. CUDA graph capture also takes less memory: 0.98 GiB against 1.52 GiB. The container image, weights, model arithmetic and memory guards are unchanged. The 4 KiB profile stays in the repository as a fallback with the 256K limit.

_Long-context prefill_

- **Indexer split (r5l).** In sequence-parallel prefill, each rank scores and selects top-k positions for its own third of the rows of the sparse-attention indexer, then all-gathers them. Previously every rank scored every row. One 4,096-token chunk is 5% faster at 64K of context, 10% at 131K and 15% at 200K. Real-text prefill improves 3.5% at 64K, 6% at 131K and 9.7% at 200K.

_Quality and correctness_

- **Decode/prefill consistency (r5k).** The ratio-2 compressor ring is now written during decode, so compressed entries for generated tokens match what prefill builds. The FP4 KV writer also rounds like DeepSeek’s reference quantizer (B12X #435). Together these cut the worst decode-versus-prefill disagreement from 8.2 to 2.5 nats.

- **Exact indexer selections (r5m).** Indexer top-k now breaks exact score ties by lowest position, as DeepSeek’s reference does. Before, two identical runs chose different positions for 78% of one layer’s prefill rows.

- **Fenced pipeline stages (r5n, r5o).** Several B12X kernels released a pipeline stage for its TMA refill while shared-memory reads from that stage could still be pending. When another kernel shared the SM, the refill could overwrite data before it was read. The effects were:

_Memory Saver (experimental; our default, opt-in for everyone else)_

- **What it fixes.** On a 64 KiB-page kernel, NVIDIA’s UVM driver backs each 256-byte GPU page table with a whole 64 KiB page. About 53,000 tables are live with the model loaded, which is 3.25 GiB per node against 0.20 GiB on 4 KiB pages. That is why switching to the 64 KiB kernel had left less memory free, not more.
- **What the patch does.** It packs 16 tables into each page, bringing the backing down to 0.22 GiB. Net usable memory is 1.78–1.86 GiB per node higher than stock 4 KiB Linux.
- **Measured impact.** In a matched screen, free memory under load rose from 6.69/7.77/7.77 to 8.55/9.56/9.55 GiB. Decode, prefill and first-token latency were level with stock 4 KiB.
- **Status.** It is source-only and must be built and signed locally (Secure Boot stays on). It is pinned to driver 580.178.04 and kernel 7.0.0-1019-nvidia-64k, and it is an independent project, unaffiliated with NVIDIA.
- **Production default.** Our production profile now runs it (above). NVIDIA ships DGX Spark on 4 KiB kernels, so the 64 KiB profile and the UVM patch are outside NVIDIA’s validated configuration.
- **Repository.** [GitHub - christopherowen/dgx-spark-memory-saver: Recover usable memory on DGX Spark with a 64 KiB kernel: an opt-in NVIDIA UVM page-table packing patch, build tools, and measured validation. · GitHub](https://github.com/christopherowen/dgx-spark-memory-saver)

**Testing**

- **Every promotion:**
  - the same content-addressed image on all three ranks;
  - quality gate 5/5;
  - a needle-in-a-haystack check, 3/3 at 153K–178K tokens on the 4 KiB promotions;
  - a clean live configuration check.

- **Correctness fixes:** each was reproduced in a co-resident stress harness before it was fixed, then confirmed at zero across 12,000 calls per shape and in the compiled SASS.
- **Indexer split:** validated with a check mode in which every rank also recomputes the full indexer and compares its rows.
- **Memory Saver:**
  - 3,584 direct CUDA allocations with hole reuse and full-buffer readback, on both the stock and patched drivers;
  - a 4.5 GiB pinned transfer, BF16 matmul and CUDA graph replay;
  - a three-rank model load with all five quality cases;
  - manual and DKMS install, removal and reboot cycles with Secure Boot enabled.
  - Long production soaks are still to come.

- **64 KiB production profile:**
  - four simultaneous 485K-token contexts, each producing a 4,096-token reply;
  - zero preemptions, with all four fully prefilled contexts coexisting for 58.85 s at 52.5% peak KV use;
  - lowest free memory 5.83 / 7.59 / 7.60 GiB, against the 3 GiB guard;
  - near-limit retrieval, 3/3 at 519,142 tokens.

**Next**

- **Deterministic output.** The goal is bit-identical temperature-0 output whatever else shares the batch, and across restarts, at production speed.
  - An experimental batch-invariant profile now passes our full trace validation: zero differing rows on all three ranks across 270 scenario pairs, 156 of them across a restart. The scenarios cover mixed traffic, chunked prefill and prefix-cache hits.
  - In a short screen it stays within 4% of the 4 KiB production profile on every workload.
  - Most of the recovery came from a new mHC kernel that sums every partial in one fixed order at every batch size.
  - Next is the full measurement matrix, then a decision on offering it as an opt-in profile.

- **Decode speed.** Routed MoE is about 60% of an eight-stream decode step’s GPU compute, so it is the main remaining lever.
- **Memory Saver.** Next steps are longer soaks of the 64 KiB default, determinism testing with it loaded, and a report to NVIDIA with the allocation trace.
- **Upstreaming.** The B12X fixes, the vLLM patches and the PyTorch `fill_defaults` change.

**On comparisons**

The caveat from Monday’s post still applies. These figures use real prompts with reasoning on, count rejected drafts, and report means rather than the best run, so they will read lower than numbers measured other ways on the same hardware.

Thanks again to Local Inference Lab for the vLLM and B12X branches this builds on. Questions, reproductions and comparison runs are welcome.

> **[GitHub - christopherowen/spark3-vllm-ds41f: Reproducible Docker/vLLM deployment, tuning, and...](https://github.com/christopherowen/spark3-vllm-ds41f)**
>
> Reproducible Docker/vLLM deployment, tuning, and benchmarks for DeepSeek V4.1 Flash on a switchless three-node DGX Spark fabric.

---

<div class="post-metadata">

**Author:** ![ptitviet](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@ptitviet](https://forums.developer.nvidia.com/u/ptitviet)\
**Post date:** [October 2, 2026, 1:20pm UTC](https://forums.developer.nvidia.com/t/new-deepseek-4-1-flash-recipe/384438/5 "2026-10-02T13:20:06Z")

</div>

Sounds nice! Thanks for the job! That makes me think about buying a third ;)

---

<div class="post-metadata">

**Author:** ![christopher\_owen](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/christopher_owen/32/469895_2.png) [@christopher\_owen](https://forums.developer.nvidia.com/u/christopher_owen)\
**Post date:** [October 5, 2026, 4:24pm UTC](https://forums.developer.nvidia.com/t/new-deepseek-4-1-flash-recipe/384438/6 "2026-10-05T16:24:26Z")

</div>

This DeepSeek V4.1 Flash deployment on DGX Spark has a new default. With r6, the model’s compute kernels are **TileLang** kernels written for SM121, with DeepSeek’s own **TileKernels** underneath and one-shot RoCE collectives from my **sparknet** library. They replace B12X’s kernels. On the same hardware, recipe and benchmark, r6 is never behind B12X, and it is ahead on single-stream step time, at high concurrency and in prefill up to 1M tokens.

Repository: [GitHub - christopherowen/spark-ds41f: Reproducible Docker/vLLM deployment, tuning, and benchmarks for DeepSeek V4.1 Flash on a switchless three-node DGX Spark fabric. · GitHub](https://github.com/christopherowen/spark-ds41f) (the project formerly known as spark3-vllm-ds41f)

**Four Sparks (TP4, 1M context, 16 sequences)**

Aggregate tok/s, mean ± 95% interval. **Bold** marks a point whose interval does not overlap the other family’s.

| Streams | 1 | 2 | 4 | 8 | 16 |
| --- | --- | --- | --- | --- | --- |
| Prose, B12X r5p | 62.1 ± 3.8% | 92.7 ± 8.0% | 141.6 ± 7.9% | 215.7 ± 7.5% | 297.3 ± 1.2% |
| Prose, TileLang r6 | 62.4 ± 1.2% | 100.2 ± 2.7% | 147.1 ± 10.2% | 222.1 ± 0.3% | **318.4 ± 1.6%** |
| Code, B12X r5p | 72.5 ± 7.5% | 111.3 ± 5.4% | 168.3 ± 9.6% | 245.7 ± 1.3% | 323.5 ± 1.5% |
| Code, TileLang r6 | 76.0 ± 0.6% | 116.8 ± 18.9% | 174.9 ± 0.6% | **264.3 ± 0.6%** | **349.8 ± 2.5%** |

| Prefill, real source text (tok/s) | 32K | 256K | 500K | 1M |
| --- | --- | --- | --- | --- |
| B12X r5p | 5,117 ± 1.7% | 4,873 ± 2.5% | 4,645 ± 1.7% | 4,026 ± 0.9% |
| TileLang r6 | **5,836 ± 4.2%** | **5,471 ± 3.1%** | **5,072 ± 1.3%** | **4,370 ± 0.6%** |

Single-stream steps are **31.79 ms** (prose) and **34.75 ms** (code), against B12X’s 34.00 and 37.72: 6.5% and 7.9% shorter.

**Three Sparks (TP3, 8 sequences at 512K)**

| Streams | 1 | 2 | 4 | 8 |
| --- | --- | --- | --- | --- |
| Prose, B12X r5p | 50.4 ± 12.2% | 75.9 ± 10.6% | 118.8 ± 12.9% | 173.3 ± 8.4% |
| Prose, TileLang r6 | 52.2 ± 0.7% | 84.4 ± 0.6% | 125.7 ± 0.5% | **198.6 ± 0.4%** |
| Code, B12X r5p | 60.1 ± 2.0% | 90.8 ± 5.0% | 136.3 ± 6.3% | 192.3 ± 3.2% |
| Code, TileLang r6 | **63.0 ± 0.5%** | 97.3 ± 8.5% | 159.2 ± 44.2% | **240.3 ± 1.0%** |

| Prefill, real source text (tok/s) | 32K | 256K | 500K |
| --- | --- | --- | --- |
| B12X r5p | 3,778 ± 14.9% | 3,632 ± 2.2% | 3,432 ± 0.6% |
| TileLang r6 | 4,306 ± 20.9% | **4,034 ± 0.3%** | **3,743 ± 0.8%** |

Single-stream steps are **38.21 ms** and **42.43 ms** , against 41.92 and 46.44: 8.9% and 8.6% shorter. Eight streams decode 15% more prose and 25% more code.

Both families run the same native FP8/FP4/BF16 weights, the same DSpark speculative decoding (five drafts), the same memory budget and the same benchmark: quality gate, prose and code with reasoning on, temperature 0, 256 output tokens, three samples per point, real source text for prefill with two repeats. Each family uses DSpark cost curves profiled for its own kernels. The r6 numbers were measured on 2026-10-05, TP4 on the four-node ring and TP3 after recabling the triangle the same afternoon. B12X’s are r5p’s acceptance benchmark from the same day, with its 500K point from a run that morning.

**Reading the intervals**

B12X’s intervals are wider because its output text changes from sample to sample, and DSpark acceptance moves with the text. TileLang produces the same text every sample. Its two wide points, code at two and four streams, come from streams interleaving differently between samples.

An earlier TP4 run this morning showed TileLang behind at two-stream code (105.0 against 111.3). That point is one fixed text, and acceptance on that one text was lower. Across 13 texts at TP3, acceptance was level (2.401 against 2.384 accepted drafts per step), with about ±5% spread per text. In r6 the same point reads 116.8 ± 18.9%, level within its interval. The intervals cover sample-to-sample spread on one text, not text-to-text acceptance. Single-stream step time doesn’t depend on acceptance, so it is the cleanest kernel comparison.

**Why TileLang**

TileLang (tile-ai/tilelang) is a Python DSL for writing GPU kernels at the tile level. You state the tiles, the shared-memory staging, the software pipeline and the tensor-core MMAs explicitly. TileLang lowers that through TVM to CUDA, JIT-compiles it and caches the result. You keep control of layout and data movement without a C++ template stack, and a kernel is a few hundred readable lines.

What made it worth trying on the Spark is that DeepSeek writes production kernels in it. Their **TileKernels** library describes itself as “dozens of highly optimized kernels” used in their internal training and inference. The v2.0.0 release ([deepseek-ai/TileKernels#34](https://github.com/deepseek-ai/TileKernels/issues/34), 2026-09-30, 229 files) covers:

- MoE routing;
- per-token, per-block and per-channel FP8/FP4 quantization with fused SwiGLU;
- Engram gating;
- manifold hyper-connections (mHC) with Sinkhorn;
- RoPE.

TileKernels lists SM90 and SM100 as its targets. Everything I use from it runs unmodified on GB10 (SM121).

Two properties come with it:

- **DeepSeek’s arithmetic.** TileKernels follows DeepSeek’s reference term for term. Where B12X and DeepSeek differ, I prefer to move toward the reference.
- **Determinism.** Every kernel adds its reductions in a fixed order, so a row’s result doesn’t depend on the batch it ran in, decode and prefill alike.

**How it works**

1. **The compiler.** TileLang 0.1.15, built from source with three patches for the SM120 family:

2. **The DS4.1 kernels** , as a vLLM kernel backend:

3. **TileKernels** supplies the MoE top-k gate, the SwiGLU and quantization casts, and mHC’s post and stream collapse. DeepSeek moved the mHC projection itself to DeepGEMM, which doesn’t run on SM12x. That projection is a TileLang kernel of mine, with the V4.1 lagged pre-mix and the stream update folded in registers.

4. **The collectives** come from sparknet (below).

5. **The switch.** The configuration’s `kernel_backend` names the family for the whole model, drafter included, and `bin/spark doctor` checks that the backend flags and environment match it. Everything else matches the B12X configuration field by field, so the comparison isolates the kernels.

**Kernel by kernel against B12X**

Rank-0 decode profile at TP4, one stream, median six-row verification step, both families on the same four Sparks. The whole step is 39.3 ms against 44.9 ms (GPU busy 30.0 against 36.8 ms).

| Kernel group | B12X | TileLang | Change |
| --- | --- | --- | --- |
| Routed experts | 21.05 ms | 16.50 ms | −22% |
| Shared expert (side stream, per call) | 52.7 + 46.7 µs | 10.0 + 19.1 µs | −71% |
| Collectives | 4.57 ms | 4.15 ms | −9% |
| mHC | 1.91 ms | 1.55 ms | −19% |
| Sparse MLA | 1.22 ms | 1.10 ms | −10% |
| Router | 0.59 ms | 0.45 ms | −24% |
| Quantization and norms | 0.44 ms | 0.32 ms | −28% |
| Fused Q-A/KV (per call) | 18.0 µs | 14.2 µs | −21% |
| DSpark main projection (per call) | 175.6 µs | 173.2 µs | −1.4% |
| Q-B (per call) | 20.8 µs | 21.9 µs | **+5%** |
| Indexer Q-B (per call) | 34.3 µs | 35.2 µs | **+3%** |

Q-B and the indexer’s Q-B still trail slightly in serving, although both win in isolation: warm in L2, cold from DRAM, and racing their own L2 prefetch. Three different Q-B tiles all measured 21.8–21.9 µs in serving, so the tile isn’t the limit. B12X stores weights tile-packed, so each tile is one contiguous read; TileLang reads row-strided weights. That layout is the next thing to try.

The shared Triton kernels (page mapping, chunk metadata, the attention output quantizer) read slightly slower in TileLang’s profile. They’re the same kernels: the L2 weight prefetch stream overlaps them more. With the prefetch off they run at or below B12X’s times, and the prefetch saves about 1.7 ms per step overall.

**A race in TileLang’s warp-specialized pipeline**

The new 16-row decode tiles failed a repeated bit check against the prefill kernel. Identical inputs gave different and wrong results, in up to every run for some tile shapes. The cause is TileLang’s warp-specialized pipeline on SM121, where producer warps stage the block scales with `cp.async` for consumer warps. TileLang’s unspecialized pipeline never failed in any configuration and is just as fast for decode. So decode tiles use it, while prefill keeps warp specialization, which is up to 17% faster there. Adding memory clobbers to TileLang’s mbarrier assembly did not fix it, so compiler hoisting isn’t the cause. The 64-row tiles in the earlier image showed no failure in 18,000 runs. I’ll report it upstream with the reproducer from the experiment.

**Numerics and determinism**

On TP3, the two families score the same 8,188 tokens equally:

| Measure | TileLang | B12X |
| --- | --- | --- |
| Mean NLL | 1.5525 | 1.5527 |
| Top-1 accuracy | 67.84% | 68.04% |
| Decode against its own prefill: logprob gap | 0.0469 | 0.0565 |
| Decode against its own prefill: argmax differs | 2.59% | 3.32% |

Quality gates passed 5/5 in every window.

B12X is still in the image, with every non-determinism and race fix from the last weeks in the tree:

- the FP4 KV writer rounds like DeepSeek’s reference quantizer;
- indexer top-k ties break by position, so selections repeat;
- dense GEMM and prefill kernels fence shared-memory stage reads before the TMA refill;
- W4A8 tiny decode stays off.

Switching back is a configuration change (r5p’s B12X configuration is kept in the repository). I’ll likely retire B12X in the next release, once the remaining kernels below are ported, so maybe grab the patches now for your recipe if you want them.

**Collectives moved to dgx-spark-networking**

The one-shot RoCE collectives (formerly RoCEnante inside B12X) now live in their own public repository, **dgx-spark-networking** (Python package `sparknet`): [GitHub - christopherowen/dgx-spark-networking: Switchless RoCE collectives, NCCL profiles and fabric tooling for DGX Spark inference recipes · GitHub](https://github.com/christopherowen/dgx-spark-networking)

It has the one-shot all-reduce and all-gather for direct-cabled Sparks, plus NCCL profiles for 2-, 3- and 4-node and switched fabrics. The TileLang family uses it at TP3 (direct triangle) and TP4 (ring), and nothing in it is specific to this model.

**Using it**

Both recipes are promoted configurations in `config/`:

- **Three Sparks (TP3):** `config/cluster.json` (64 KiB pages; `config/cluster-4k.json` for a 4 KiB kernel). `docs/replicate.md` walks through the hosts, the image build and the first start.
- **Four Sparks (TP4, 1M context):** `config/cluster-tp4.json`, selected with `--cluster-config config/cluster-tp4.json`.

Four Sparks, briefly:

1. Cable a ring, each node to the next with two ConnectX-7 paths, and give every cable its own pair of /24 subnets. Check each path with a 9000-byte ping and that RoCE GID index 3 holds the IPv4 address.
2. Write a four-node map from `config/examples/nodes-ring4.json` as `config/nodes-ring4.local.json`: ranks in ring order, management IPs, and each peer’s RoCE devices.
3. Edit `config/cluster-tp4.json` (the promoted TP4 profile) for your site: head address, home directory and interface names. Then run `bin/spark --cluster-config config/cluster-tp4.json doctor`.
4. Build once (`bin/spark build prepare && bin/spark build image --apply`), copy the image to the other nodes, download the model on each, then `bin/spark --cluster-config config/cluster-tp4.json cluster start --apply`.

The TP4 section of `docs/replicate.md` has the same steps.

**What is still missing**

- **Kernels still on B12X under the TileLang family:** the WO projections, a BF16 GEMV, attention rotary, the KV cache writers, the compressor and the checkpoint loader. Also outside the family for now: Engram (TileKernels has Engram kernels to evaluate), the DSpark context-KV projection and the vision tower.
- **Q-B and indexer Q-B in serving:** a tile-contiguous weight layout for the decode projections, kept bit-identical with prefill.
- **mHC:** fold the per-token finalize into the projection, for one launch per sublayer.
- **Tooling:** `bin/spark tuning` still speaks B12X’s transport settings; teach it sparknet’s.
- **Upstream:** the TileLang SM120 patches and the warp-specialization race report.

**On the experience**

Compared with the FlashInfer and CUTLASS development I did for gpt-oss-120b back in the day, this was a pleasure. TileLang kernels are short Python I can read and change in an afternoon, and the generated CUDA is there to inspect when an MMA, or a pipeline, does something unexpected. Most of the time went into the model, not the toolchain.

**Credits**

- the tile-ai team for TileLang;
- DeepSeek for TileKernels and for publishing kernels used in production;
- Local Inference Lab for the B12X and vLLM integration branches, which remain the reference this work was measured against.

The experiments, configurations and raw benchmark reports are in the repository:

- `experiments/2026-10-05-tilelang-r6` for r6 and both benchmarks;
- `experiments/2026-10-05-tilelang-decode-kernels` for the kernel-by-kernel work, the race census and the profiles;
- `2026-10-03-tilelang-kernels` through `2026-10-04-tilelang-1m` for the module work.

Setup: DGX Spark (GB10), driver 580.178.04, kernel 7.0.0-1019-nvidia-64k, DeepSeek V4.1 Flash native weights. Benchmarks are temperature 0, 256 output tokens, reasoning on, three samples per point, aggregate throughput, 95% intervals; prefill is real source text, two repeats.

Questions, and results from other Spark setups, are welcome.

---

<div class="post-metadata">

**Author:** ![paxren2020](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@paxren2020](https://forums.developer.nvidia.com/u/paxren2020)\
**Post date:** [October 5, 2026, 6:34pm UTC](https://forums.developer.nvidia.com/t/new-deepseek-4-1-flash-recipe/384438/7 "2026-10-05T18:34:18Z")

</div>

Thanks for keeping the TP3 configuration supported! As soon as my third node arrives, this recipe is going to be the very first thing I try out. Appreciate the detailed write-up!

---

<div class="post-metadata">

**Author:** ![eugr\_nv](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/eugr_nv/32/508945_2.png) [@eugr\_nv](https://forums.developer.nvidia.com/u/eugr_nv)\
**Post date:** [October 5, 2026, 9:21pm UTC](https://forums.developer.nvidia.com/t/new-deepseek-4-1-flash-recipe/384438/8 "2026-10-05T21:21:20Z")

</div>

BTW, DSV4F TP3 config is supported by spark-vllm-docker too if you happen to use it.

---

<div class="post-metadata">

**Author:** ![christopher\_owen](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/christopher_owen/32/469895_2.png) [@christopher\_owen](https://forums.developer.nvidia.com/u/christopher_owen)\
**Post date:** [October 5, 2026, 10:54pm UTC](https://forums.developer.nvidia.com/t/new-deepseek-4-1-flash-recipe/384438/9 "2026-10-05T22:54:37Z")

</div>

Thank you for contributing to my new deepseek 4.1 flash recipe thread :).

Did you give it a shot?

---

<div class="post-metadata">

**Author:** ![el8](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/el8/32/536468_2.png) [@el8](https://forums.developer.nvidia.com/u/el8)\
**Post date:** [October 6, 2026, 1:32am UTC](https://forums.developer.nvidia.com/t/new-deepseek-4-1-flash-recipe/384438/10 "2026-10-06T01:32:26Z")

</div>

Please clarify that your sparknet library’s one-shot collectives are a vendored derivative of b12x’s RoCEnante, with your switchless-routing extensions and subsequent backend work. RoCEnante hasn’t moved out of b12x. The repository documents this relationship, but the forum announcement currently reads as though the upstream project relocated.

---

<div class="post-metadata">

**Author:** ![lukea](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/lukea/32/34575_2.png) [@lukea](https://forums.developer.nvidia.com/u/lukea)\
**Post date:** [October 6, 2026, 2:06am UTC](https://forums.developer.nvidia.com/t/new-deepseek-4-1-flash-recipe/384438/11 "2026-10-06T02:06:03Z")

</div>

Personally I have no issue with forks (or I wouldn’t have picked the license I did), but renaming the project, declaring our pun “dead” and then engaging in multiple rounds of self-promotion on social media is in astonishingly bad taste.

RoCEnante (and the rest of b12x) will be moving to FlashInfer officially today/tomorrow, and we’ll begin the full upstream integration process shortly thereafter.

---

<div class="post-metadata">

**Author:** ![christopher\_owen](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/christopher_owen/32/469895_2.png) [@christopher\_owen](https://forums.developer.nvidia.com/u/christopher_owen)\
**Post date:** [October 6, 2026, 12:30pm UTC](https://forums.developer.nvidia.com/t/new-deepseek-4-1-flash-recipe/384438/12 "2026-10-06T12:30:03Z")

</div>

Luke, the jokes on Twitter were meant in good humour \*. Here, I prefer to keep the discussion professional. I share the code, patches, experiments and results so others can inspect, reproduce and build on the work. Hopefully it’s useful for somebody.

To clarify “moved out of B12X”: I meant the implementation used in my recipe. Sparknet’s one-shot collectives derive from RoCEnante, with my switchless-routing extensions and subsequent work. Attribution is present in the package itself: the [NOTICE](https://github.com/christopherowen/dgx-spark-networking/blob/main/NOTICE) explicitly credits Luke Alonso and the Local Inference Lab contributors, and the provenance documentation identifies the source revision and changes applied. I recognise that the announcement could have expressed that relationship more precisely.

The separation has a technical purpose, explained in the [dedicated networking thread](https://forums.developer.nvidia.com/t/dgx-spark-networking-open-source-switchless-roce-collectives-nccl-profiles-and-fabric-tooling-for-multi-spark-inference/385049): pursuing GPU-initiated networking independently of the model’s compute kernels. The planned direct gpu to network path moves preparation of network work requests onto the GPU, initially retaining a CPU proxy to ring the NIC doorbell. The current collective transport still uses the host proxy; the GPU path is ongoing development. Experiments this morning saw a 1.5% lift. All of the experimentation is visible in the repos, but I will make announcements if it is a net gain.

My intention is also to remove B12X from this recipe entirely as the remaining kernels are replaced with TileLang implementations, using DeepSeek’s TileKernels where applicable. That transition is unfinished, and the release post lists the remaining B12X components. The attribution for code derived from RoCEnante remains relevant throughout that work. My work supporting sm121a in TileLang is in PRs at that project.

The networking package includes several additional changes:

- **Switchless collectives:** per-peer NIC routing, four-node ring support, and bidirectional relaying to reach the opposite node. Dispatch limits are separated from buffer capacity so transport selection can be tuned independently.

- **NCCL correctness:** Stanislav Bardyuk’s AArch64 memory-ordering fence fixes a send-path race that can hang the proxy. His contribution is credited explicitly.

- **NCCL performance:** my subsequent patches balance channels across both ring directions and PCIe roots, expose the allocation floor, and reduce thread-block size before dropping channels for small calls. The selected ring profile uses four channels and a 4 MiB buffer, alongside size-based selection between one-shot collectives and NCCL.

On the tested four-node ring, the balanced NCCL policy reduced a 2 MiB all-reduce from 320 to 232 µs. The documentation also reports regressions for some smaller reduce-scatter cases and the thermal trade-off that led to choosing four channels.

The recipe’s other tuning includes smaller decode tiles, improved prefill weight access, sequence-parallel prefill, speculative-decoding changes and memory savings. Those changes and their measurements are published in this thread.

I appreciate the upstream foundation and welcome concrete technical feedback or attribution corrections. The planned FlashInfer integration sounds like a useful development for everyone working on this hardware.

Personally, I had difficulties working with the Flashinfer team in the past when I was doing the gpt-oss-120b work last year so I have been steering away from it - but maybe that is different now with gentlemen like yourselves involved.

If anyone is in Berlin for GTC, I’d be happy to discuss tech stuff over a beer (or X shitposting if you want to be mad with me ;)).

I respect you and appreciate you. I’m just having fun and building stuff for myself and sharing it. At least with this Deepseek 4.1 Flash recipe, I have been having success as measured by myself.

\* please see my recent video posted on X of Deepseek 4.1 flash internals in the theme of 1980’s BBC Hitchhikers Guide to the Galaxy for reference to my brand of humour. Alternatively watch the video posted from a week earlier in the theme of Portal the game / “How it is made” for another example. I will keep my humour on X and not here.
