When antirez/ds4 shipped in early May as an MLX-first engine for DeepSeek-V4-Flash, I forked it within days, before it had CUDA support, and wrote my own CUDA backend for it. Upstream added its own soon after, but it wasn’t going to chase Blackwell serving performance the way I wanted, so I kept running with consumer Blackwell as the premiere target. Ten weeks, 409 commits, and 90,000+ added lines later, the fork makes this model genuinely fast to serve on a single DGX Spark, and it’s public for the first time today. One command:
curl -sSL https://raw.githubusercontent.com/entrpi/ds4-on-spark/main/install.sh | bash -s -- --start
That installs, builds for the Spark’s GPU, downloads the model (~91 GiB), smoke-tests, and serves an OpenAI-compatible endpoint on :8000.
Versus the upstream engine (same box, same GGUF, speculation off in this comparison; it only widens the gap):
Highlights
It’s fast. Prefill runs ~2× the upstream engine, and speculative decode (DSpark: a small draft model checked by the full model, so output quality is untouched at any temperature) reaches 27–35 tok/s on structured work like math, Q&A, and code, vs ~20 tok/s plain. Across a 9-workload benchmark suite the mean is 1.38× plain.
It can’t make things slower. Speculation only pays off when the draft model guesses well, and on difficult prose it historically cost you speed. The fork watches each request and switches speculation off mid-request the moment it stops paying. The worst case I could produce (continuing War & Peace, a test chosen specifically to make the draft model fail) measured 0.96× plain, where always-on speculation drops to 0.72×.
Agents get the speed too. Requests that think and call tools (the shape every agent framework produces) ride the fast batched path with speculation, instead of falling back to a slow serial path the way tool-call grammar usually forces. tool-eval-bench in thinking mode runs 27 % faster wall-clock (50 → 37 minutes) at the same score, and a real agent framework (Hermes) ran end-to-end with every generation on the fast path and zero fallbacks.
Deep context on one box. A 518K-token orchestrator conversation and a 248K-token subagent served concurrently: 766K live tokens on one Spark. I did a full day of deep-context needle-in-a-haystack testing, and the model always found the needle: ten runs each at 250K and 500K context, with the needle buried at a different depth every time. The first deep prefill is real time (~13 min at 250K), but prefix caching makes every later turn on that conversation start in ~1–2 seconds.
It’s tested like a release. Every release candidate has to pass the same battery: tool-eval-bench (all 69 scenarios on fast + thinking paths, plus adversarial and error-injection modes, plus three full repeats to check consistency), the deep-context and needle testing above, an hour of churn (one huge pinned conversation while a dozen short-lived clients hammer the server; it never lost the deep conversation, and memory stayed flat), and a full re-run of the quality evals against my June baseline (GSM8K 96.8, HumanEval 90.9, MMLU 63.9, needle 70/70). All of it green on the exact build the installer ships. Two crash bugs this testing caught along the way are fixed in the release.
The speed is diligence, not one trick. It comes from closing every efficiency gap I could find across the whole stack, from the 2-bit math kernels up through scheduling and serving, with work landing three days out of every four and each change kept only if it beat the previous build on the same benchmark. As far as I can tell, no engine running a large mixture-of-experts model below 4-bit has pushed this class of GPU closer to its physical limits. There’s more to come, but the majority of the available headroom is already captured.
The point is bigger than this one model. The models that fit consumer Blackwell at NVFP4 are already well served; the most capable open MoE models don’t fit at 4-bit on a 128 GB box (this one is 81 GiB at 2-bit; the same weights at FP4 would be roughly 140 GiB). This fork is meant as an exemplar of what’s possible in the class above that line: served fast, batched, deep-context, with quality holding, on one box.
caveats
- Speculative gains depend on content: code/math/Q&A win big, creative prose sits at parity.
- Decode slows with very deep context (~146 ms/tok at 248K, ~177 at 519K); the deep-context win is capacity and instant warm turns, not raw speed.
- This is the 2-bit-experts quant; quality holds up remarkably well (numbers above), but it’s not the unquantized model.
Repos: ds4-on-spark (installer + full benchmarks) · Entrpi/ds4 (the fork; full change list) · drafter GGUF
Credits: antirez/ds4 is the foundation and reference engine. DSpark follows the same family of ideas as Modal’s DFlash. Model by deepseek-ai, 2-bit recipe by antirez.



