1x Spark tuned DSpark for DeepSeek-V4-Flash: 35 tok/s, 800+ prefill, and fast multi-agent serving

When antirez/ds4 shipped in early May as an MLX-first engine for DeepSeek-V4-Flash, I forked it within days, before it had CUDA support, and wrote my own CUDA backend for it. Upstream added its own soon after, but it wasn’t going to chase Blackwell serving performance the way I wanted, so I kept running with consumer Blackwell as the premiere target. Ten weeks, 409 commits, and 90,000+ added lines later, the fork makes this model genuinely fast to serve on a single DGX Spark, and it’s public for the first time today. One command:

curl -sSL https://raw.githubusercontent.com/entrpi/ds4-on-spark/main/install.sh | bash -s -- --start

That installs, builds for the Spark’s GPU, downloads the model (~91 GiB), smoke-tests, and serves an OpenAI-compatible endpoint on :8000.

Versus the upstream engine (same box, same GGUF, speculation off in this comparison; it only widens the gap):

Highlights

It’s fast. Prefill runs ~2× the upstream engine, and speculative decode (DSpark: a small draft model checked by the full model, so output quality is untouched at any temperature) reaches 27–35 tok/s on structured work like math, Q&A, and code, vs ~20 tok/s plain. Across a 9-workload benchmark suite the mean is 1.38× plain.

It can’t make things slower. Speculation only pays off when the draft model guesses well, and on difficult prose it historically cost you speed. The fork watches each request and switches speculation off mid-request the moment it stops paying. The worst case I could produce (continuing War & Peace, a test chosen specifically to make the draft model fail) measured 0.96× plain, where always-on speculation drops to 0.72×.

Agents get the speed too. Requests that think and call tools (the shape every agent framework produces) ride the fast batched path with speculation, instead of falling back to a slow serial path the way tool-call grammar usually forces. tool-eval-bench in thinking mode runs 27 % faster wall-clock (50 → 37 minutes) at the same score, and a real agent framework (Hermes) ran end-to-end with every generation on the fast path and zero fallbacks.

Deep context on one box. A 518K-token orchestrator conversation and a 248K-token subagent served concurrently: 766K live tokens on one Spark. I did a full day of deep-context needle-in-a-haystack testing, and the model always found the needle: ten runs each at 250K and 500K context, with the needle buried at a different depth every time. The first deep prefill is real time (~13 min at 250K), but prefix caching makes every later turn on that conversation start in ~1–2 seconds.

It’s tested like a release. Every release candidate has to pass the same battery: tool-eval-bench (all 69 scenarios on fast + thinking paths, plus adversarial and error-injection modes, plus three full repeats to check consistency), the deep-context and needle testing above, an hour of churn (one huge pinned conversation while a dozen short-lived clients hammer the server; it never lost the deep conversation, and memory stayed flat), and a full re-run of the quality evals against my June baseline (GSM8K 96.8, HumanEval 90.9, MMLU 63.9, needle 70/70). All of it green on the exact build the installer ships. Two crash bugs this testing caught along the way are fixed in the release.

The speed is diligence, not one trick. It comes from closing every efficiency gap I could find across the whole stack, from the 2-bit math kernels up through scheduling and serving, with work landing three days out of every four and each change kept only if it beat the previous build on the same benchmark. As far as I can tell, no engine running a large mixture-of-experts model below 4-bit has pushed this class of GPU closer to its physical limits. There’s more to come, but the majority of the available headroom is already captured.

The point is bigger than this one model. The models that fit consumer Blackwell at NVFP4 are already well served; the most capable open MoE models don’t fit at 4-bit on a 128 GB box (this one is 81 GiB at 2-bit; the same weights at FP4 would be roughly 140 GiB). This fork is meant as an exemplar of what’s possible in the class above that line: served fast, batched, deep-context, with quality holding, on one box.

caveats

  • Speculative gains depend on content: code/math/Q&A win big, creative prose sits at parity.
  • Decode slows with very deep context (~146 ms/tok at 248K, ~177 at 519K); the deep-context win is capacity and instant warm turns, not raw speed.
  • This is the 2-bit-experts quant; quality holds up remarkably well (numbers above), but it’s not the unquantized model.

Repos: ds4-on-spark (installer + full benchmarks) · Entrpi/ds4 (the fork; full change list) · drafter GGUF

Credits: antirez/ds4 is the foundation and reference engine. DSpark follows the same family of ideas as Modal’s DFlash. Model by deepseek-ai, 2-bit recipe by antirez.

More of a deep-dive…

Decode, per workload

DSpark decode on the llama-benchy 9-workload suite, wall-clock non-streaming tok/s (GB10, 128 GB):

workload plain t/s DSpark t/s gain accept
stepwise_math 20.1 34.5 1.71× 89 %
qa_factual 19.6 31.9 1.63× 91 %
summarize 19.6 31.0 1.59× 87 %
code_cpp 20.4 30.0 1.47× 80 %
translation 20.4 27.9 1.37× 76 %
code_python 20.5 26.4 1.29× 72 %
explain_concept 20.5 25.2 1.23× 69 %
long_code_review 18.8 22.1 1.17× 66 %
creative_short 20.6 19.9 0.96× 58 %
suite mean 20.1 27.7 1.38× n/a

The break-even law, and why speculation can’t hurt you here

The table spans 58–91 % draft acceptance, and that spread is the whole story. One 5-row verify forward costs ≈ 2.17 plain decode steps (measured, and depth-flat), so below ~56 % acceptance a verify step costs more than it pays; classic always-on speculation bottoms out at 0.72× plain on adversarial prose. That is what usually makes spec decode a footgun in production.

The fork ships two guards, both on by default:

  • a terminal yield-quench controller: every verify step accumulates the realized regret (2.22 − tokens_committed); a request that can’t pay its way gets speculation turned off for the rest of that request and runs at true plain speed. The threshold was calibrated by replaying per-step traces from real runs, and when I force-quench a request and measure it, it runs at 1.000× plain; it really is plain decode.
  • a kv-depth gate at 64k: deep-context requests hand off to plain batched decode, which stays 1.15–1.2× upstream out to 128k.

Net effect on the hardest corpus I have (the War & Peace continuation in the OP’s chart, ~52 % accept, deliberately the floor test): worst case 0.96× plain instead of 0.72×, with all the structured-content wins intact. The dashed line in that same chart is C-source continuation: mean 1.11×, peaks 1.4× at 47k depth. Content decides the win, not depth.

On losslessness: the target model’s verify forward is the only token source (DFlash-style), so output quality is the model’s own. The accept loop samples the genuine token from the verify-row logits with the request’s own sampler and RNG, so this holds at any temperature, not just greedy.

How agents got onto the fast path

Tool-call grammar needs a mid-stream sampler swap, which is why engines typically route requests with tools plus thinking (or any temperature > 0) onto a slow serial path. The fork runs that sampler override inside the continuous-batching path itself, so agent traffic (thinking, tools, any temp) rides continuous batching + DSpark.

The numbers behind the OP’s claim: running all 69 tool-eval-bench scenarios in thinking mode went from all-serial to all-batched at the same score, and wall clock dropped from 50 to 37 minutes (−27 %). Driving the server with Hermes end-to-end: every generation rode the speculative path at 80–97 % acceptance, committing 3.4–4.5 tokens per verify step, and nothing fell back to the serial path.

One more knob for agent fleets: DS4_SERVER_DEFAULT_TEMP sets the temperature for requests that omit one (most frameworks do); explicit temperatures are untouched. Lets you run a fleet greedy-by-default without patching the framework.

Deep context, precisely

The OP’s 766K claim, unpacked. The admission/placement stack (VMM demand-mapped KV, memory-budgeted banks, and a prefix cache that pins deep conversations against eviction) serves:

  • 766,437 tokens live at once: a 518K-token orchestrator and a 248K-token subagent, both resident in KV and decoding concurrently at ctx 524288;
  • needle-in-a-haystack, 20 for 20: ten runs each at 248K and 519K actual tokens, needle buried at a different depth every time (5 %–95 %), always found exactly. Cold TTFT was dead-flat across depth (843 s ± 0.4 at 248K, 2434 s ± 2 at 519K) and memory stayed flat across the 9-hour matrix;
  • the next turn on a 518K-token conversation starts in 1.2 s: the prefix cache serves it warm, in place. Building that conversation the first time is ~41 min of real prefill, paid once;
  • deep decode, honestly: ~48 ms/tok shallow → 146 ms/tok @248K → 177 ms/tok @519K. Decode is not context-flat at these depths; that’s the real curve.

The F32-KV ceiling is ~766K concurrent tokens per box with the full model resident. FP8/FP4 compressed-KV storage (bit-lossless, already shipped as opt-in flags) is the road past 1M.

Concurrency: what a fleet of agents gets

Same ship-default server, 1→16 concurrent requests (pp=2048, tg=256 each):

Aggregate decode rises 18 → 30 tok/s and saturates around 4 concurrent requests. The regimes are deliberate: a solo stream gets DSpark speculation (the single-stream point in this chart uses the adversarial prose corpus, so it shows parity; on structured content a solo stream reaches 27–35 tok/s, per the table above); concurrent traffic hands off to plain continuous batching, which wins above N=1 on this hardware. In practice a single agent decodes at DSpark speeds comparable to the whole box’s batched aggregate, and a fleet gets ~30 tok/s of shared decode plus warm-start/fork prefix caching for TTFT.

What’s in the stack

  • D2R prefill kernels: IQ2_XXS / Q2_K / Q8_0 expert weights dequantized directly into tensor-core fragments from a weight server’s aligned SoA artifacts, plus token-tile HMMA attention and an L2-aware expert-major CTA schedule. The kernel work took prefill from 305 to 800 tok/s @12k on GB10; ~2× upstream (~4× on an RTX PRO 6000, sm_120).
  • CUDA-graph decode: per-layer graph capture of the batched decode step.
  • DSpark: 3-layer block drafter fused to the target’s hidden states (Q2K, ~6.5 GiB, prebuilt on HF), verified by the target every step. Plus the quench controller and kv gate above.
  • Serving: continuous batching (mid-flight admit/evict, chunked prefill interleaved with live decode), prefix caching (warm start ≈7× faster TTFT on shared prefixes, fork-by-copy fan-out for parallel branches, partial-prefix reuse when a request diverges mid-prompt, and a pinned tier so a deep orchestrator conversation is never evicted by short-lived requests), OpenAI- and Anthropic-shape APIs.
  • Weight server: one resident process owns the 81 GiB model (VMM-backed); engine restarts and A/B runs import it in seconds instead of minutes.

The roofline receipts

The OP claims this pushes sm_120-class Blackwell about as close to its limits as a below-4-bit big-MoE engine has gotten. The receipts: my own microbenchmarks put GB10’s practical ceilings around 250 INT8 TFLOPS and 125 F16 TFLOPS, and the prefill GEMMs now operate in that regime. An isolated-kernel roofline study concluded that the remaining single-stream decode gap is the memory substrate itself, not software. What’s left on the table: FP8/FP4 KV decode parity, and a projected 15-30 % of prefill.

The release gates, in full

I delayed this announcement once already: my own tool-calling tests crashed the exact configuration I was about to ship, twice. Both crashes are fixed (root-caused to an out-of-bounds eager decode in a compressed-KV mirror kernel on large admission chunks: a deterministic placement bug, not a race), and every release candidate now has to pass the same battery before it ships:

  • tool-eval-bench, all 69 scenarios on both the fast and thinking paths, scoring 81–86/100 across repeated runs, plus opt-in Hard Mode (15 adversarial/stateful scenarios: 73/100, both times I ran it), error-injection (every tool call has a 20 % chance of returning a failure: 82/100, both times), and repeatability (three full-suite runs: mean 84.7 ± 2.3; a scenario passes at least one of the three runs 81 % of the time, and passes all three 71 % of the time);
  • the deep-context capacity test (the 766K/518K numbers above are its output);
  • an hour of churn: one pinned ~240K-token conversation held live while 12 short-lived clients cycled through the server continuously, 41 rounds in all. Every few minutes I asked the deep conversation to recall a secret code planted in its context. It answered correctly every time, from cache, in 1.8 s; it was never evicted; server memory drifted 1.0 GiB over the hour;
  • a full re-run of the quality evals against my June baseline: GSM8K 96.8 (June 97.6), MMLU 63.9 (63.5), HumanEval 90.9 (87.8), MBPP 89.0 (90.0), IFEval strict 81.7 (83.4; that drift is run-to-run noise, identical in a same-day control build: the release engine itself moved exactly 1 item of 541), and needle-in-haystack 70 of 70 exact across the inline/64K/130K tiers.

Fine print the OP didn’t have room for

  • DSpark engages on solo streams; concurrent traffic runs plain continuous batching (which wins above N=1 on this hardware). Deep-context requests (>64k) decode plain by design.
  • Loading a deep conversation the first time is real prefill time (~13 min at 250K, ~41 min at 518K). The prefix cache makes every subsequent turn start in seconds.
  • Everything here is measured on GB10 (sm_121); RTX PRO 6000 / 5090-class sm_120 builds and was measured for prefill, sm_100 datacenter untested.
  • Re-running the installer on an existing install upgrades in place (see the upgrade notes); --no-dspark serves plain continuous decode instead.

Full benchmarks, the roofline model, the break-even math, and every fork-side change: ds4-on-spark README and the fork CHANGELOG. Benchmarks via eugr/llama-benchy.

Where this is heading next. The engine is in good shape, but the roadmap is not empty; here are the top priorities, in order.

1. Deep-context decode performance. I think there’s still a lot of room to optimize deep-context performance, so that’s one of the top priorities going forward. Decode today runs ~48 ms/tok shallow but 146 ms/tok at 248K and 177 ms/tok at 519K, and unlike prefill, that deep band has never had the full profiling-and-kernel treatment; it’s the least-scrutinized hot path left in the engine. I don’t expect it to go flat, but I expect the curve to come down.

2. Past 1M tokens on one box. The v0.2 ceiling is ~766K concurrent tokens because the KV cache is stored at F32 width. FP8/FP4 compressed-KV storage already ships as opt-in flags, it’s bit-lossless, and it cuts KV memory ~2.4×, which puts >1M concurrent tokens on a single Spark within reach. What holds it back from being the default is speed, not correctness: the current implementation pays an F32 round-trip and a scalar unpack on the decode path, which costs about a third of deep decode throughput. That’s implementation overhead, not physics, and closing it is a kernel job of exactly the kind the prefill arc already went through. At parity it becomes the default and the ceiling moves past 1M.

3. The last of the prefill headroom. Prefill went from 305 to 800 tok/s at 12k over the fork’s life, and the roofline math says roughly 15-30 % is still on the table. The remaining levers are identified; they’re queued behind the two items above.

4. Graceful behavior at the memory ceiling. Today, a box holding several idle cached deep conversations will refuse a brand-new deep request rather than evict one of them; the refusal is clean (an immediate 503, never a silent slow path), but eviction would be better. The plan is budget-aware eviction: idle cached conversations yield their memory to new demand, oldest first, pinned ones last. Same theme, smaller: the couple of remaining failure modes that can still take the whole server down under extreme memory pressure should fail only the offending request instead.

None of this changes the defaults you get today; the OP’s install command and numbers stand. If you’re running workloads where any of these bite (or you want a benchmark run against something specific), reply here; the gate suite grows by exactly this kind of request.

We’re all expecting V4-Flash official release from DeepSeek any day now, so I’ll jump on that to get it into this system ASAP when the weights drop too.

This is crazy! ! ! Prefill speed 800! ! ! Isn’t he the king of a single dgx spark! ! Don’t doubt! ! He definitely is! ! !

thank you, just trying it and so far I am impressed :) Just one thing and probably stupid question when I list models /v1/models I see 2 ids:

deepseek-v4-flash
deepseek-v4-pro

Is there any difference to which I am pointing?

thanks

Hi! Could you clarify what you mean by “V4-Flash official release”?

Since DeepSeek-V4-Flash weights are already public on Hugging Face, are you referring to a new checkpoint, a stable (non-preview) release, or something else?

That’s the preview release. There’s no final release yet.

Thanks for trying it, and glad it’s holding up. Not a stupid question at all; that one’s on the server, not you.

The model field never switches weights. Your box has exactly one model loaded (DeepSeek-V4-Flash, the 81 GiB quant the installer fetched), and every request is served by it no matter which id you send. The second id shows up because the endpoint hardcodes its list: the engine also supports V4 Pro weights, and /v1/models advertised both regardless of what’s actually loaded. That’s misleading, so I’ve pushed a fix that lists only the loaded model; it’ll be in the next release.

The ids that do change behavior on your install:

  • deepseek-chat turns thinking off (the fast path; what you want for tools and agents)
  • deepseek-reasoner or the default deepseek-v4-flash serve with thinking on
  • or set think: false (or thinking: {type: disabled}) explicitly and use whichever id you like

So pointing at deepseek-v4-flash is already right; switch to deepseek-chat when you want non-thinking speed.

I am super impressed!!! Thank you so much for your efforts to make it feasible to use on a single Spark!

I am particularly impressed by your needle-in-a-haystack test results. Do you think the good results will be affected if we eventually used FP8 for the KV cache?

Again, very impressive work! Thank you so much!

@entrpi Hi! What is the recommended way to monitor real generation throughput for all live ds4-server requests?

I can follow the server log with:

tail -f /home/dgx/ds4-server.log

but on the DSpark continuous-batch path I only see lines like CONT_MTP_ACCEPT with drafter acceptance and tokens-per-step. I do not see actual generation tok/s per request or aggregate tok/s across concurrent requests.

Is there an existing flag, endpoint, or log setting to emit:

  • per-request decode tok/s;
  • aggregate batch tok/s;
  • prefill tok/s;

for normal live server traffic?

Thank you! Short answer: no, FP8 KV won’t affect retrieval quality here.

The longer answer. In this engine the FP8 KV cache is not the usual lossy round-trip of full-precision activations. DeepSeek-V4’s compressed KV values are already the model’s own low-precision codes; storing the cache at F32 was just storing those same values wide. The FP8/FP4 storage keeps the codes directly, and I verified the A/B: the FP8-stored cache returns byte-identical values to the F32-stored cache. Same bits in, same bits out, so there is nothing for retrieval to degrade.

Measured on top of that, on the FP8/FP4 build specifically: deep needle probes still exact at 248K and 519K tokens, the tool-calling suite in band, and the engine’s speculative self-check passing bit-exact at temp 0.

It already ships in v0.2 as opt-in flags if you want to try it (DS4_CUDA_FP8_KV=1 DS4_CUDA_FP4_INDEX=1 in the server env). It cuts KV memory about 2.4×, which is the road past 1M concurrent tokens on one Spark. The only reason it isn’t the default yet is speed, not quality: the current decode path pays an F32 round-trip and a scalar unpack that costs about a third of deep decode throughput. That’s implementation overhead rather than physics, and closing it is priority #2 on the roadmap a few posts up; at parity it becomes the default.

Good ask. Two of the three exist today; the third is a real gap and I’ll close it.

Per-request decode tok/s. The most accurate place to measure is the response itself: every response carries an OpenAI-style usage block with prompt_tokens and completion_tokens, plus cached_tokens so you can see when the prefix cache served part of your prompt. For streaming, pass stream_options: {“include_usage”: true} and you get the same block on the final event. completion_tokens divided by the wall time from first token is the per-request decode rate; that is exactly how every number I’ve posted in this thread is measured (non-streaming, usage-counted wall tok/s).

Prefill tok/s. Boot the server with DS4_CONT_GEN_TIMING=1 in its environment. Every admission then logs a line like:

ds4: cont admit bank=2 pos_base=0 suffix=130255 wall=226.93s
suffix is the number of tokens actually prefilled and wall is the time it took, so suffix / wall is your prefill rate for that request. A nonzero pos_base tells you how many tokens the prefix cache spared you. The same flag also logs a summary line per continuous-batch cycle.

Aggregate batch tok/s. You’re right: this one doesn’t exist for live traffic yet. The CONT_MTP_ACCEPT lines you found only cover drafter acceptance, and I’ve been deriving aggregate rates client-side (summing concurrent per-request rates). Fair request, and it prompted a proper design rather than a patch. Next release gets a real observability surface, all rendered from one internal metrics core:

a timings block in every response next to usage (TTFT, prefill tok/s with the cached split, decode tok/s, speculative acceptance), so per-request numbers need no log parsing at all;
GET /metrics in Prometheus text format, so a fleet can point Grafana at it;
GET /v1/stats as human-readable status text, so watch curl works as a poor man’s top.
The log lines above will keep working, but they’re debug plumbing; the endpoints will be the interface.

I tried to get one of you earlier attempts downloaded and running and failed. This went a lot better.

I went the manual route, and it installed and reported a successful smoke test.

  • git clone
  • ./install.sh
  • ~/code/ds4/ds4-server --host 0.0.0.0 --cuda -m ~/gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf -c 32768

I didn’t think the server was working at first. It loads quickly, and sparkview reports its only using 29GiB of memory – but the cache is 89GiB – but its all there, responding to queries. I plan to have a good go with it tomorrow.

How large a context can I realistically run with this?


BTW: congrats – super impressive!

Nice release — the per-commit prefill ledger (305→800) and the yield-quench design are the most rigorous writeup I’ve seen on this box. We run a sibling antirez/ds4 CUDA fork (DS4S) on one GB10 serving the same model, so a few calibration questions to line our numbers up with yours: (1) the 800 tok/s @12k is labeled engine-side — what does llama-benchy report through :8000 with the default 4096-token chunked/interleaved prefill at the same depth? Your June wsclient CSV showed ~400 there, and served-vs-engine is the number most readers will actually get. (2) For decode, our plain single-stream measures 14.7–16.5 t/s on a ~91 GB mixed-quant build and ~19.7 ms/row of that is pure weight-bandwidth, so your 18–20 on the 81 GiB 2-bit build plus graph capture looks roofline-consistent — but the headline 35 is the stepwise-math speculative peak at 89% acceptance; what does the suite mean (we read 27.7) serve at on the adversarial-prose leg at, say, 32k context? (3) At c≥2 with DSpark gated off, do you have per-request p50/p95 rather than aggregate? For reference our DSpark-equivalent (draft=5, α=77.2%) serves 24.5–26.4 t/s single-stream. Happy to trade traces — the D2R expert-major schedule is something we’re looking at for our MXFP4 build.

I compared two locally deployed models on the full HumanEval+ benchmark:

Model HumanEval HumanEval+ Total generation time
DeepSeek V4 Flash 98/164 — 59.8% 97/164 — 59.1% 1 hour 33 minutes 52 seconds
Qwen 122B 153/164 — 93.3% 147/164 — 89.6% 14 minutes 22 seconds

Both models were tested under the same conditions: one greedy completion per task, temperature 0.

Of course, a single HumanEval+ run does not prove much on its own. In regular conversations and agentic tasks, the quantized DeepSeek model seems to work reasonably well and calls tools correctly. However, working at around 30 tokens per second in Hermes feels painfully slow. The same applies to coding in VS Code — there is simply too much waiting involved. I guess fast models have already spoiled us :)

For now, Qwen 122B remains the best option for me, both for agentic workflows in Hermes and for coding in VS Code.

I launched Qwen 122B using the dFlash setup described in this NVIDIA forum thread:

Yeah I made that thread too. I have used Qwen 122B more than anything else for single Spark as well. It’s the only good model that runs at comparable speed to normal API models.

For this model, I’m mostly anticipating v4 final and a likely later v4.2 will be big upgrades.

Based on all the tests you have done, what model would you recommend for someone like me with one spark? I used the qwen3.5-122B and was on the qwen3.6-27B, but I started to get tool errors using OpenCode and OpenHands. I am thinking of going with this model you tuned- well done by the way but want to hear from you. My use case is to run long agentic tasks with developer tools such as OpenCode and OpenHands. My OpenHands setup has not been able to run smoothly for more than two days, and I am still learning the ropes, so please bear with me.

Great questions. Thoughts, in order:

1. Prefill: your ~400 from June and my 800 are probably both right. The prefill arc landed in stages, and the jump to 800 tok/s at 12k came from the D2R tensor-core GEMM work plus the CTA-schedule and attention retiling that followed. Measured through :8000 with llama-benchy today, I see ~795 tok/s at 2k prompts decaying to ~550 at 128k (that’s the four-panel chart in the OP). If your June measurement was against my fork at the time, 400 was accurate then; if it was against upstream, that’s roughly where upstream still is.

2. One llama-benchy calibration warning before comparing decode numbers. Stock llama-benchy over-counts generation tok/s by ~1.9× for servers that do not return per-delta token_ids: the fallback re-tokenized each ~5-character SSE fragment separately and lost cross-boundary merges (128 real tokens counted as 248). I hit this, fixed it in my llama-benchy fork (5a05ab4 for the token count, 2114071 for peak_gen_tps, which stayed ~2× inflated even after the first fix because per-fragment timestamps survived the correction), and got burned badly enough along the way that it manufactured a false regression in one of my own comparison charts and I had to redo the numbers. If you’re benchmarking with stock benchy against a server that doesn’t emit token_ids, I’d rerun with the fork, or self-check by comparing benchy’s counted tokens against usage.completion_tokens from the same response. Every number I’ve posted in this thread is non-streaming usage-counted wall tok/s for exactly this reason (my ds4 fork also serves per-delta token_ids, so patched benchy counts it exactly). Depending on your benchy and whether your fork emits token_ids, your own numbers may deserve a second look in either direction.

3. Decode: 35 is a sustained whole-run rate, but yes, it’s speculative and content-dependent. 34.5 tok/s is the wall-clock mean over full generations on the stepwise-math workload at 89 % draft acceptance, not an instantaneous burst. The full spread is published in the table in post 2: 0.96× to 1.71× across nine workloads, suite mean 27.7 vs 20.1 plain. And that table may actually undersell the favorable end: the highest acceptance I’ve measured anywhere is on agentic traffic, where tool-call structure and thinking traces are highly predictable for the drafter. End-to-end agent runs (Hermes driving the server) sustain 80-97 % acceptance at 3.4-4.5 tokens per verify step, and agents are one of the most common real workloads for a box like this. With speculation off entirely, this fork decodes 18.8-20.6 tok/s on those workloads at shallow context; engine-measured against upstream at tg=128 it’s 17.6 vs 13.9 at 2k. So your 14.7-16.5 on a 91 GB mixed quant reads as upstream-class decode. The gap to this fork is not one trick; it’s an accumulated decode arc: per-layer CUDA-graph capture of the whole batched step, aligned-quant dispatch tiers reading repacked artifacts in place, split-K vectorized decode matmuls, fused gate+up+swiglu MoE decode, fused router and compressor stages, GPU-side embed gather. A large chunk of the fork’s several hundred commits are performance work of exactly this kind, each kept only if it beat the previous build on the same box.

4. p50/p95 with DSpark off: good ask, and about to get easy. v0.2.1 (tagging as we speak) adds a timings block to every response: TTFT, prefill tok/s with the cached split, decode tok/s per request. Percentiles become a one-liner over any load you drive. I’ll include a DSpark-off p50/p95 table when I post the v0.2.1 numbers later today.

Welcome, and good timing: agentic traffic is exactly what this fork is tuned for, though I haven’t driven OpenCode or OpenHands against it myself yet (my end-to-end agent testing has been with Hermes). Both tools speak OpenAI-compatible APIs, so they should point at this server directly.

I’ve queued work to characterize OpenCode and OpenHands specifically against this server (request shapes, streaming and tool-call compatibility, retry behavior, where reliability needs hardening) and I’ll publish a per-tool config recipe plus any fixes that come out of it. If you try either one before that lands and hit anything odd, post the request shape or error here and I’ll chase it; that kind of report can help grow the tool-calling gate suite to improve reliability.

Coming from Qwen: the quality numbers for this 2-bit build are in the eval table upthread (GSM8K 96.8, HumanEval 90.9); for agentic coding it holds up well in my testing, but I don’t really use small models for coding.

As for a general recommendation, that’s hard because smaller models are so use case dependent, that you really are best just trying them against what you want to achieve. You can build your own personalized benchmark if you can reliably characterrize your work. Generally I think this and Qwen 122B are the most robust for the single Spark config though.