C1 1058pp/s, 52 tg/s on 1x dgx spark on Deepseek V4 flash 0731 full 256 experts

Thanks for this!

Ran this recipe as check of the numbers in the OP.

Concurrency Aggregate tok/s Per-request median
1 57.3 57.3
8 62.7 10.1
64 114.5 12.8

Opus 4.8:
Config: MAX_NUM_SEQS=12, MAX_MODEL_LEN=262144, GPU_MEMORY_UTILIZATION=0.85, otherwise recipe defaults. For reference, the same harness on the same box measures the Entrpi/ds4 v0.5.5 stack (IQ2XXS + DSpark drafter) at 21.5 single / 62 @c8 — so this is ~2.7× single-stream at comparable aggregate, with headroom at c16 the ds4 stack doesn’t reach.

Tool calling: 3/3 on our agent-style test (multi-step call with correct args, chained follow-up after a tool result, and restraint on a no-tool question) — the DSML output parses cleanly through the recipe’s own deepseek_v4 parser. Quality spot-checks (code correctness, instruction following, arithmetic reasoning at T=0) came out equal-or-better versus the IQ2XXS quant; notably the reasoning probe finishes its answer where the 2-bit imatrix quant tends to exhaust its budget.

Two gotchas worth knowing before you launch:

  1. Empty env strings crash the kernel compiler. The compose file passes several variables as ${VAR:-}, which Docker delivers as empty strings, not unset vars. On our box that surfaced as a crash loop deep in the CUTE DSL during MLA kernel compile — KeyError: ‘’ in an enum lookup — before the server ever became healthy. Fix: set them explicitly in .env (at minimum CUTE_DSL_ARCH=sm_121a on Spark, plus GPU_MEMORY_UTILIZATION and MAX_MODEL_LEN). After that, clean boot every time.
  2. Don’t repeat entrypoint flags in EXTRA_VLLM_ARGS. The entrypoint already sets the correct --tool-call-parser deepseek_v4 (and --served-model-name, --port, …). Anything you add via EXTRA_VLLM_ARGS is appended after those, and vLLM’s arg parsing lets the last occurrence win — it even logs a “Found duplicate keys” warning that’s easy to miss. I initially added --tool-call-parser deepseek_v3 out of habit and got raw DSML markup in content with tool_calls: null, which looks exactly like a broken model. It isn’t — it’s a config foot-gun. Only add flags, never override, and grep the boot log for “duplicate keys” if tools misbehave.