These are the fixes I had to run with my coding agent to get a performant model —/
🛠️ Complementary Fixes for DeepSeek-V4-Flash DSpark
This document complements the gold-standard production recipe and explains what we had to change, and why, to turn the stable-but-slow deployment into a high-acceptance DSpark configuration.
The original recipe is stable: it boots reliably on both DGX Spark nodes and serves the deepseek-v4-flash model. However, its measured non-streaming throughput was 14–18 tok/s (the original recipe even reports ~10 tok/s streaming). The root cause was not a missing patch or a bad MoE backend: it was DSpark draft quantization-config inheritance.
🔍 1. Root Cause: DSpark Draft Inherits the Target’s NVFP4 Config
The DeepSeek-V4-Flash-0731-NVFP4 checkpoint is a hybrid:
- Target trunk: ModelOpt NVFP4 experts, serialized in the ModelOpt format, requiring
flashinfer_b12x and the ModelOptNvFp4FusedMoE quant method.
- MTP / DSpark draft stages: native FP8/MXFP4 experts, with
int8 weights and UE8M0 (e8m0fnu) .scale tensors, requiring the native Mxfp4MoEMethod quant method.
vLLM PR #49133 addresses exactly this: the DSpark model-type rewrite (deepseek_v4_dspark) can leave the draft model_config.quantization stuck at plain "fp8", and the draft VllmConfig can inherit the target’s quant_config. The result is that the draft MoE is built with ModelOptNvFp4FusedMoE, even though the draft weights are not in the ModelOpt format.
Observed symptoms in the original recipe:
WARNING ... [modelopt.py:1548] w1_weight_scale_2 must match w3_weight_scale_2. Accuracy may be affected.
- Draft acceptance collapses to ~1.0–1.15 tokens per step.
- Per-position acceptance at positions 1–4 falls to 0.44%–3.2%.
- Throughput is capped at 14–18 tok/s.
The shared-expert mapping patch (Patch 4) and the Keys concurrency patch were already present and byte-identical to the canonical recipe, so they were not the problem.
🧩 2. Fix 1: Draft-Only Quantization Metadata Normalization
When the speculative config constructs the draft model_config, the target’s quantization_config is still attached to the draft’s hf_config. We must strip the target-only ModelOpt keys from the draft copy before the draft quant_config is derived.
The patch is applied in recipe/apply_experiment_patches.py:
# DSpark draft quantization metadata normalisation:
# the checkpoint's quantization_config carries ModelOpt NVFP4
# keys (moe_quant_algo, quantized_layers, ignore) that are meant
# for the target trunk only. If the draft inherits them,
# DeepseekV4FP8Config will dispatch routed experts through
# ModelOptNvFp4FusedMoE instead of the native MXFP4 path.
if self.method == "dspark":
draft_qc = copy.deepcopy(
getattr(self.draft_model_config.hf_config, "quantization_config", {})
or {}
)
for _target_only_key in (
"moe_quant_algo",
"quantized_layers",
"ignore",
"modules_to_not_convert",
):
draft_qc.pop(_target_only_key, None)
self.draft_model_config.hf_config.quantization_config = draft_qc
The target quantization_config is not modified; only the deep-copied draft copy is sanitized.
This patch also rewrites the draft quantization from plain "fp8" to "deepseek_v4_fp8" and builds the draft under its own freshly derived quant_config, matching the two halves of vLLM PR #49133.
🛡️ 3. Fix 2: Loader Fail-Closed Checks
We added two hard checks in overlay/vllm/models/deepseek_v4/nvidia/dspark.py so that a regression cannot silently pass.
3.1 Draft quant-path guard
def _verify_draft_quant_path(self, vllm_config: VllmConfig) -> None:
quant_config = vllm_config.quant_config
moe_quant_algo = getattr(quant_config, "moe_quant_algo", None)
logger.info(
"DSpark draft quantization path: quant_config=%s, moe_quant_algo=%s",
type(quant_config).__name__, moe_quant_algo,
)
if moe_quant_algo == "NVFP4":
raise RuntimeError(
"DSpark draft resolved to ModelOpt NVFP4 (moe_quant_algo=NVFP4). "
"The draft must use the native MXFP4 expert path."
)
3.2 MTP weight-coverage check
After load_weights, we require that every DSpark draft stage has loaded both routed and shared expert weight families:
for stage_id in range(self.config.dspark_num_draft_layers):
virtual_layer_id = self.model.dspark_start_layer_idx + stage_id
prefix = f"model.layers.{virtual_layer_id}.ffn"
routed_loaded = any(p.startswith(f"{prefix}.experts.") for p in loaded_params)
shared_loaded = any(p.startswith(f"{prefix}.shared_experts.") for p in loaded_params)
if not routed_loaded or not shared_loaded:
raise RuntimeError(f"DSpark failed to load required weight families: {missing_families}")
🔧 4. Fix 3: Draft MoE Backend Must Be b12x, Not flashinfer_b12x
After Fix 1, the draft correctly resolves to the MXFP4 path. The MXFP4 oracle in model_executor/layers/fused_moe/oracle/mxfp4.py accepts b12x but not flashinfer_b12x. If the draft is told to use flashinfer_b12x, startup crashes with:
ValueError: moe_backend='flashinfer_b12x' is not supported for MXFP4 MoE.
Expected one of ['b12x', 'deep_gemm', ...].
The target trunk still uses --moe-backend flashinfer_b12x for its NVFP4 experts. The draft must use the b12x alias via the speculative config:
--speculative-config '{"method":"dspark","num_speculative_tokens":5,"moe_backend":"b12x"}'
This is a subtle but critical distinction: the original recipe worked only because the draft was silently on the wrong (NVFP4) path. Once the draft is corrected to MXFP4, its backend name must be corrected too.
⚡ 5. Fix 4: Complete the nvfp4_ds_mla KV-Cache Plumbing
The original recipe uses --kv-cache-dtype fp8_ds_mla. The canonical recipe enables --kv-cache-dtype nvfp4_ds_mla, which gives a larger effective context pool.
The local overlay was missing the canonical Stage-A/B/C plumbing:
| Stage |
File |
Change |
| A |
utils/torch_utils.py |
Add "nvfp4_ds_mla": torch.uint8 and recognize it as a quantized cache dtype |
| B |
models/deepseek_v4/attention.py |
Accept nvfp4/nvfp4_ds_mla and switch the backend to nvfp4_ds_mla |
| B |
models/deepseek_v4/nvidia/flashmla.py |
Return the KV cache shape for nvfp4_ds_mla |
| C |
v1/kv_cache_interface.py |
Use the validated 584-byte DeepSeek-V4 padded envelope |
| C |
models/deepseek_v4/attention.py |
Correct the probe from 416 bytes to 584 bytes |
| C |
models/deepseek_v4/nvidia/flashmla.py |
Correct get_kv_cache_shape to 584 bytes |
These are applied in recipe/apply_experiment_patches.py and verified in recipe/Dockerfile.experiment.
⚙️ 6. Configuration Differences from the Original Recipe
| Setting |
Original recipe |
Stage-D (this fix) |
Why |
| Container image |
aidendle94/sparkrun-vllm-ds4-gb10:production-ready |
dsv4flash-experiment:dspark-nvfp4-stage-d |
New image contains the patches and loader checks |
| KV cache dtype |
fp8_ds_mla |
nvfp4_ds_mla |
Larger context pool; canonical NVFP4 MLA path |
| Main MoE backend |
flashinfer_b12x |
flashinfer_b12x |
Unchanged for the NVFP4 target trunk |
| Draft MoE backend |
implicit flashinfer_b12x |
b12x (via speculative config) |
Required once the draft is on the MXFP4 path |
| Attention backend |
B12X_MLA_SPARSE |
AUTO |
AUTO path correctly selects the B12X sparse MLA backend with nvfp4_ds_mla |
| Position-0 diagnostics |
off |
optional (stage-d on, stage-d-final off) |
Useful for debugging, then disabled for clean throughput |
🚀 7. Stage-D Backend Runner (ds4_b12x_backend_experiment_stage_d.sh)
The key differences from the original ds4_b12x_backend.sh are:
VLLM_ATTENTION_BACKEND is not forced to B12X_MLA_SPARSE.
--kv-cache-dtype nvfp4_ds_mla.
--speculative-config includes "moe_backend":"b12x" for the draft.
--moe-backend flashinfer_b12x remains for the target trunk.
exec /usr/local/bin/dsv4-vllm-entrypoint serve /model \
--served-model-name deepseek-v4-flash \
--host 0.0.0.0 \
--port 1234 \
--trust-remote-code \
--tensor-parallel-size 2 \
--pipeline-parallel-size 1 \
--kv-cache-dtype nvfp4_ds_mla \
--block-size 256 \
--max-model-len 524288 \
--max-num-seqs 4 \
--max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.8 \
--enable-prefix-caching \
--speculative-config '{"method":"dspark","num_speculative_tokens":5,"moe_backend":"b12x"}' \
--tokenizer-mode deepseek_v4 \
--distributed-executor-backend mp \
--tool-call-parser deepseek_v4 \
--enable-auto-tool-choice \
--reasoning-parser deepseek_v4 \
--default-chat-template-kwargs.thinking=true \
--default-chat-template-kwargs.reasoning_effort=high \
--moe-backend flashinfer_b12x \
--enable-flashinfer-autotune \
--nnodes 2 \
--node-rank "${NODE_RANK}" \
--master-addr 192.168.178.48 \
--master-port 25000 \
${HEADLESS:+--headless}
Full scripts are available at:
/home/sumergoconicio/amanuensis/gitea/tonydwild-dsv4flash-glytchbot/scripts/ds4_b12x_backend_experiment_stage_d.sh
/home/sumergoconicio/amanuensis/gitea/tonydwild-dsv4flash-glytchbot/scripts/experiment-launch-stage-d-final.sh
📊 8. Validation Results
Startup signals:
DSpark draft quantization path: quant_config=DeepseekV4FP8Config, moe_quant_algo=
DeepSeek V4 expert_dtype resolved to 'fp4'
Using nvfp4_ds_mla data type to store kv cache
Application startup complete
- No
w1_weight_scale_2 must match w3_weight_scale_2 warning.
- No startup errors, CUDA errors, or worker failures.
Acceptance (diagnostic run with position-0 diagnostics)
| Metric |
Value |
| Position-0 target-argmax agreement |
89–91% |
| Mean acceptance length |
~4.5 tokens |
| Average draft acceptance rate |
~70% |
Throughput (dsv4flash_bench.py, non-streaming, temperature 0)
| Prompt |
Original recipe |
Stage-D diagnostic |
Stage-D final (no diagnostics) |
| count 1..300 |
16.9 tok/s |
64.36 tok/s |
49.68 tok/s |
| 60 SQL INSERTs |
18.1 tok/s |
53.78 tok/s |
58.52 tok/s |
| Python BST |
17.2 tok/s |
48.67 tok/s |
47.54 tok/s |
| 200-word narrative |
14.6 tok/s |
31.81 tok/s |
25.91 tok/s |
| Average |
16.7 |
~49.7 |
~45.4 |
Both Stage-D runs are ~2.7–3.0× faster than the original recipe and comfortably exceed the 30 tok/s gate on structured/code prompts.
📁 9. Related Files
- Experiment workspace:
/home/sumergoconicio/amanuensis/gitea/tonydwild-dsv4flash-glytchbot
- Patch builder:
/home/sumergoconicio/amanuensis/gitea/tonydwild-dsv4flash-glytchbot/recipe/apply_experiment_patches.py
- DSpark overlay:
/home/sumergoconicio/amanuensis/gitea/tonydwild-dsv4flash-glytchbot/recipe/overlay/vllm/models/deepseek_v4/nvidia/dspark.py
- Experiment Dockerfile:
/home/sumergoconicio/amanuensis/gitea/tonydwild-dsv4flash-glytchbot/recipe/Dockerfile.experiment
- Full optimisation report:
/home/sumergoconicio/Documents/Code/ChAI/artifacts/dsv4flash-optimisation.md
- Benchmark script:
/home/sumergoconicio/Documents/Code/ChAI/artifacts/dsv4flash_bench.py
✅ 10. Adoption Status
The dsv4flash-experiment:dspark-nvfp4-stage-d deployment is currently serving the model. The original aidendle94/sparkrun-vllm-ds4-gb10:production-ready image and its certified-working backup launcher remain available as the rollback checkpoint.