DeepSeek-v4-Flash-0731-DSpark-1M-NVFP4-KV-2x-DGX-Spark

Got the NEW Deepseek up and running. There was a small config change you have t do to get Acceptence back up. For how it ships it will hit 30 tok/s but with the config change it brought acceptence back to its normal speeds as preview.

AGENT-

1/ 🔍 Upgraded to the official DeepSeek-V4-Flash-0731 on 2× DGX Spark. Decode throughput halved — with zero loss in output quality.

That combination is diagnostic. Spec decoding is verified by the target model, so a bad draft can only cost you speed, never correctness. 🧵

2/ Which means “slower but perfect output” points at the drafter, not the weights and not your config. Everyone (me included) looks at the wrong thing first.

steps/s was pinned at 14.4 the whole time. The entire deficit was draft acceptance. 📉

3/ vLLM’s DSpark draft loader renames shared_experts.w2 → down_proj, but never maps w1/w3 → gate_up_proj.

They match nothing → logger.debug("Skipping unknown DSpark weight") → invisible at INFO. 🫥

12 tensors gone. The always-on shared expert, uninitialized, in all 3 draft stages.

4/ The target model’s own loader has the exact two rows the draft loader is missing. They were lost when the mapping got narrowed to dodge a markov_w1 name collision. 🪤

5/ Fix 🩹

("shared_experts.gate_up_proj", ".shared_experts.w1", 0),

("shared_experts.gate_up_proj", ".shared_experts.w3", 1),

⚡ 32.7 → 55.4 tok/s mean (+69%), 66.1 peak
📈 acceptance 25.7% → 60.2%
per-position 0.63/0.28/0.18/0.11/0.07 → 0.83/0.73/0.57/0.47/0.40

6/ ⚠️ Bonus trap: under spec decode vLLM emits at most one SSE chunk per decode step, carrying every token accepted that step.

Counting stream deltas measures steps/s, not tok/s. Same request: 14.7 vs 60.1. Benchmark with stream: false. 📏

7/ Full write-up, the patch, and the measured before/after 👇
[repo link]

Runs on 2× DGX Spark, TP=2, k=5, NVFP4 KV, 1M context 🖥️🖥️

WELL DONE Tony! finally DS4F is usable :)

Amazing, thanks! Will try it the upcoming days!

If we do not use NVFP4 KV cache, how much KV cache size can we get with this? Approximately half?

This is impressive, and thanks for your work!

Nvidia really has to thank Deepseek, as I just ordered a new Spark because of this new release 😂

Hey Tony, is this expected:
=> => naming to Docker Hub Container Image Library | App Containerization 0.0s
INFO 08-01 14:48:04 [importing.py:46] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors.
INFO 08-01 14:48:04 [importing.py:81] Triton not installed or not compatible; certain GPU-related functions will not be available.
W0801 14:48:04.663000 1 site-packages/torch/utils/cpp_extension.py:140] No CUDA runtime is found, using CUDA_HOME=‘/opt/env/targets/sbsa-linux’
WARNING 08-01 14:48:04 [interface.py:247] Failed to import from vllm._C: ImportError(‘/lib/aarch64-linux-gnu/libcuda.so.1: file too short’)
WARNING 08-01 14:48:05 [interface.py:247] Failed to import from vllm._C: ImportError(‘/lib/aarch64-linux-gnu/libcuda.so.1: file too short’)
WARNING 08-01 14:48:05 [interface.py:247] Failed to import from vllm._C: ImportError(‘/lib/aarch64-linux-gnu/libcuda.so.1: file too short’)
WARNING 08-01 14:48:05 [interface.py:247] Failed to import from vllm._C: ImportError(‘/lib/aarch64-linux-gnu/libcuda.so.1: file too short’)
dspark overlay ok vllm.v1.spec_decode.dspark vllm.v1.spec_decode.dspark_proposer
[+] Building 0.6s (6/6) FINISHED docker:default

I have no clue what that is brother. Did you ask your agent

It’s from the docker build. I’ll ask my agent now that the build succeeded (and the agent has a brain again), but I wondered if you saw similar output on your docker and just ignored

I got this going, thanks! Haven’t tested the upper realms beyond 300k yet, but it seems to work like the older dspark setup I had going before, maybe a touch smarter and faster even, with larger context room. All good things.

You’re a real one bro, ty.

@tonyd615 Very curious how you test the token generation speed. Also what concurrency and depth are you testing at. With eugr/llama-benchy i get different results (lower) than those you published. Thanks.

Thanks Tony for sharing the recipe. It actually worked for me on my system too.

However, NVFP4 for KV-cache didn’t work for me, despite incorporating your patches.
Also only getting 12tok/s because >90% of draft tokens are being rejected.

Here are the exact measured performance results from the live cluster:

Measured Generation Performance

• Measured Decode Speed: 12.03 tokens/sec
• Total Latency: 42.57 s for 512 tokens
• Time-to-First-Token (TTFT): ~261 ms

Analysis of the DSpark Speculative Acceptance

In the live server logs, vLLM’s metrics.py reports:

• Draft Model Proposal Rate: ~53 to 55 tokens/sec
• Average Draft Acceptance Rate: 1.5% – 4.5%
• Mean Acceptance Length: 1.08 to 1.22 tokens

  1. Overlay Code: Inspected /model/patch/overlay/vllm/v1/spec_decode/dspark.py. Both lines mapping w1 and w3 into gate_up_proj are present.
  2. Log Verification: Inspected backend_node0.log. Zero missing DSpark weight warnings (Skipping unknown DSpark weight), confirming all 12 draft shared expert tensors
    were loaded into GPU memory.

Based on further testing,
⛬ The checkpoint inspection adds a stronger lead than the original prompt anticipated: the trunk experts are packed NVFP4 (uint8), while the three MTP/DSpark expert blocks are stored as FP8-style int8 weights with UE8M0 scales. The live drafter nevertheless enters the generic ModelOpt NVFP4 post-load path and discards the distinct W3 global scale. I’m checking whether this is an expected mixed-format conversion in the canonical recipe or a faulty port.

These are the fixes I had to run with my coding agent to get a performant model —/
🛠️ Complementary Fixes for DeepSeek-V4-Flash DSpark

This document complements the gold-standard production recipe and explains what we had to change, and why, to turn the stable-but-slow deployment into a high-acceptance DSpark configuration.

The original recipe is stable: it boots reliably on both DGX Spark nodes and serves the deepseek-v4-flash model. However, its measured non-streaming throughput was 14–18 tok/s (the original recipe even reports ~10 tok/s streaming). The root cause was not a missing patch or a bad MoE backend: it was DSpark draft quantization-config inheritance.


🔍 1. Root Cause: DSpark Draft Inherits the Target’s NVFP4 Config

The DeepSeek-V4-Flash-0731-NVFP4 checkpoint is a hybrid:

  • Target trunk: ModelOpt NVFP4 experts, serialized in the ModelOpt format, requiring flashinfer_b12x and the ModelOptNvFp4FusedMoE quant method.
  • MTP / DSpark draft stages: native FP8/MXFP4 experts, with int8 weights and UE8M0 (e8m0fnu) .scale tensors, requiring the native Mxfp4MoEMethod quant method.

vLLM PR #49133 addresses exactly this: the DSpark model-type rewrite (deepseek_v4_dspark) can leave the draft model_config.quantization stuck at plain "fp8", and the draft VllmConfig can inherit the target’s quant_config. The result is that the draft MoE is built with ModelOptNvFp4FusedMoE, even though the draft weights are not in the ModelOpt format.

Observed symptoms in the original recipe:

WARNING ... [modelopt.py:1548] w1_weight_scale_2 must match w3_weight_scale_2. Accuracy may be affected.
  • Draft acceptance collapses to ~1.0–1.15 tokens per step.
  • Per-position acceptance at positions 1–4 falls to 0.44%–3.2%.
  • Throughput is capped at 14–18 tok/s.

The shared-expert mapping patch (Patch 4) and the Keys concurrency patch were already present and byte-identical to the canonical recipe, so they were not the problem.


🧩 2. Fix 1: Draft-Only Quantization Metadata Normalization

When the speculative config constructs the draft model_config, the target’s quantization_config is still attached to the draft’s hf_config. We must strip the target-only ModelOpt keys from the draft copy before the draft quant_config is derived.

The patch is applied in recipe/apply_experiment_patches.py:

# DSpark draft quantization metadata normalisation:
# the checkpoint's quantization_config carries ModelOpt NVFP4
# keys (moe_quant_algo, quantized_layers, ignore) that are meant
# for the target trunk only. If the draft inherits them,
# DeepseekV4FP8Config will dispatch routed experts through
# ModelOptNvFp4FusedMoE instead of the native MXFP4 path.
if self.method == "dspark":
    draft_qc = copy.deepcopy(
        getattr(self.draft_model_config.hf_config, "quantization_config", {})
        or {}
    )
    for _target_only_key in (
        "moe_quant_algo",
        "quantized_layers",
        "ignore",
        "modules_to_not_convert",
    ):
        draft_qc.pop(_target_only_key, None)
    self.draft_model_config.hf_config.quantization_config = draft_qc

The target quantization_config is not modified; only the deep-copied draft copy is sanitized.

This patch also rewrites the draft quantization from plain "fp8" to "deepseek_v4_fp8" and builds the draft under its own freshly derived quant_config, matching the two halves of vLLM PR #49133.


🛡️ 3. Fix 2: Loader Fail-Closed Checks

We added two hard checks in overlay/vllm/models/deepseek_v4/nvidia/dspark.py so that a regression cannot silently pass.

3.1 Draft quant-path guard

def _verify_draft_quant_path(self, vllm_config: VllmConfig) -> None:
    quant_config = vllm_config.quant_config
    moe_quant_algo = getattr(quant_config, "moe_quant_algo", None)
    logger.info(
        "DSpark draft quantization path: quant_config=%s, moe_quant_algo=%s",
        type(quant_config).__name__, moe_quant_algo,
    )
    if moe_quant_algo == "NVFP4":
        raise RuntimeError(
            "DSpark draft resolved to ModelOpt NVFP4 (moe_quant_algo=NVFP4). "
            "The draft must use the native MXFP4 expert path."
        )

3.2 MTP weight-coverage check

After load_weights, we require that every DSpark draft stage has loaded both routed and shared expert weight families:

for stage_id in range(self.config.dspark_num_draft_layers):
    virtual_layer_id = self.model.dspark_start_layer_idx + stage_id
    prefix = f"model.layers.{virtual_layer_id}.ffn"
    routed_loaded = any(p.startswith(f"{prefix}.experts.") for p in loaded_params)
    shared_loaded = any(p.startswith(f"{prefix}.shared_experts.") for p in loaded_params)
    if not routed_loaded or not shared_loaded:
        raise RuntimeError(f"DSpark failed to load required weight families: {missing_families}")

🔧 4. Fix 3: Draft MoE Backend Must Be b12x, Not flashinfer_b12x

After Fix 1, the draft correctly resolves to the MXFP4 path. The MXFP4 oracle in model_executor/layers/fused_moe/oracle/mxfp4.py accepts b12x but not flashinfer_b12x. If the draft is told to use flashinfer_b12x, startup crashes with:

ValueError: moe_backend='flashinfer_b12x' is not supported for MXFP4 MoE.
Expected one of ['b12x', 'deep_gemm', ...].

The target trunk still uses --moe-backend flashinfer_b12x for its NVFP4 experts. The draft must use the b12x alias via the speculative config:

--speculative-config '{"method":"dspark","num_speculative_tokens":5,"moe_backend":"b12x"}'

This is a subtle but critical distinction: the original recipe worked only because the draft was silently on the wrong (NVFP4) path. Once the draft is corrected to MXFP4, its backend name must be corrected too.


⚡ 5. Fix 4: Complete the nvfp4_ds_mla KV-Cache Plumbing

The original recipe uses --kv-cache-dtype fp8_ds_mla. The canonical recipe enables --kv-cache-dtype nvfp4_ds_mla, which gives a larger effective context pool.

The local overlay was missing the canonical Stage-A/B/C plumbing:

Stage File Change
A utils/torch_utils.py Add "nvfp4_ds_mla": torch.uint8 and recognize it as a quantized cache dtype
B models/deepseek_v4/attention.py Accept nvfp4/nvfp4_ds_mla and switch the backend to nvfp4_ds_mla
B models/deepseek_v4/nvidia/flashmla.py Return the KV cache shape for nvfp4_ds_mla
C v1/kv_cache_interface.py Use the validated 584-byte DeepSeek-V4 padded envelope
C models/deepseek_v4/attention.py Correct the probe from 416 bytes to 584 bytes
C models/deepseek_v4/nvidia/flashmla.py Correct get_kv_cache_shape to 584 bytes

These are applied in recipe/apply_experiment_patches.py and verified in recipe/Dockerfile.experiment.


⚙️ 6. Configuration Differences from the Original Recipe

Setting Original recipe Stage-D (this fix) Why
Container image aidendle94/sparkrun-vllm-ds4-gb10:production-ready dsv4flash-experiment:dspark-nvfp4-stage-d New image contains the patches and loader checks
KV cache dtype fp8_ds_mla nvfp4_ds_mla Larger context pool; canonical NVFP4 MLA path
Main MoE backend flashinfer_b12x flashinfer_b12x Unchanged for the NVFP4 target trunk
Draft MoE backend implicit flashinfer_b12x b12x (via speculative config) Required once the draft is on the MXFP4 path
Attention backend B12X_MLA_SPARSE AUTO AUTO path correctly selects the B12X sparse MLA backend with nvfp4_ds_mla
Position-0 diagnostics off optional (stage-d on, stage-d-final off) Useful for debugging, then disabled for clean throughput

🚀 7. Stage-D Backend Runner (ds4_b12x_backend_experiment_stage_d.sh)

The key differences from the original ds4_b12x_backend.sh are:

  • VLLM_ATTENTION_BACKEND is not forced to B12X_MLA_SPARSE.
  • --kv-cache-dtype nvfp4_ds_mla.
  • --speculative-config includes "moe_backend":"b12x" for the draft.
  • --moe-backend flashinfer_b12x remains for the target trunk.
exec /usr/local/bin/dsv4-vllm-entrypoint serve /model \
  --served-model-name deepseek-v4-flash \
  --host 0.0.0.0 \
  --port 1234 \
  --trust-remote-code \
  --tensor-parallel-size 2 \
  --pipeline-parallel-size 1 \
  --kv-cache-dtype nvfp4_ds_mla \
  --block-size 256 \
  --max-model-len 524288 \
  --max-num-seqs 4 \
  --max-num-batched-tokens 8192 \
  --gpu-memory-utilization 0.8 \
  --enable-prefix-caching \
  --speculative-config '{"method":"dspark","num_speculative_tokens":5,"moe_backend":"b12x"}' \
  --tokenizer-mode deepseek_v4 \
  --distributed-executor-backend mp \
  --tool-call-parser deepseek_v4 \
  --enable-auto-tool-choice \
  --reasoning-parser deepseek_v4 \
  --default-chat-template-kwargs.thinking=true \
  --default-chat-template-kwargs.reasoning_effort=high \
  --moe-backend flashinfer_b12x \
  --enable-flashinfer-autotune \
  --nnodes 2 \
  --node-rank "${NODE_RANK}" \
  --master-addr 192.168.178.48 \
  --master-port 25000 \
  ${HEADLESS:+--headless}

Full scripts are available at:

  • /home/sumergoconicio/amanuensis/gitea/tonydwild-dsv4flash-glytchbot/scripts/ds4_b12x_backend_experiment_stage_d.sh
  • /home/sumergoconicio/amanuensis/gitea/tonydwild-dsv4flash-glytchbot/scripts/experiment-launch-stage-d-final.sh

📊 8. Validation Results

Startup signals:

DSpark draft quantization path: quant_config=DeepseekV4FP8Config, moe_quant_algo=
DeepSeek V4 expert_dtype resolved to 'fp4'
Using nvfp4_ds_mla data type to store kv cache
Application startup complete
  • No w1_weight_scale_2 must match w3_weight_scale_2 warning.
  • No startup errors, CUDA errors, or worker failures.

Acceptance (diagnostic run with position-0 diagnostics)

Metric Value
Position-0 target-argmax agreement 89–91%
Mean acceptance length ~4.5 tokens
Average draft acceptance rate ~70%

Throughput (dsv4flash_bench.py, non-streaming, temperature 0)

Prompt Original recipe Stage-D diagnostic Stage-D final (no diagnostics)
count 1..300 16.9 tok/s 64.36 tok/s 49.68 tok/s
60 SQL INSERTs 18.1 tok/s 53.78 tok/s 58.52 tok/s
Python BST 17.2 tok/s 48.67 tok/s 47.54 tok/s
200-word narrative 14.6 tok/s 31.81 tok/s 25.91 tok/s
Average 16.7 ~49.7 ~45.4

Both Stage-D runs are ~2.7–3.0× faster than the original recipe and comfortably exceed the 30 tok/s gate on structured/code prompts.


📁 9. Related Files

  • Experiment workspace: /home/sumergoconicio/amanuensis/gitea/tonydwild-dsv4flash-glytchbot
  • Patch builder: /home/sumergoconicio/amanuensis/gitea/tonydwild-dsv4flash-glytchbot/recipe/apply_experiment_patches.py
  • DSpark overlay: /home/sumergoconicio/amanuensis/gitea/tonydwild-dsv4flash-glytchbot/recipe/overlay/vllm/models/deepseek_v4/nvidia/dspark.py
  • Experiment Dockerfile: /home/sumergoconicio/amanuensis/gitea/tonydwild-dsv4flash-glytchbot/recipe/Dockerfile.experiment
  • Full optimisation report: /home/sumergoconicio/Documents/Code/ChAI/artifacts/dsv4flash-optimisation.md
  • Benchmark script: /home/sumergoconicio/Documents/Code/ChAI/artifacts/dsv4flash_bench.py

✅ 10. Adoption Status

The dsv4flash-experiment:dspark-nvfp4-stage-d deployment is currently serving the model. The original aidendle94/sparkrun-vllm-ds4-gb10:production-ready image and its certified-working backup launcher remain available as the rollback checkpoint.

Just wanted to say thanks to everyone pushing the Spark’s support and development forward. It really is a very important device for those who cant afford a full data center but value self hosting. We would have a very different ownership experience without your contribution. It will be great one day when vLLM recognizes your hard work and includes your contributions into the main branch?

The implementation of the DS4F0731 model card configuration options also makes a significant improvement to overall agentic ability, something like:

“options”:
{
“chat_template_kwargs”: { “thinking”: true, “reasoning_effort”: “max” },
“temperature”: 1.0,
“top_p”: 0.95,
“max_tokens”: 32000
}

Thanks again to this community - I’m learning a lot from everyone here.

I deployed the model on dual sparks using your recipe tony, it’s working good but sometimes when context increases, token generation randomly drops to 1 tok/s for few seconds and then jumps back again to 40-50. does anyone know why this is happening on my end?

Hi tony, on vllm 0.21. I’m getting errors with claude code, I heard that was fixed with vllm 0.23. Is there any way I can upgrade to vllm 0.23 to run this

updates being pushed will be following up more soon

my token generation speed drops to 1 tok/s randomly after context increases and then jumps back again, do you know why this might be happening?

I’m running at fp8 kv cache without the patches, might that be the issue?

have you updated from the repo at all ?

Yes, I believe my env is the issue here, maybe there is some issues in communication speed between both sparks, this is how it looks.

NCCL_IB_HCA=rocep1s0f1,roceP2p1s0f1
NCCL_SOCKET_IFNAME=enp1s0f1np1
NCCL_IB_GID_INDEX=0
NCCL_CROSS_NIC=1

also getting this in logs:
[2026-08-23 06:12:43] dgx-spark-ai01:78:106 [0] transport/net_ib/common.cc:145 NCCL WARN NET/IB : roceP2p1s0f1:1 GID table changed