DDTree plus diffusion drafting (DFlash) to optimize GB10

Starting a thread now. It’s likely (supported by Mitko Vasilev’s testing) that this combination will revolutionize the GB10 platform.

DDTree page:

DDTree repo:

The Concept

Recently, diffusion drafting models became available which efficiently project forward many tokens - these clearly have a lot of potential. The bottleneck becomes checking/verification rather than token drafting. However, anything amiss with the drafted sequence, due to a wrong pick or unusual vocab, and the drafted sequence is rejected fairly early.

This is because baseline DFlash builds an entire projected draft sequence but then only that specific sequence is checked. Any wrong calls usually end up having the rest of the tree pruned.

DDTree - or Diffusion Draft Tree - builds an entire tree with probabilities at every position rather than one fixed sequence and then verifies iteratively based on per-position predictions. FAR higher acceptance rates. Verification budget remains the overall constraint.

See the animation on the linked official page above, which shows it intuitively.

What you see is far higher acceptance rates, better performance, resilience for unusual or atypical vocabularies.

What it costs is a small amount of bonus VRAM for the drafted tree instead of a single static sequence. We have that to spare.

Implementation

Proof of concept code in above repo. Mitko Vasilev has implemented into vLLM and per X posts with logs shows 80+ tok/s with Qwen3.5-27B AWQ quant on GB10. Not sure if links to X posts are allowed here but you can search for yourself. He says he plans to post this to his vllm-turboquant repo.

Watch for activity here:

Has anyone been able to replicate these speeds?

Your best bet today is Lucebox-hub. It’s a hyper specific ultra light harness because it’s designed for a specific 4 bit GGUF and one RTX 3090. But it does actually now support consumer Blackwell with DFlash and DDTree.

I’ve compiled and run it, it works… but like I said rough around the edges. You have to dig into the code to figure out why it’s default capped at returning 512 tokens, it has a very low context cap because of DFlash attention being shared, so while you can crank it on GB10 it will really harm throughput. They have an Issue on that which, once fixed, may warrant a full reassessment.

He states :

• Qwen3.6-35B-A3B NVFP4 + DFlash:
91–97 tok/s single-stream

I’m getting numbers close to those with FP8 Quant + DFlash + vLLM tune (My before+after here: Introducing vLLM-Tune — Kernel tuning CLI for vLLM on DGX Spark - #9 by azampatti )

Running vLLM tune on my spark right now, I will send my results in the thread once it completes

This is awesome, but as yet I do not believe vLLM has DDtree, so there’s more potential once it gets added!