Starting a thread now. It’s likely (supported by Mitko Vasilev’s testing) that this combination will revolutionize the GB10 platform.
DDTree page:
DDTree repo:
The Concept
Recently, diffusion drafting models became available which efficiently project forward many tokens - these clearly have a lot of potential. The bottleneck becomes checking/verification rather than token drafting. However, anything amiss with the drafted sequence, due to a wrong pick or unusual vocab, and the drafted sequence is rejected fairly early.
This is because baseline DFlash builds an entire projected draft sequence but then only that specific sequence is checked. Any wrong calls usually end up having the rest of the tree pruned.
DDTree - or Diffusion Draft Tree - builds an entire tree with probabilities at every position rather than one fixed sequence and then verifies iteratively based on per-position predictions. FAR higher acceptance rates. Verification budget remains the overall constraint.
See the animation on the linked official page above, which shows it intuitively.
What you see is far higher acceptance rates, better performance, resilience for unusual or atypical vocabularies.
What it costs is a small amount of bonus VRAM for the drafted tree instead of a single static sequence. We have that to spare.
Implementation
Proof of concept code in above repo. Mitko Vasilev has implemented into vLLM and per X posts with logs shows 80+ tok/s with Qwen3.5-27B AWQ quant on GB10. Not sure if links to X posts are allowed here but you can search for yourself. He says he plans to post this to his vllm-turboquant repo.
Watch for activity here: