Boy – Qwen 3.8-Next looks like an ideal model for two DGX Sparks. Sparse architecture but takes advantage of the generous available memory. They explicitly mention the model in their breathless press release:
"Beyond rack-scale deployment, Qwen3.8-Flash-Next also runs on local NVIDIA hardware, including NVIDIA DGX Station, NVIDIA DGX Spark clusters, "
So why not, you know, publish a configuration that actually works? Particularly in the NVFP4 format that’s perfect for this model.
Now I know I can download some zero star vibe-coded hack of a vllm container with random patches to sort of get it running. Or worse, go down an entirely new rabbit hole with SGLang. But WHY does a company that literally makes so much money they don’t know what to do with it completely abuse its enthusiast base?