1st 4-bit, 117GB - but for MX Vontra/Qwen3.8-Flash-Next-MLX-4bit at main
4-bit looks like it will fit native in our Spark (before they get the offload working) ~100GB
1st 4-bit, 117GB - but for MX Vontra/Qwen3.8-Flash-Next-MLX-4bit at main
4-bit looks like it will fit native in our Spark (before they get the offload working) ~100GB
What I’m reading there is that vLLM can’t offload to disk yet, so it’s not ready for 1x DGX Spark users yet.
and likely won’t be - its enterprise-first engine, they keep everything in ram, its aplenty. unless heavy mods are added that facilitate that
OK, I will be sure that we can runs this model on our single spark!!.
Full FP8 = 180GB
NVFP4 + BF16 Ngram = 186GB
NVFP4 + FP8 Ngram =135 GB ← Radixar
Now we have two way
1 Quantization Ngram to 6 5 4 bit → this will be lower size to around < 105GB
2 Offload Ngram to SSD
FP8 version should be run on 2x dgx spark perfectly…any one ran it successed ?
I will try with sglang later today, curious to hear if anoyone already had success.
Got it running with eager, now playing around
@tonyd615 is live streaming right now on yt messing qwen
Got a recipe?
Not really, far from finished haha
For the GGUF from unsloth, here are the requirements: Qwen3.8-Flash-Next: How to Run Locally | Unsloth Documentation.
EDIT:
llama.cpp is not ready yet:
model: add Qwen3.8-Flash-Next (qwen4exp) by danielhanchen · Pull Request #27742 · ggml-org/llama.cpp
Has anyone already had success with a single DGX Spark and NVFP4 from RadixArk?
Unsloth: Run Qwen3.8-Flash-Next Locally on 75GB unified memory (guide)
If smaller quants lose vision, I’m unsure of its validity to unseat DSF4 on a single gb10.
ugh I can’t sell this 5090 fast enough to pick up another lil buddy box.
Unsloth also release Q4 : unsloth/Qwen3.8-Flash-Next-GGUF · Hugging Face . I’am going to perform some test now.
better quality or higher decode… tempting to jump at GLM but we need proper spec decode for it…
What was your verdict with flash next? Tony managed to run it on 2 sparks at 50 t/s but I assume sglang
Glm is big, so cache likely be under 200k tokens
56 tk/s with EP and 2.4K prefill on MTP3
Edit: but was text only, had some early issues with boot… now going for vision as well
very close to jump to GLM… have a feeling that might be better to work on