Qwen3.8-Flash-Next

1st 4-bit, 117GB - but for MX Vontra/Qwen3.8-Flash-Next-MLX-4bit at main

4-bit looks like it will fit native in our Spark (before they get the offload working) ~100GB

What I’m reading there is that vLLM can’t offload to disk yet, so it’s not ready for 1x DGX Spark users yet.

and likely won’t be - its enterprise-first engine, they keep everything in ram, its aplenty. unless heavy mods are added that facilitate that

OK, I will be sure that we can runs this model on our single spark!!.

Full FP8 = 180GB

NVFP4 + BF16 Ngram = 186GB

NVFP4 + FP8 Ngram =135 GB ← Radixar

Now we have two way

1 Quantization Ngram to 6 5 4 bit → this will be lower size to around < 105GB

2 Offload Ngram to SSD

FP8 version should be run on 2x dgx spark perfectly…any one ran it successed ?

I will try with sglang later today, curious to hear if anoyone already had success.

Got it running with eager, now playing around

@tonyd615 is live streaming right now on yt messing qwen

Got a recipe?

Not really, far from finished haha

For the GGUF from unsloth, here are the requirements: Qwen3.8-Flash-Next: How to Run Locally | Unsloth Documentation.

EDIT:

llama.cpp is not ready yet:
model: add Qwen3.8-Flash-Next (qwen4exp) by danielhanchen · Pull Request #27742 · ggml-org/llama.cpp

Has anyone already had success with a single DGX Spark and NVFP4 from RadixArk?

Unsloth: Run Qwen3.8-Flash-Next Locally on 75GB unified memory (guide)

If smaller quants lose vision, I’m unsure of its validity to unseat DSF4 on a single gb10.

ugh I can’t sell this 5090 fast enough to pick up another lil buddy box.

Unsloth also release Q4 : unsloth/Qwen3.8-Flash-Next-GGUF · Hugging Face . I’am going to perform some test now.

Already getting interesting - here is a nice path forward for us: “The full 95.37 GiB PLE sidecar must never become resident in the 128 GB unified-memory pool.”

Also on Code Arena 3.8-Flash-Next is just 2 positions off of Claude Fabel 5! Just WoW…

better quality or higher decode… tempting to jump at GLM but we need proper spec decode for it…

What was your verdict with flash next? Tony managed to run it on 2 sparks at 50 t/s but I assume sglang

Glm is big, so cache likely be under 200k tokens

56 tk/s with EP and 2.4K prefill on MTP3

Edit: but was text only, had some early issues with boot… now going for vision as well

very close to jump to GLM… have a feeling that might be better to work on