Working hard on implementing an atlas-version here:
Let me know if you care to test the recipe when its up, Qwen3.6 runs real smooth :)
Working hard on implementing an atlas-version here:
Let me know if you care to test the recipe when its up, Qwen3.6 runs real smooth :)
Got RadixArk/Qwen3.8-Flash-Next-NVFP4 working on 2 x GX10 using Vllm. There were issues with PLE offload to begin with, fixed that and its loading with 262k context now.
Hardware: 2x NVIDIA DGX Spark (GB10, sm_121a, aarch64), RoCE fabric between them
The two bugs:
Fix: detect FP8 PLE via ple_embedding_dtype == “float8_e4m3fn”; remove weight_scale param
61.7 GiB/rank, KV cache 3.16M tokens (12x at 262K), coherent generation, tool calls + reasoning work.
vllm serve /model \
--served-model-name qwen38-flash-next-nvfp4 \
--quantization modelopt_fp4 \
--tensor-parallel-size 2 \
--pipeline-parallel-size 1 \
--nnodes 2 \
--node-rank 0 \
--master-addr 10.10.10.1 \
--master-port 29501 \
--distributed-executor-backend mp \
--enforce-eager \
--max-model-len 262144 \
--max-num-seqs 8 \
--gpu-memory-utilization 0.85 \
--tool-call-parser qwen3_coder \
--enable-auto-tool-choice \
--reasoning-parser qwen3 \
--host 0.0.0.0 \
--port 8888
Let the fun commence :D
need some test results :)
Got the FP8 version running on 2xSpark, using the official vllm/vllm-openai:qwen38-flash-next image for arm!
Couldn’t get FP8 cache working yet, needs some vLLM patching (or maybe more).
That’s just a first successful launch, not yet refined. With MTP=3 and max context size 262K we get 365K KV-cache (BF16), and a very shallow llama-bench shows this:
| model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) |
|---|---|---|---|---|---|---|
| Qwen/Qwen3.8-Flash-Next-FP8 | pp2000 | 2077.78 ± 50.02 | 967.84 ± 22.99 | 963.76 ± 22.99 | 967.84 ± 22.99 | |
| Qwen/Qwen3.8-Flash-Next-FP8 | tg64 | 37.57 ± 5.02 | 45.00 ± 5.10 |
Trying it right now in a benchmark of mine, going around 40-45 tps. Looking good.
Setup: 2× DGX Spark (GB10, 128 GiB unified memory each), one GPU per node. Tensor parallelism across both nodes (TP=2 is the minimum validated configuration and is the only multi-GPU option when each node has a single GPU). Use the official vLLM image for this model (vllm/vllm-openai:qwen38-flash-next, vLLM ≥ 0.28) — PyPI installs and other forks are not compatible. The FP8 checkpoint is ~173 GiB; each node ends up with ~88 GiB of weights resident.
Result once working: 262,144-token native context, MTP3 speculative decoding, TP2. Expect a ~20-25 min boot: checkpoint streaming + a one-time kernel JIT pass.
Issues you will hit, in order, and what to do:
Container dies instantly with unrecognized arguments: infinity. The official image sets ENTRYPOINT=["vllm","serve"] with no default CMD. If your launcher keeps containers alive with docker run ... IMAGE sleep infinity (a common pattern), it actually runs vllm serve sleep infinity — sleep is eaten as the model name, infinity is an unknown flag, PID 1 exits and takes the cluster down. Fix: override the entrypoint for the keep-alive (--entrypoint sleep IMAGE infinity). Images with a pass-through entrypoint don’t have this problem.
VLLM_PLE_CPU_OFFLOAD=1 is rejected: ValueError: ... Unsupported settings: nnodes=2. The N-gram embedding CPU-offload path is single-node only, and with one GPU per GB10, TP2 always means 2 nodes. The 51B N-gram table shard (~25.5 GiB/node at FP8) must live in GPU memory. Don’t set it.
--kv-cache-dtype fp8 is rejected: NotImplementedError: Qwen3.8-Flash-Next QSA requires a BF16 main KV cache. The QSA attention implementation checks the global cache dtype, so even --kv-cache-dtype-skip-layers cannot work around it. Leave KV in BF16. Impact: only 1 in 4 layers (QSA) accumulate real KV — the Gated DeltaNet layers hold a constant ~72 MiB — so a BF16 cache roughly halves your token capacity versus FP8.
Memory is the real constraint. ~88 GiB of weights per node leaves little headroom. The startup check requires “free memory ≥ gpu_memory_utilization × total” at init, and the desktop environment already consumes ~12 GiB — so 0.90 fails and even 0.89 fails with anything else on the GPU. Working value: 0.82 → ~365K KV tokens (1.39 concurrent requests at full 262K context). Nothing else must be running on the GPUs during bring-up. If you need more context headroom, lower --max-model-len rather than pushing utilization.
Default torch.compile wedges both machines on cold boot. The default compilation mode (VLLM_COMPILE/Inductor) generates so much code and allocator churn that both nodes run out of memory and stop responding to SSH (one may even drop ICMP). Only a power-cycle recovers them. Workaround: run --enforce-eager --compilation-config '{"mode":0}' — fully eager, no cudagraphs. Functional and stable, just slower. If you re-enable compilation later, persist the torch cache to a host mount so the storm only happens once — and be ready with a power-cycle plan for that first boot.
Don’t mistake normal boot phases for hangs: ~7 min of weight loading (172.78 GiB checkpoint; the second node streams slower and the head waits at the TP barrier), then ~10 min of Triton kernel JIT (single-threaded, CPU 40-95%, GPU idle — this is normal), then KV allocation and Application startup complete.
I have also tested on a single DGX Spark the Unsloth UD-Q4_K_XL: 111 GB (biggest that could fit).
Seems to get about 24 tk/s
I will give it some real tasks soon… the speed is a bit underwhelming, but I am sure that it is going to improve.
Using llama cpp or vLLM ? could you share your config ?
I’m using below :
/home/fossod/llama.cpp-qwen4exp-canary-20260826/build-gb10/bin/llama-server
–model /home/fossod/models/Qwen3.8-Flash-Next-GGUF-canary/UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf
–host 127.0.0.1
–port 9203
–alias qwen38-flash-next-canary
–ctx-size 8192
–n-gpu-layers 999
–parallel 1
–flash-attn on
–kv-unified
–mmap
–no-mmproj
–fit on
–fit-target 8192
–jinja
–metrics
llama.cpp 0.3.0-dev
build 10656
commit 035e22731
Really great to see it working decently so quickly after release!
Well, now we wait for the latest version to pass the Tool Evanch Bench test.
I’m sure it won’t score higher than 94 points.
Latest version of tool eval bench has brought down the baseline to under 90 for all the models I run, so if run the bench with the latest I’d expect anything more than 89 would be a positive result.
Take not just the latest version, but including the latest pr requests.
@oxbyte congrats! Trying to catch up. Followed your recipe, weights are humongous and taking time with download speed
Pulling weights now too. Image is here, recipe ready. 100 mb/s. Need to kick myself in a butt and get in touch with telecom, optics is here for almost a year..
BTW 35 t/s sounds like under optimized, it should be 50-60. Much improvements possible. If it can do 500-600k with yarn it becomes interesting
trying to start up @oxbyte 's recipe but running into couple errors with ngram weights. working with DSV40731 trying to start it. once ready will put write up here.
are you using the specific vllm image?
Yes — we’re using vllm/vllm-openai:qwen38-flash-next (the exact image from oxbyte’s post), confirmed version 0.1.dev20073+g8e685d198, on both nodes, TP=2
Key Findings
Throughput:
22 tok/s with NVMe PLE offload, Q4_K_XL (103.7 GB GGUF, 76.9 GB resident)
Prefill: 355-662 tok/s depending on context length
KV cache at f16 = only 24 KB/token (GDN architecture is incredibly efficient)
Full 262K token context = ~6 GB KV cache
Memory breakdown:
Model file: 103.7 GB, but only ~76.9 GB resident (51.2 GB PLE served from NVMe)
Process RSS: ~1.4 GB
Page cache: ~26 GB
KV cache (full 262K): ~6 GB
Total effective usage: ~110 GB in 128 GB — tight but fits
Speculative decoding results (n-gram modulation):
Reproduce a file with one change: 97.4 tok/s, 94.7% acceptance ✅
Targeted bug fix in a file: 68.6 tok/s, 81.4% acceptance ✅
Add a function to a file: 31.4 tok/s, 56.9% acceptance
Free-form prose: 22.1 tok/s, 5.8% acceptance
This is a game-changer for coding agents. When the output already appears in context (reproducing files, targeted bug fixes), n-gram speculation 4.4x amplifies throughput.
What’s being fixed right now:
–parallel 1 limitation → concurrency patch in progress
Graph reuse (canreuse-qwen4exp.patch) → +2.8% already landed
PLE quantization staging → rowband-ple-quant.patch solved
vLLM INT4 PLE-offload → merging soon
Stuck my working build into a docker image. Its the base vllm image and I just patched ple_layer as was having issues getting it running. All the hurdles overcome along the way listed here. Hopefully helps! Loving this model so far x00byte/Qwen3.8-Flash-Dual-Spark-Recipe: Working config for running the new qwen 3.8 flash next on 2 x gx10 / spark
ON A SINGLE DGX SPARK @ Context 256k and 45-90 t/s decode
Got it working thanks to 0xBakeer:
My recipe uses llama.cpp with some patches : full q4 quant, the auxillary table is offloaded to nvme:
@styles01 did you mean to write 19-22 t/s decode?