● 1x concurrency:
┌──────────────────────────────────────────┬────────┬──────────────────┬──────────────┬─────────────────┬─────────────────┬─────────────────┐
│ model │ test │ t/s │ peak t/s │ ttfr (ms) │ est_ppt (ms) │ e2e_ttft (ms) │
├──────────────────────────────────────────┼────────┼──────────────────┼──────────────┼─────────────────┼─────────────────┼─────────────────┤
│ nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4 │ pp2048 │ 3183.54 ± 521.83 │ │ 668.34 ± 143.80 │ 666.86 ± 143.80 │ 668.42 ± 143.80 │
├──────────────────────────────────────────┼────────┼──────────────────┼──────────────┼─────────────────┼─────────────────┼─────────────────┤
│ nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4 │ tg32 │ 37.00 ± 0.10 │ 38.20 ± 0.10 │ │ │ │
└──────────────────────────────────────────┴────────┴──────────────────┴──────────────┴─────────────────┴─────────────────┴─────────────────┘
2x concurrency:
┌──────────────────────────────────────────┬─────────────┬─────────────────┬──────────────────┬──────────────┬────────────────┬─────────────────┬─────────────────┬─────────────────┐
│ model │ test │ t/s (total) │ t/s (req) │ peak t/s │ peak t/s (req) │ ttfr (ms) │ est_ppt (ms) │ e2e_ttft (ms) │
├──────────────────────────────────────────┼─────────────┼─────────────────┼──────────────────┼──────────────┼────────────────┼─────────────────┼─────────────────┼─────────────────┤
│ nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4 │ pp2048 (c2) │ 3736.79 ± 27.41 │ 1982.94 ± 113.32 │ │ │ 1037.38 ± 59.17 │ 1036.14 ± 59.17 │ 1037.45 ± 59.16 │
├──────────────────────────────────────────┼─────────────┼─────────────────┼──────────────────┼──────────────┼────────────────┼─────────────────┼─────────────────┼─────────────────┤
│ nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4 │ tg32 (c2) │ 63.47 ± 0.39 │ 34.30 ± 1.77 │ 65.51 ± 0.41 │ 35.41 ± 1.82 │ │ │ │
└──────────────────────────────────────────┴─────────────┴─────────────────┴──────────────────┴──────────────┴────────────────┴─────────────────┴─────────────────┴─────────────────┘
4x concurrency:
┌──────────────────────────────────────────┬─────────────┬─────────────────┬──────────────────┬───────────────┬────────────────┬──────────────────┬──────────────────┬──────────────────┐
│ model │ test │ t/s (total) │ t/s (req) │ peak t/s │ peak t/s (req) │ ttfr (ms) │ est_ppt (ms) │ e2e_ttft (ms) │
├──────────────────────────────────────────┼─────────────┼─────────────────┼──────────────────┼───────────────┼────────────────┼──────────────────┼──────────────────┼──────────────────┤
│ nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4 │ pp2048 (c4) │ 3588.59 ± 42.63 │ 1384.37 ± 432.26 │ │ │ 1628.74 ± 485.91 │ 1627.42 ± 485.91 │ 1628.78 ± 485.89 │
├──────────────────────────────────────────┼─────────────┼─────────────────┼──────────────────┼───────────────┼────────────────┼──────────────────┼──────────────────┼──────────────────┤
│ nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4 │ tg32 (c4) │ 53.88 ± 1.85 │ 20.59 ± 5.93 │ 114.40 ± 1.20 │ 28.80 ± 1.17 │ │ │ │
└──────────────────────────────────────────┴─────────────┴─────────────────┴──────────────────┴───────────────┴────────────────┴──────────────────┴──────────────────┴──────────────────┘
Exciting work, and hot thread!
So I’m getting mixed signals from the responses in this thread.
I’m seeing few reports along the lines of ‘only getting X tokens/sec’ but unsure if that’s about the same, or too little of an improvement compared to vanilla/eugr vLLM.
Further, it’s also unclear if this means more NVFP4 models are able to run with Avarok vLLM due to compatibility. I think it’s useful to 1) compile a list of LLMs to determine which NVFP4 LLMs work, then 2) compare which LLMs are faster with Avarok vLLM.
I want to get started on 1), but I can’t seem to rack my head around this:
=== Starting Interactive Bash Shell ===
=== Starting Interactive Bash Shell ===
...
=== Starting Interactive Bash Shell ===
=== Starting Interactive Bash Shell ===
This is just using Docker compose, using avarok/dgx-vllm-nvfp4-kernel image.
This has been a very interesting thread to follow along. I have downloaded and run both the OP’s dgx-vllm and eugr’s excellent spark-vllm-docker. However, while I can see the memory advantage, I am not really seeing a real world performance improvements. I have a single DGX Spark and using a Qwen3-Coder-Next-FP8 recipe I am getting usable performance and accuracy with my existing Opencode toolchain to run documentation, integration test writing and codebase research.
Like many others I bought into the NVFP4 marketing hype when I purchase my Spark and experienced waves of buyers remorse as I realised it was much more bleeding edge than I was expecting. However I have experienced this before with other technology lifecycles and I am really appreciative of all the work this community is putting into fulfilling the promise of this device. I too think its a bit of an iphone like watershed moment, with the backing of multiple vendors, and the right software stack it really has potential.
So far I have tried
- GadflyII/Qwen3-Coder-Next-NVFP4
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4
- nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4
In an OpenWebUI chat window they all seem to work. From my personal experience with writing tasks Qwen3 Next 80B is a step up in sophistication with the quality of its responses. Not as good as Kimi K2.5 but usable for many tasks. The avorok Qwen3-Next-80B-A3B-Instruct-NVFP4 outlined here is giving me 50 ish tokens a second which feels performant enough for a self hosted model running on your desktop and consuming less than 100w.
However if I add --enable-auto-tool-choice --tool-call-parser <parser> to vllm and run simple tool calls with Opencode the “marginal” degradation in performance is noticeable. None of the models feel as responsive as Qwen3-Coder-Next-FP8 and vary from error-prone to unusable.
Reading comments like the following leaves me all the more confused.
flash3 – If you are confused too - it seems that Avarok has solved a problem that does not exist anymore (software fallback for e2m1), only exists in older versions of CUTLASS and is “solved” in newer ones by just using cpu instruction sets that the compiler does not want you to use if you haven’t said the magic word “family”.
along with…
uger – Just in case you haven’t seen the other thread, this doesn’t require any special builds and works with any recent vLLM build, including our community Docker out of the box.
I guess my question is: Where do you think all of this heading? Will NVFP4 get more performant and accurate with time as the software stack catches up and the models are tuned to this quantisation – or is the accuracy likely to remain how it is and all we can expect is a speed improvement or just less memory overhead? I ask because this feels like a huge rabbit hole to go down and I have a setup that while not what I was expecting is adequate. This community clearly has a wealth of expertise – so what future improvements do you think we should be expecting from this work? What contributions are helpful from the less technical or newer members?
Thanks
So, that means it isn’t a port, it’s a rewrite of the memory pipeline from TMEM to SMEM, and go through the gencode versioning hell @flash3 described. Looks like native FP4 may not move the needle much over what Marlin already does and that raises the question of whether it’s worth doing at all. Fortunately, the community has produced more visible diagnostic work on this specific problem than NVIDIA has publicly yet.
Right there with you,
I’m hoping this turns out to be something other than vaporware: [BUG] [Blackwell] Enable FP4/tcgen05 support for sm_121 (DGX Spark) in CuTe DSL · Issue #2947 · NVIDIA/cutlass · GitHub
ok, I’m seeing slightly lower performance with a fresh build:
./launch-cluster.sh -t vllm-node-20260222-2 --solo -e VLLM_NVFP4_GEMM_BACKEND=marlin -e VLLM_TEST_FORCE_FP8_MARLIN=1 exec vllm serve nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4 --gpu-memory-utilization 0.7 --host 0.0.0.0 --port 8888 --max-model-len 128000 --load-format fastsafetensors
v0.16.0rc2.dev386+gc645e9a21.d2026022
llama-benchy --base-url "http://spark2:8888/v1" --model nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4 --pp 2048 --tg 32 --runs 5
| model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) |
|---|---|---|---|---|---|---|
| nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4 | pp2048 | 3891.17 ± 491.88 | 541.09 ± 78.25 | 536.17 ± 78.25 | 541.18 ± 78.24 | |
| nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4 | tg32 | 41.34 ± 1.33 | 42.68 ± 1.37 |
llama-benchy (0.3.1)
date: 2026-02-22 21:38:51 | latency mode: api
Frankly, I’m not sure. Given that FP8 quants work at speeds closer to expected on this platform, there is a hope that we can get good performance from FP4 as well. But consumer Blackwell (sm120 which has the same arch as our sm121 sparks) has been out for over a year now, and we still lack proper support. There are some efforts underway, including from NVidia, so we’ll see, I guess.
As for the accuracy, it depends on the quantization. W4A4 NVFP4 quants that are currently prevalent, I believe, have lower accuracy than AWQ. W4A16, in theory, should be slightly more accurate than AWQ quants, but a lot depends on how the quantization was done, and whether any calibration (and on what data!), or better, QAT was performed.
Right now, outside of a few models, like Nemotron (where it’s native), NVFP4 quants don’t offer any advantage over AWQ on Spark. Qwen3-Next series seem to also fall in this category as AWQ quants underperform compared to NVFP4 (when using Marlin) for some reason, which is not the case for most other models. But all 4-bit quants of Qwen3-Next also lose to FP8 quants, so it’s a weird edge case.
| model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) |
|:-----------------------------------------|-------:|----------------:|-------------:|--------------:|---------------:|----------------:|
| nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4 | pp2048 | 4340.68 ± 36.00 | | 473.50 ± 3.90 | 471.85 ± 3.90 | 473.59 ± 3.91 |
| nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4 | tg32 | 41.20 ± 0.04 | 42.54 ± 0.05 | | | |
I just unplugged the power adapter and waited for about 5 minutes before restarting the DGX SPARK. Now the performance is back to normal, and I didn’t rebuild the image during this process.
I suspect it might be related to the NVIDIA DGX Spark Field Diagnostics I ran yesterday.
So to paraphrase - it feels like you are describing a future patchwork of recipes where specific quants and settings achieve optimal performance for specific models – but the path to get there will remain somewhat agnostic? NVFP4 is not a solution, just one of many options to consider. This makes the 1 petaFLOP FP4 claim feel all the more illusory – maybe some model in some lab setting achieved it ¯\_(ツ)_/¯ but that doesn’t generalise to all models.
Beyond deep technical rabbit holes…
Think like a salesman: you have a dead horse. The horse was supposed to be something of your own, so you named it My-Horse.
Since people keep walking by and pointing at clearly visible problems (a kid asked “is it dead?”), you stop advertising the horse itself and instead promote its effect on the rider: rider rides really fast with this.
On the other side, you’re basically throwing every idea at the wall hoping something sticks. QAD, QAT, QTT… all things that are supposed to revive your dead horse. Who wants something dead — if not outright meat, then at least fresh or aged, but certainly not to slaughter yourself.
But in the meantime, a vegetarian fanbase has formed that genuinely believes the dead horse is alive and visits it every day. Nobody actually wants to ride it, but the idea is just so fascinating…
My-Horse makes you sooo fast.
[DGX Spark (SM121) Software Support is Severely Lacking - Official Roadmap Needed - #42 by flash3]
Well, 1 PF was a bit of a marketing gimmick and even carried an asterisk next to it saying that it is for sparse ops only.
Also, even with optimized FP4 kernels you won’t get inference (token generation) performance much higher than you can get using AWQ quants now (ignoring qwen3-next where AWQ is underperforming for some reason). That’s because inference is mostly memory bound, so no matter how much compute you throw at it, it will not be faster, at least for a single query. To get more on Spark, you need a cluster.
Where we will see performance gain is prompt processing speed (which is mostly compute bound) and high-concurrency scenarios where generation is batched, so GPU is more fully utilized.
Yes and no. If NVIDIA or community implement proper FP4 support in vLLM or SGLang, NVFP4 models could become a preferred quant. But what I said before still holds true, quant type is just one of the variables.
I mean, it’s not even limited to Spark - it is true for any hardware where you run quantized models. There will always be an optimal combination of quant type/settings for any given hardware type. One could hope for the sane default though.
PFLOP is count of ops per floatingpoint on a unit.
why floating point?
In the GPU years, there was a lot of cheating with benchmarks. Ultimately, floating point became the gold standard — somewhat harder to compute, more precise than INT, but also significantly more bits.
And this perceived advantage has now been taken to absurd levels by our dead horse salesman. 4-bit, the least precise on the market and initially not hardware-supported on DGX at all, now sort of.
The mantissas of NVFP4 don’t leave much room with so few bits (just one).
finally this is NVFP4 in its simplest form (very close approximation, without blockscale), taken from a post above. thats the dead horse.
the magic is in E4M3 that scales this fine enough.
this no real floating point. this is a mapping of 0-7 (3 bit) to a special profile 0..6. i can not believe that anyone would handle this with a floating point instruction.
this shows the dynamic behaviors. think about the idea of sparse modelling - to null any weight that is nearly null, on nvfp4 this could lead to a lot of nulling because models trained for nvfp4 can differentiate this “range” very well. but it may become rather a lobotomy?
in this “vegetarian dead horse non-riders hype” you need to know that everything begins with training. moonshot has successfully trained models in int4. so it depends. and it depends rather on proper software and framework support.
There’s a key difference between a dynamic FP8 quant and NVFP4:
FP8 (dynamic mode) can often be nearly negligibly different from the base model in nearly every way and does not require calibration data.
NVFP4 requires calibration data. You do lose something in the process, unless natively trained that way, but the result will be nearly transparent when guided by your calibration data. That’s why I recently suggested that we need good recipes for doing our own quants locally on our own data.
Think of dynamic FP8 like a generalist and NVFP4 like a specialist. When you download a NVFP4 quant which someone else did, you’re at the mercy of their calibration data for your perceived performance.
exactly.
you cut out what you don’t need.
The retrainer always loved Faust? so he expects the same from the reduced model. After quantization it can quote Goethe perfectly well, but fails at basic arithmetic.
and using the dead horse it fails even more and definitely slower.
I understand the concepts here very well. I designed FPGA serial multipliers in VHDL back in 1996. It still feels like a magic trick that they can get 4 bits to work so well. Even BitNet was kind-of crazy. I guess I bought into the idea because of the logarithmic scaling. It made it seem plausible, and knowing how hardware multipliers work you expect a speed-up as long as the scaling mechanism doesn’t eat into the cost saving - it should come out ahead. So in theory it sounds plausible to get a 1.5 x improvement on FP8 with “marginal” degradation. I just haven’t experienced that yet.
I had already seen the benchmarks when I got my Spark. I have a Mac and got the Spark because of how it could handle concurrent workloads - that is my area of interest. NVFP4 plays into my memory requirements of long running, highly concurrent workloads. I am not as interested in clusters right now. Nevertheless the models now are so much better than they where when I started my project a year ago. So I am super happy with everything thus far, I just thought I would have a larger memory budget - and kind-of built that into my assumptions. So I am just trying to figure out what is the most preferment and optimal solution and what the picture might look like 2/4/6 months from now.
There are those who would argue that the horse is just sleeping.
then give int4 autoround a try: [FP4 on DGX Spark — Why It Doesn't Scale Like You'd Expect - #125 by flash3]
its even faster than nvfp4.
barely FP :)
Yeah, I know, you don’t have to explain that to me :)
