NVFP4 on DGX Spark / GB10 is broken. I bought 9 of these for this feature. Requesting NVIDIA's official roadmap and response

Sorry, OP, I know that you wanted a reply from NVIDIA, but I decided to chime in regardless, because you are probably not going to get any in this thread, and this is why:

You may like it or not, but no legal council worth their money would allow employees to respond to a request like this in a public forum.

Now, back to the topic. I can only comment on what I see from the outside, but I interact with various people in different related Open Source projects, including NVIDIA folks, and I know that they listen to the feedback on the forums and want to improve things. Yes, I wish the support was there from day 1, but the platform is definitely not abandoned, and better software support is in the works.

So, for your questions specifically:

  1. I don’t think it will ever reach parity with SM100 as it lacks dedicated TMEM, but I hope we can get closer in terms of raw performance. How close, I don’t know - I didn’t really have time to dig deeper into low-level CUDA stuff, even though it really interests me.

  2. I’ve seen the number somewhere in the technical publications, but maybe I’m imagining it. I’d like to see it too.

  3. As I said, they are working on it already (and community too). In particular, both vLLM and Flashinfer maintainers are very responsive to related requests coming directly from NVIDIA or even a few mortals like me.

  4. It’s always been available in CUDA technical documentation. SM121 is identical to SM120 except for unified memory and associated quirks.

  5. The B12x backend should improve performance of Nemotron-H models, hopefully. I haven’t had a chance to test it myself though.

The GB300 does not have a unified memory architecture like the DGX SPARK; instead, it features a structure where GPU VRAM and system RAM are separate, similar to traditional workstations. The GPU utilizes HBM, while the system RAM uses LPDDR5X.

Yup, after seeing the actual specs I agree. A lot of the sales pages are just calling it all “unified memory”, though, with some not even mentioning the separate GPU memory at all. Very confusing. But back onto topic! :-)

To be honest, I also think NVFP4 is a scam. At the very least, current-gen NVFP4 doesn’t seem to offer any advantages over other quantization methods. There might be a difference in high-concurrency data center serving, but on the DGX SPARK, it hasn’t provided any benefits in terms of speed, quality, or capacity when compared to formats like AutoRound or AWQ.

Initially, this fact made me quite angry. However, looking around, the DGX SPARK is still the best choice for those wanting to run LLMs locally. Browsing Reddit and various LLM communities, I always get the feeling that if I had bought a Mac Studio or a 395, I would have faced much more frustrating and aggravating situations. Even if Nvidia lied to us, they are still doing much better and providing faster support than similar solutions from other companies.

Of course, I fully agree that it would be much better if there were a clear roadmap for the DGX SPARK.

I don’t think that NVFP4 is a scam per-say, but I think it was vastly underestimated on what it would take to getting it up-and-running properly on SM121. As if we’re in a bizarre episode of The Twilight Zone, NVIDIA seems to be steadfastly disinterested in dedicating any resources to solving this.

I even publicly offered to buy Uncle Jensen a brand new fancy leather jacket if he would help stoke the pace of NVFP4 adoption on SM121… but sadly (and surprisingly) even that did not really move the needle.

NVFP4 is a VERY clever mathematical approach, and it should beat anything else out there, pound for pound. It’s really why I bought my Spark cluster… to run extremely large models. In retrospect, I probably should have just gotten another RTX PRO 6000.

On the other hand, I went in fully aware of the bandwidth limitation of Spark and have made peace with it as I’m not overly concerned with speed - I just want q8 quality in q4 size, so I can work with gigantic models.

Sad where we still are, STILL, but I’m fairly optimistic that the talented llama.cpp team will pick up the slack and get it working fairly soon.

It looks like a “VERY clever mathematical approach.” Indeed. Proof! Why are there so many attempts in quantization algorithms to handle this cleverness? Tell me how to handle outliers correctly — and is it important to differentiate that sharply near zero?

It’s a cool idea. OK. It’s an idea so different from normal quantization that you might become envious if your hardware doesn’t support it. Is it really necessary to have this hardware supported if it’s that clever? Or does this approach also need poor software support to show that only this hardware can do it — even though the hardware support doesn’t exist across all devices of the “clever” generation?

As long as you don’t question it, I agree. It’s so cool — the shiny object I never knew I needed to properly run AI.

Sighz… yes, that’s what came to mind the minute I saw that. Would be cheaper to just refund the box than to involve legal.. Funny thing is, you could probably sell the box for more than or close to what you paid.

I know that there is a recipe “book” , but would be great if it’s 1) Updated and 2) More details . I spend hours building containers to figure out that “no, the default options didn’t work but there is these other pytorch, cuda,patch, XML, etc versions that I could specify that would work or offer better performance on this model but would then bomb out on all other models”. Reading through all of the forums (which is fun except when you are trying to get things done).
Anyway, glad to see that there are so many members that’s taking their precious time out to help rest of us out.

your dreams may become true …

for USD 2999 only, while supplies last, only at selected olympus outlets

Form factor is wrong, I was expecting it to be in the form of a CX7 . But the issue isn’t the bandwidth, I bought it knowing the bandwidth… it’s the mainly the software…

i bought a golden jewelry box with a datacenter inside, is there software included?

I paying OPUS to fix the problems with the SPARK.

These are the key elements that influenced my decision to purchase the Nvidia DGX Spark. We purchased one as soon as they were available in Australia.

My personal experience of the first two weeks of ownership was terrible. NVFP4 was a horror show, slow, and the models sucked at doing anything useful. It took a while for me to realise I had to un-think what I had read and what I had expected from all that promotion and commentary. I felt rather foolish that I had been taken in by it all. I thought I had done adequate research prior to making my purchase.

1 Petaflop claim - without qualification

NVFP4 paper sold me on Gen 5 tensor core

Unboxing by a local AI engineer

Typical marketing hype not grounded in reality – as expected

Instead I had to find out what models and setups actually worked and what those setups could actually do. Once I did, with the help of people here, my experience got a lot better. The NVFP4 support is also improving, but still lagging slightly behind int4 in performance.

On balance I am glad I got the DGX Spark because of this community, and because it forced me to delve a lot deeper than I was expecting to need to. So we purchased another one.

A lot of people will purchase these GB10 system with very little understanding of what they are in for. If you are a new user, its not unreasonable to feel mad about this situation.

The people in this community are pure tech geeks, individuals who choose to believe in the possibilities of local and desktop computing. Even though NVIDIA hasn’t fully delivered on its promises, the posts everyone shares are very rational, simply asking NVIDIA to provide a roadmap. This little box has a certain power—it gathers a group of people who trust NVIDIA while also being pure, passionate, and creative. If it weren’t for everyone’s (especially eugr, albond, tenary) creativity and craftsmanship, which gave this little box the ability to run massive LLMs, I would have rather taken the loss and returned it (I’m in a large developing country without robust consumer protection laws). The best community, one that retains the spirit of exploration and sharing from the classic internet, is paired with a product from a mega-corporation whose capabilities don’t match their marketing. This is the sorrow of this era. What’s even more sorrowful is that its competitors, AMD and Apple, still seem to be asleep, leaving us with no alternatives.

OMG.

I have to come clean. I had the NVIDIA Shield tablet, and I use the Shield Pro (?) as a set-top box. And the support really is strong. Software updates. It runs. Top-notch. Really. Product design very good.

But honestly — I would never have gotten the idea to start hacking around inside a Shield (even though it’s doable), but on the other hand, taking this golden box seriously as a consumer device… Setup was excellent, no question. There are genuinely much, much worse products out there. It’s just the sizing…

If anything, I bought based on specs, hoping the numbers wouldn’t be as bad in reality as they looked (memory bandwidth would allow so much more, if only…). Yeah, with these kinds of new approaches, you’re jumping into the deep end a bit. But the divergence, and the absence of what was advertised…

Good thing I treated it as a challenge (see my other posts) and didn’t take NVFP4 seriously. And thanks to Intel. Did I mention Intel already? And thanks to the Österreicher MARLIN team too.

I think Janus will turn out to be right. The façade is crumbling. Simply too much, too great, too hip. And that’s just a shame. Maybe “shame” hits it pretty well.

A decode appliance is the answer. I don’t think it makes sense to attach it to the unit.

Add it to the list, and let’s make it the #1 viewed post and Maybe they will look at it.

bump

Official answer from NVIDIA on NVFP4:

You did an amazing job on this post. I have over 30 years in IT and I felt like a complete novice just trying to get this platform to do what was promised. My dual Spark system is doing OK but definitely not living up to expectations. Yet I remain hopeful that it might.

Keep fighting the good fight! I appreciate everyone’s efforts.

Cheers!

Been out of the forum for a couple of weeks and humbly apologize for not keeping up with it. I tried going through the threads and am still confused about it, so here goes: what’s the current status of NVFP4 support in the Spark? It is at least as fast as FP8 with lower memory footprint?