Ant Ling-3.0-Flash 124B-A5B... new fast model for one Spark!

Ant group has just released a 124B-A5B model which beats their last 1T model across nearly all benchmarks…

Ling-3.0 starts with native hybrid-linear attention: KDA and MLA layers stacked 5:1.
KDA gives fine-grained control over long-range memory, while 1/64 expert activation makes MoE compute more efficient.
It supports 256K context natively and can scale to 1M.

They haven’t released weights yet or posted a schedule, but since they’re releasing these arch details we should assume soon. Seems to be a new tradition this year to wait 10 days to ship weights, so that’s a nominal assumption.

A5B means this should be much faster than our other ~100B options.

Looks promising indeed, and like you said, because of the low number of activated weights, it should be quite fast too.

It appears to benchmark close to DS4F, which is promising, but the new Laguna model also did and so far (perhaps still too early to say) that one looks benchmaxed.

Ling comes from Alibaba I understand, and typically at least the Qwen team (also Alibaba) did not appear to strongly benchmax, so the hope is that the weights are released soon and that real-world performance (quality) is indeed in line with something like DS4F.

As of now DeepSeek 4 flash is the best model for single DGX Spark, but it is slow. And a5b should surely be lightning fast.

They just confirmed open weights soon, probably after Aug 3

INT4 auto-round or NVFP4 of this model will make this a good new contender to the current single spark GOAT qwen3.5 122B

Weights have just dropped! inclusionAI/Ling-3.0-flash · Hugging Face

Looks like there’s an NVFP4 out olka-fi/Ling-3.0-flash-NVFP4 · Hugging Face

another good model to PrismaQuant it dont you think?

Just eyeballing the numbers, looks like ~70-90GB total for the model size? I wonder how it performs compared to Qwen 122B, and if we can actually run it single spark. Per the model card, it doesn’t seem like the NVFP4 version (… as usual) performs as well as it should given the supposedly superior quant over the 77GB smaller version.

Yes, A5B is the sweet spot, so maybe they HAVE been listening. It’s small enough to be fast and hopefully large enough to be intelligent.

Waiting for an INT4 Autoround vLLM receipt from @eugr_nv This maybe a good upgrade to current workhorse for single spark qwen3.5 122b

Why INT4 Autoround? My understanding is that NVFP4 gets a speed bump on blackwell hardware?
I tried to find some comparison benchmarks for existing models but didn’t find much.

Full instruction set for sm121 Nvfp4 does not match older set that are targeting DC gpus and despite native Nvfp4 current implementation fallback go software emulation affecting speed and as per many observations - quality

As far as I have been reading, nvfp4 does not work well on GB10 yet, as Orand pointed out

In my subjective experience, every NVFP4 model I have tried, as 0rand has pointed out, has had disappointing performance. Speed might be fine, but reasoning is oddly bad.

Interesting. I’m pretty new here, and new to Spark, but as far as I understood this had been sorted out in the last 6 months or so. But, I may be incorrect/misinformed. NVFP4 for Qwen 3.6 27b seems to work pretty well for me. But I’ll definitely try to find some INT4 Autoround quantizations to try out now.

Try running fp8 for comparison.

Downloading 4 bit for Mac. Down have much hopes but we will see. Hard to beat 27b fp8 and 35b fp8. If it’s better than Laguna s2.1 it might be interesting

Is that what this is?

They released the official INT4 an FP4 quants today:

Getting 50-80 tok/s decode and 2,500+ prefill on one Spark!