2+ trillion, about to open weight. Maybe a small companion but not very likely imo.
i talked to reps in Europe and they dont want to release the 122B+ models that would fit perfectly for us… of course might be 35B MoE and 27B Dense, these would still be candidates to challenge DS4 Flash GA, Mimo v2.5 and Minimax M3.
Looking forward to it and maybe they will surprise us…
They will surprise us by putting on Huggingface only the model of 2T that no one will be able to load, I think it’s over the little models it doesn’t help them to do it…
I need a new 122b sooo bad :)
Like I said before, a mid-point would be great too, like 80B-A8B or something along those lines, that might be even better :)
small models these days coming from google, but optimized for phones :)
frontier now 1T+ parameters, so, why would they invest time to build small, if they can build big and someone will distill it to old horses.
New 27B dense is on the way, holding out hopes for a 35B-A3b MOE…
Agreed, seriously holding out hope for a 122B on 3.8, but some other mid quants would be nice too. I just seriously hope it is not just “benchmaxed”.
Currently, at least from my perspective, the 122B Parameter Model serves me perfectly. It really seems the sweet spot on a singe DGX. Perhaps we’ll get a 3.8 successor - who knows?
for now what ive seen only 27B dense is confirmed
I think they correctly address current ram/dram shortage and target most accessible size that performs well: 27b in int4 can fit 24gb cards and mid-tier apple (32-64gb), AMD lpdramm platforms
If 27B will be a huge step up in how well it performs vs its 3.6 version, there might not be a point in my Spark anymore, to be honest. My 5090 runs 27B with 200k context and pretty good performance on just a plain jane llama.cpp installation. I’m sure a vllm rendition would perform even better.
Currently 3.5 122B is still the best model I have found in general intelligence for my task. Laguna I’m mystified about because it’s like it performs incredibly well… once. Then never again. 35 A3B is okay but kind of pointless when 122B exists (as far as quality), and my experience with lower quant 731 versions of DSV4F have been mixed. Maybe the 731 version has been minmaxed for coding tasks.
Not sure what to do on a single spark anymore. That’s why I hope they release a 122B or some variant in between, otherwise it’s going to feel kind of unnecessary soon.
Best reason to own one spark is that it’s a half of 2 node cluster. That’s a game changer. I have m5 max 128gb but mostly use 27b and 35b there - better than 122b if you can plug a cloud model for tough questions, leaves more run for cache and many other apps.
BTW checked today gigabyte atop top 4T price, same as I own, same seller, +60% from what I paid for my second spark 3 months ago
Well at that point, I’m going to be around 7k deep into the platform to run however much better DSV4F is. My current spark cost me 3k and another one will cost 4k. I guess we’ll have to see. I’m still holding out hope for a 122B quant (or something between 35 and 122).

