Longcat-Next

So.. I did a thing. I saw this model, Longcat-Next, a 74b parameter MOE model with an estimated 3b active parameters. It’s a model clearly built on top of Longcat-Flash-Light, which I previously had gotten running quantized, before most serving methods supported it. I decided to do it again.

What makes this model special is, it can read text, decipher words from audio, and examine images, even from video. Which is pretty cool. Also it generates text, the spoken word, and images. Also cool. Oh, and voice cloning.

Much more complicated project… BUT I did it. it works, all the modalities, understanding and generation. Fits on one DGX Spark. Should work with other GB10 computers.

If you’re interested then give it a shot. Please let me know if you find any bugs to fix, or improvements to make.

5 Likes