Alibaba’s Qwen team just teased Qwen4 today. Here’s what you need to know.
ModelScope posted a countdown for Qwen3.8-Flash-Next, an open-weight release built on the next-gen architecture that will power the upcoming Qwen4 model family. Alibaba says it’s shipping this architecture early so developers can prepare before Qwen4 itself lands.
Qwen3.8-Flash-Next is a multimodal MoE model. It pairs a large parameter count with a smaller active count per token, plus a huge n-gram embedding table for fast local token lookups, on top of GDN and QSA mechanisms.
Key numbers:
- 125B main-model parameters
- 51B additional N-gram embedding parameters
- 6B parameters active per token
- Open weights expected within about a day of the announcement
No official benchmark numbers or exact release date have been published yet, this is still a preview ahead of the full Qwen4 family.
I sure do hope it’s at Least 512K of context window.