Qwen3.8-Flash-Next

Alibaba’s Qwen team just teased Qwen4 today. Here’s what you need to know.

ModelScope posted a countdown for Qwen3.8-Flash-Next, an open-weight release built on the next-gen architecture that will power the upcoming Qwen4 model family. Alibaba says it’s shipping this architecture early so developers can prepare before Qwen4 itself lands.

Qwen3.8-Flash-Next is a multimodal MoE model. It pairs a large parameter count with a smaller active count per token, plus a huge n-gram embedding table for fast local token lookups, on top of GDN and QSA mechanisms.

Key numbers:

  • 125B main-model parameters
  • 51B additional N-gram embedding parameters
  • 6B parameters active per token
  • Open weights expected within about a day of the announcement

No official benchmark numbers or exact release date have been published yet, this is still a preview ahead of the full Qwen4 family.

I sure do hope it’s at Least 512K of context window.