Key facts
- Alibaba's Qwen team is releasing Qwen 3.8-Flash-Next on Wednesday.
- The model has 125 billion total parameters, activating 6 billion per token.
- It is described as a preview of the Qwen 4 architecture.
- The model is multimodal and expected to use a mixture-of-experts design.
- The weights are not yet live on ModelScope or Hugging Face.
Alibaba's Qwen team is set to release Qwen 3.8-Flash-Next on Wednesday, a model with 125 billion total parameters that activates only 6 billion per token. The team has framed this release as a preview of the upcoming Qwen 4 architecture, rather than a finished flagship product. While official benchmark scores have not yet been published, the model is described as multimodal and built upon the Qwen 4 architecture. It is expected to employ a mixture-of-experts (MoE) design, a system where the network is divided into specialized sub-models, with only the relevant ones activated for specific tasks. This MoE approach allows a large model to operate with the computational cost of a much smaller one, making advanced capabilities accessible on more common hardware. Alibaba has released this early build to allow developers to prepare for the full Qwen 4 family. The weights for the model are not yet available on platforms like ModelScope or Hugging Face.
