Alibaba's Qwen team released Qwen3.8-Flash-Next, an open-weight preview of the design it plans to use for Qwen4, which holds 125 billion parameters but uses only 6 billion for each piece of text it produces. The team frames the change as a way to cut the cost of running long agent jobs, and the performance figures come from its own testing. The weights carry a qwen-community licence that may not meet the EU AI Act's free and open-source exemption, since Alibaba has said it wants to charge its largest commercial users.
What changed
Alibaba's previous model, Qwen3.7-Plus, held 397B parameters and activated 17B per token.
What it unlocks
Running an Alibaba open-weight model that previews the Qwen4 design on roughly a third of the previous active compute, including on memory-limited accelerators.
- 125B parameters, 6B active per token
- 51B in a separate embedding
- prior model: 397B total, 17B active
What you need to act on it
- accepting the qwen-community licence
- self-hosting or a hosted provider
- legal review of the licence under the EU AI Act
Sources