Qwen3.8-Flash-Next open-sourced: 125B total, 6B activated, previewing Qwen4 architecture
Qwen3.8-Flash-Next 开源:Qwen4 架构早期预览
Qwen released Qwen3.8-Flash-Next weights as an early preview of the Qwen4 architecture. The model has 125B total parameters, activates only 6B per token, and carries an extra 51B N-gram embedding table that can be offloaded to host memory. Four architectural changes: attention uses Gated DeltaNet plus Qwen Sparse Attention for long-sequence compression and sparse block selection; residuals become four-branch gated residuals; embeddings add N-gram lookup for cheap capacity scaling; the optimizer switches to Muon. Training cost is roughly 1/9 of Qwen3.7-Plus, yet it scores higher on coding and office benchmarks. API pricing is $0.16 per million input tokens and $0.47 per million output tokens. Native context is 262K, extendable to 1M with YaRN. Take the scores with a grain of salt—they come from Qwen's own tech report; wait for community reproduction.
Why it matters: Early Qwen4 architecture preview: 125B total params, 6B activated, attention layers replaced with Gated DeltaNet plus sparse attention. Concrete new mechanisms. A flagship Chinese model architecture release with direct relevance for inference and open-source work. Score not hi...