Skip to content
AI HOT (Curated Pool)

Alibaba Qwen releases Qwen3.8-Flash, a 125B MoE model activating only 6B per token, as an early preview of the Qwen4 architecture

Qwen3.8-Flash 开源,Qwen4 架构预览

Alibaba Qwen open-sourced Qwen3.8-Flash with full weights. It's a multimodal MoE model with 125B total parameters, activating only 6B per token. Training cost is 1/9 of Qwen3.7-Plus while outperforming it across the board. Production API pricing is $0.16/1M input tokens and $0.47/1M output tokens, with 262K native context expandable to 1M. The model also serves as an early preview of the Qwen4 architecture.

Why it matters: Alibaba Qwen open-sources Qwen3.8-Flash, a 125B MoE model activating only 6B per inference, with 1/9 the training cost of its predecessor and claimed performance gains, plus a Qwen4 architecture preview. Domestic flagship release with concrete numbers — HKR all hit. Not 90+ be...

Read the original ↗Export Markdown