Skip to content
AI HOT (Curated Pool)

Alibaba Qwen releases Qwen3.8-Flash: a 125B multimodal MoE activating only 6B per token, trained at 1/9 the cost of Qwen3.7-Plus

通义千问发布 Qwen3.8-Flash,多模态 MoE 模型,Qwen4 架构早期预览

Qwen3.8-Flash is an early preview of the Qwen4 architecture: 125B total params, only 6B active per token. Native context is 262K, extendable to 1M. Training cost is just 1/9 of Qwen3.7-Plus, with better coding and office-task performance. Weights are open. The post doesn't disclose specific benchmark scores or license details.

Why it matters: Alibaba Qwen drops Qwen3.8-Flash as an early Qwen4 architecture preview: 125B total params, 6B active, trained at 1/9 the cost of Qwen3.7-Plus. Weights are open. The efficiency numbers are concrete, but the post doesn't disclose specific benchmarks or the open-source license, ...

Read the original ↗Export Markdown