Alibaba open-sources Qwen3.8-2.4T-A95B: 2.4T MoE, 95B active, native 256K context
阿里开放 Qwen3.8-2.4T-A95B 模型权重:2.4T MoE、激活 95B、原生 256K 上下文
Alibaba's Qwen team open-sourced its first Qwen-Max-level weights. Qwen3.8-2.4T-A95B has 2.4T total parameters with 95B active per token, native 262K context expandable to 1.01M tokens. It uses a 512-expert MoE, routing 10 experts plus one shared expert per token, and includes multi-token prediction training. The model targets coding, office tasks, research, and long-horizon agent workflows. Benchmarks against Opus 4.8, Fable 5, and GPT 5.6 Sol show mixed results, with top scores on PaperBench and IFBench among listed models. Post-training combines combinatorial environment scaling, a unified reward system, and an online data balancer to reduce gradient variance. The post does not disclose the open-source license or inference hardware requirements.
Why it matters: Alibaba's first full open release of a cloud-grade flagship — 2.4T total params, 95B activated, native 256K context — puts it in the top tier. Hits all three HKR axes and triggers the domestic flagship model positive signal. Held back from 90+ because we only have the announce...