Skip to content
AI HOT (Curated Pool)

MiniMax open-sources M3 model weights (428B total, 23B active) with lower long-context cost

MiniMax 开源 M3 模型权重及 MSA 技术论文

MiniMax open-sourced M3 model weights last Friday—428B total parameters, 23B active—along with the MSA sparse attention paper that cuts long-context inference cost. M3 is the first open-source model trained with interleaved text and image data from the pre-training stage. Two weeks post-release, it ranked #1 among open-source models on the Artificial Analysis Intelligence Index and GDPval-AA, reached Pareto-optimal on Code Arena WebDev, and topped Chinese models on Vals.AI. Output speed improved from ~30 TPS to ~80 TPS, with another 30–40% planned. A usage dashboard was added to the Token Plan backend.

Why it matters: MiniMax open-sourced a 428B MoE model with interleaved image-text pretraining and two #1 open-source rankings in two weeks — enough signal for featured. Held back from p1 because the post is a first-party announcement without third-party benchmarks or concrete MSA cost numbers...

Read the original ↗Export Markdown