MiniMax-M3 open-sourced: a 428B MoE model with 23B activated parameters
MiniMax released MiniMax-M3 weights on Hugging Face. It's a mixture-of-experts model with ~428B total parameters and ~23B activated per inference. The post doesn't disclose training data, benchmarks, or minimum VRAM for local runs.
Why it matters: MiniMax dropped full weights for M3 on Hugging Face — a 428B MoE model activating only 23B per forward pass, putting it in the top tier of open-weight efficiency plays. No benchmarks or hardware requirements disclosed yet, which caps the score, but the weight release alone is ...