Skip to content
r/LocalLLaMA

MiniMax-M3 open-sourced: a 428B MoE model with 23B activated parameters

MiniMaxAI/MiniMax-M3 · Hugging Face

MiniMax released MiniMax-M3 weights on Hugging Face. It's a mixture-of-experts model with ~428B total parameters and ~23B activated per inference. The post doesn't disclose training data, benchmarks, or minimum VRAM for local runs.

Why it matters: MiniMax dropped full weights for M3 on Hugging Face — a 428B MoE model activating only 23B per forward pass, putting it in the top tier of open-weight efficiency plays. No benchmarks or hardware requirements disclosed yet, which caps the score, but the weight release alone is ...

Read the original ↗Export Markdown