This one's worth opening because Moonshot wrote their tech report as a worked example of scaling under real hardware constraints, not a parameter flex. 2.78T total, 104.2B active, 93 layers, native 1M context—the numbers aren't the point. The point is how they made them run on actual GPU clusters.
On the sequence axis: 69 KDA layers handle cheap recurrent history propagation, 24 Gated MLA layers do expensive global correction at a fixed 3:1 ratio. This keeps the KV cache manageable. It's structurally similar to DeepSeek V4's HCA+CSA approach—both build a "near-detailed, far-coarse" information economy inside the model, just with different implementations.
For depth, Block AttnRes compresses 93 layers into 9 block-level addressing sources. They sacrificed fine-grained cross-layer reads within each 12-layer block to keep cross-device communication under control. On the MoE side, LatentMoE halves the transmission vector before all-to-all communication, and the saved bandwidth lets them activate more experts—hitting 16/896 routing.
Training signals: AgentENV uses physical verifiers instead of an LLM Judge. It doesn't care how smooth the model's output sounds—it checks whether the code passed unit tests, whether the database record actually changed. Dynamic harness swapping across Kimi Code, Claude Code, Codex, and others forces the model to learn tool intent rather than memorize tool names. That's a solid design choice for agent generalization.
Post-training uses MOPD with a domain × inference-effort 2D matrix: 9 teachers across three domains and three effort levels, distilled into one model. It's finer-grained than single-dimension domain splitting, but the report doesn't show performance gaps between the 9 teachers or ablation results for the final student—I'd hold off on judging how much this actually helps.
Deployment runs QAT throughout: MXFP4 for routed expert weights, MXFP8 for activations. They paid the quantization cost during training so the model learns to work in low precision, rather than forcing it at inference time. Practical.
Overall, the report's value isn't a single breakthrough. It's a detailed look at how to make scaling decisions when VRAM, bandwidth, communication, and latency are all screaming at you. If you're building MoE architectures or long-context systems, the sections on communication compression, load balancing, and sequence trade-offs are worth a close read.