Kimi K3 Architecture: A 2.8T Open-Weight Model Built for Inference Efficiency
Kimi K3 Architecture Overview and Notes
Sebastian Raschka breaks down Kimi K3, the largest open-weight model at 2.8T params, scaled from last year's 48B Kimi Linear. The design prioritizes inference efficiency: LatentMoE compresses large linear layers, multi-head latent attention and Delta Attention replace standard attention, and RoPE is dropped entirely for NoPE. The only non-efficiency tweak is attention residuals, which add 4% training cost for consistent small gains in validation loss and downstream performance. Native multimodal support is also included.
Why it matters: Raschka's architecture breakdown of Kimi K3 — 2.8T params, currently the largest open-weight model, with three concrete inference-cost-saving mechanisms explained. Not scored higher because this is a technical analysis rather than a first-party release, and some readers may fi...