Skip to content
Hacker News front page

Kimi K3 Architecture: A 2.8T Open-Weight Model Built for Inference Efficiency

Kimi K3 Architecture Overview and Notes

Sebastian Raschka breaks down Kimi K3, the largest open-weight model at 2.8T params, scaled from last year's 48B Kimi Linear. The design prioritizes inference efficiency: LatentMoE compresses large linear layers, multi-head latent attention and Delta Attention replace standard attention, and RoPE is dropped entirely for NoPE. The only non-efficiency tweak is attention residuals, which add 4% training cost for consistent small gains in validation loss and downstream performance. Native multimodal support is also included.

Why it matters: Raschka's architecture breakdown of Kimi K3 — 2.8T params, currently the largest open-weight model, with three concrete inference-cost-saving mechanisms explained. Not scored higher because this is a technical analysis rather than a first-party release, and some readers may fi...

Read the original ↗Export Markdown