Kimi Linear: A Hybrid Linear Attention That Beats Full Attention
Kimi Linear: An Expressive, Efficient Attention Architecture
Moonshot AI's Kimi team released a tech report on Kimi Linear, a hybrid linear attention architecture. Its core, Kimi Delta Attention (KDA), extends Gated DeltaNet with finer-grained gating to use limited RNN memory more effectively. They trained a 3B-active, 48B-total MoE model mixing KDA and MLA layers. Under the same recipe, it outperforms pure MLA across all benchmarks, cuts KV cache by up to 75%, and boosts 1M-context decoding throughput 6x. The team open-sourced the KDA kernel, vLLM integration, and model checkpoints.
Why it matters: Moonshot AI drops an architecture-level tech report with a concrete hybrid linear attention mechanism and a 48B MoE model. Not scoring higher because it's an arxiv preprint with no product timeline — real-world impact depends on community reproduction and third-party benchmarks.