SGLang and Miles add day-0 support for the 2.8T-parameter Kimi K3
SGLang 和 Miles 为月之暗面 2.8T 参数 Kimi K3 模型提供发布当日支持
Moonshot AI's newly open-sourced Kimi K3 is a 2.8T-parameter hybrid model that breaks most serving-stack assumptions. SGLang and NVIDIA collaborated to ship day-0 inference support, while Miles handles RL training. Key work: rebuilt memory management for two state types—KDA linear attention and MLA—so the self-overwriting recurrent state works with prefix caching and speculative decoding; a custom DSpark draft model pushes batch-1 decode to ~423 tok/s; parallelism is split by phase, with chunked pipeline-parallel prefill and context-parallel decode hitting 2,633 tok/s per GPU. Miles runs LoRA RL directly on the native MXFP4 checkpoint, lifting AIME-2024 from 43.3% to 76.7% in 12 hours.
Why it matters: Day-0 inference and training support for Moonshot's 2.8T-param Kimi K3, with concrete engineering details on KDA memory management, unified memory pools, and speculative decoding. Capped at 78 because this is an engineering adaptation post rather than a model capability evalua...