Western labs adopt DeepSeek's KV cache optimization
What happened
On September 30, 2026, a Hacker News front-page report said KV cache optimizations published by DeepSeek are being adopted by Western labs. Its MLA architecture compresses the cache roughly 15x, the report said. DeepSeek-V4.1-Flash uses CSA2, cross-layer cache reuse and FP4 caching to cut the global KV cache to 890 bytes per token, about 437x less cache than DeepSeek-V1 in long sessions. The report also said Claude Opus 5.5 and GPT-6.1 Sol have quietly adopted the optimization.
Written by AI from the coverage · updated 55 minutes ago
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pageThe AI Race Just Got Awkward
DeepSeek 公开的 KV cache 优化正被西方实验室采用:其 MLA 架构将缓存压缩约 15x,DeepSeek-V4.1-Flash 借助 CSA2、跨层缓存复用与 FP4 缓存把全局 KV cache 降至每 token 890 字节,相比 DeepSeek-V1 在长会话场景下缓存占用降低约 437x。
Heat over time
Not enough continuous observations to draw a trend yet.