Skip to content
Hacker News front pageallisdust

The AI Race Just Got Awkward

DeepSeek 公开的 KV cache 优化正被西方实验室采用:其 MLA 架构将缓存压缩约 15x,DeepSeek-V4.1-Flash 借助 CSA2、跨层缓存复用与 FP4 缓存把全局 KV cache 降至每 token 890 字节,相比 DeepSeek-V1 在长会话场景下缓存占用降低约 437x。

Read the original ↗Export Markdown