Skip to content
Computing Life · Share · Yage

DeepSeek V4.1 Flash shifts the long-context cost battle from compute to memory

DeepSeek V4.1 Flash:算力优化触顶之后,长上下文的战争转向内存

DeepSeek released V4.1 Flash, compressing the global KV cache to about 1/4 and persistent KV cache to 1/8 of the previous generation, while cutting cache-hit input prices by roughly 60%. The tech report argues that sparse attention has already squeezed compute costs low; what now drags down long-running agent tasks is HBM filling up, SSD offloading, and bus transfers. Flash tackles this with 4-bit storage, cross-layer global-cache reuse, and dropping sliding-window disk writes, shifting the cost center from compute to the memory hierarchy. On deployment, DeepSeek initially planned to route all V4 Pro traffic to Flash immediately, but pushed the cutover to Sept 14 after developer pushback. The report also flags potential position-selection bias from layer reuse and degradation risks in extreme long-context cache reconstruction. All throughput and reduction figures are self-reported, not independently verified.

Why it matters: DeepSeek V4.1 Flash isn't a routine price cut — it compresses KV cache to 1/4–1/8 of the previous gen and slashes cache-hit input pricing by 60%. The tech report argues that sparse attention already tamed compute; the bottleneck is now VRAM and bus transfers. For agent builder...

Read the original ↗Export Markdown