DeepSeek-V4.1 Flash: Pushing the Limits of KV Cache Compression
DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression
This technical report breaks down DeepSeek-V4.1 Flash's architecture, which compresses KV Cache by another 4x. The 552B-parameter model activates only 8B params during prefill and 16B during decode, using just 20 of its 40 layers for prefill. Compression tactics include cross-layer KV sharing (CSA2), FP4 KV Cache, and sparse attention indexer optimizations. The author argues the changes are so substantial it should be called DeepSeek-V5 Flash. The post does not disclose training cost or release timeline.