Skip to content
AI HOT (Curated Pool)

DeepSeek V4.1-Flash cuts KV cache memory for AI agents to a quarter of its predecessor

DeepSeek 发布 V4.1-Flash,大幅降低 AI Agent 的 KV cache 内存需求

DeepSeek released V4.1-Flash, a 552B-parameter model built to slash memory costs for AI agents. Its KV cache in fast GPU memory is about a quarter the size of V4-Flash, and the offloaded portion shrinks to roughly an eighth. The model splits into an encoder and decoder: only 8B parameters activate per token during input processing, versus 16B during text generation, nearly halving input compute. It supports 1M-token contexts and stores the main KV cache in FP4. On the DeepSWE v1.1 coding benchmark it scores 74.2%, narrowly beating Anthropic Opus 5 and OpenAI GPT-5.6 Sol, but it still trails on complex scientific tasks and image analysis. Weights are on Hugging Face under the MIT license. The post does not disclose inference latency or specific hardware requirements.

Why it matters: DeepSeek drops V4.1-Flash targeting agent memory costs — KV cache down to 1/4 of predecessor. Concrete architecture numbers, not vapor. Held at featured rather than p1 because only one source so far (no cross-source cluster yet) and the post doesn't disclose real latency/throu...

Read the original ↗Export Markdown