Skip to content
AI HOT (Curated Pool)

DeepSeek V4.1-Flash: 1M context, FP4 KV cache, and cross-layer attention reuse

DeepSeek AI 发布 DeepSeek-V4.1-Flash:1M 上下文、FP4 KV 缓存与跨层注意力复用

DeepSeek released V4.1-Flash, targeting long-context efficiency. It supports a 1M-token context window, uses FP4 KV cache to cut memory, and reuses attention across layers to reduce compute. The post does not disclose benchmark scores, parameter count, license, or API pricing—only the technical features are described.

Why it matters: DeepSeek drops V4.1-Flash with 1M context, FP4 KV cache, and cross-layer attention reuse — a concrete engineering combo that's worth a look. But no params, benchmarks, license, or pricing are disclosed, so we can't gauge real competitiveness. That gap keeps it at the featured ...

Read the original ↗Export Markdown