DeepSeek-V4.1-Flash: 552B MoE multimodal model with KV cache compression
DeepSeek-V4.1-Flash 发布:552B MoE 多模态模型主打 KV cache 压缩
DeepSeek released V4.1-Flash, a 552B MoE multimodal model. The key feature is KV cache compression, which cuts memory usage during long-context inference. The paper just hit arXiv and doesn't disclose compression ratios or benchmarks yet, but the title says 'Pushing the Limits' — this is about inference efficiency.