DeepSeek open-sources V4-Flash-Vision-Exp, its first vision model, with multimodal agent performance near Opus-4.8
DeepSeek-V4-Flash-Vision-Exp 模型已开源,多模态 Agent 能力接近 Opus-4.8
DeepSeek released V4-Flash-Vision-Exp on Hugging Face under MIT License—the first V4 model that accepts image inputs. The repo includes a minimal PyTorch inference implementation covering the vision encoder, MoE, DFlash Attention, and other core modules. It handles JPEG, PNG, GIF, and WebP for tasks like image captioning, screenshot OCR, and chart reading. Text-only performance matches the stable V4-Flash; multimodal agent benchmarks show a big jump, nearing Opus-4.8. This is an experimental version—it hit the API on Aug 21 and now has open weights.
Why it matters: DeepSeek's first multimodal V4 model, MIT-licensed, directly targeting Claude Opus-4.8 on agent tasks — a significant update from a major Chinese lab. Score held back because it's an experimental release and the post doesn't disclose specific benchmark numbers or comparison de...