Skip to content
Trending storyPast story

DeepSeek releases V4.1 Flash model with optimized long-context performance

13 reports5 sourcesupdated 15 days ago

What happened

From the coverage

DeepSeek 今天上线了 V4.1-Flash,这是他们全新模型架构里最小的一个,自带看图能力。新架构的设计目标是拉高能力天花板、跑得更快、吞吐量更大,后续会放大到更大参数的模型上。跑分方面,GPQA Diamond 90.9,Codeforces 3471,HLE 36.8,在编程和 Agent 任务上看着挺能打。API 这边,模型名改成 dee...

From AI HOT 精选

Coverage

Follow the reports to see the story from different sides.

Sep 12
  1. Latent SpacePick
    DeepSeek V4.1-Flash: a 763B encoder-decoder MoE with 8B prefill, 16B decode, and native vision

    DeepSeek dropped V4.1-Flash on Sep 10. Despite the 4.1 label, Sebastian Raschka called it a V5-level rewrite. It's a 763B total-parameter MoE with a causal encoder-decoder split: 8B active for prefill, 16B for decode, yielding 1–2% sparsity and up to 8× smaller KV cache vs V4 Flash. Native vision is built in, and V4 Pro has been quietly retired. The post doesn't include benchmark tables but argues current evals miss the point—the real advance is context efficiency for long-running agents.

  2. AI HOT (Curated Pool)Pick
    DeepSeek V4.1-Flash open-sourced: CED architecture cuts prefill cost for coding agents

    DeepSeek released open weights for V4.1-Flash, a 552B MoE model with a Causal Encoder-Decoder architecture tuned for coding agents. It splits compute asymmetrically: 8B active params during prefill, 16B during decode, plus improved KV cache efficiency. On Terminal Bench 2.1 it hits 90.6; on Automation-Bench it scores 54.8—better than V4-Pro but still failing roughly half of complex workflows, so keep a human in the loop. It is also DeepSeek's first non-experimental model with native image input. Chartography reaches 78.9, but ZeroBench logical reasoning over images is only 49. DeepSeek has already retired V4-Flash traffic and will reroute V4-Pro traffic to V4.1-Flash starting September 14.

Sep 11
  1. AI HOT (Curated Pool)
    DeepSeek V4.1 Flash Tested: Price Drop, Native Vision, Game & City Gen

    The article body is blocked by WeChat; only the title remains. It claims DeepSeek V4.1 Flash has a big price drop, native vision, and was tested on game and city generation tasks. The post does not disclose the exact price cut, vision specs, or generation quality.

  2. Computing Life · Share · YagePick
    DeepSeek V4.1 Flash shifts the long-context cost battle from compute to memory

    DeepSeek released V4.1 Flash, compressing the global KV cache to about 1/4 and persistent KV cache to 1/8 of the previous generation, while cutting cache-hit input prices by roughly 60%. The tech report argues that sparse attention has already squeezed compute costs low; what now drags down long-running agent tasks is HBM filling up, SSD offloading, and bus transfers. Flash tackles this with 4-bit storage, cross-layer global-cache reuse, and dropping sliding-window disk writes, shifting the cost center from compute to the memory hierarchy. On deployment, DeepSeek initially planned to route all V4 Pro traffic to Flash immediately, but pushed the cutover to Sept 14 after developer pushback. The report also flags potential position-selection bias from layer reuse and degradation risks in extreme long-context cache reconstruction. All throughput and reduction figures are self-reported, not independently verified.

  3. AI HOT (Curated Pool)Pick
    DeepSeek V4.1 Flash scores 40 on Intelligence Index, surpassing DeepSeek V4 Pro 0813 as new flagship

    Artificial Analysis reports DeepSeek V4.1 Flash hits 40 on the Intelligence Index, edging out DeepSeek V4 Pro 0813 as DeepSeek's top-scoring model. It uses 8B active parameters for input and 16B for output, supports 1M token context, and is MIT-licensed. The post doesn't disclose inference speed or pricing, so I'd hold off on cost-performance claims for now.

Sep 10
  1. AI HOT (Curated Pool)Pick
    DeepSeek-V4.1-Flash lands on SiliconFlow, a 552B MoE with 1M context window

    SiliconFlow launched DeepSeek-V4.1-Flash on Day 0. It's a 552B MoE model with ~8B active params during prefill and ~16B during decode, native vision, and a 1M-token context window. KV cache footprint is about 1/4 of V4 Flash, which helps with deployment cost. MIT license keeps commercial use straightforward.

  2. AI HOT (Curated Pool)Pick
    DeepSeek V4.1-Flash cuts KV cache memory for AI agents to a quarter of its predecessor

    DeepSeek released V4.1-Flash, a 552B-parameter model built to slash memory costs for AI agents. Its KV cache in fast GPU memory is about a quarter the size of V4-Flash, and the offloaded portion shrinks to roughly an eighth. The model splits into an encoder and decoder: only 8B parameters activate per token during input processing, versus 16B during text generation, nearly halving input compute. It supports 1M-token contexts and stores the main KV cache in FP4. On the DeepSWE v1.1 coding benchmark it scores 74.2%, narrowly beating Anthropic Opus 5 and OpenAI GPT-5.6 Sol, but it still trails on complex scientific tasks and image analysis. Weights are on Hugging Face under the MIT license. The post does not disclose inference latency or specific hardware requirements.

  3. AI HOT (Curated Pool)Pick
    DeepSeek V4.1-Flash drops with native vision and a big price cut

    DeepSeek released V4.1-Flash with a new Causal Encoder-Decoder architecture and native vision understanding — no separate vision model needed. It's a 552B MoE, activating 8B params for input and 16B for output. The post doesn't disclose the exact price cut or benchmark numbers, so I'd wait for third-party evals before getting excited.

  4. AI HOT (Curated Pool)Pick
    DeepSeek Releases V4.1-Flash: New Causal Encoder-Decoder Architecture with Native Vision

    DeepSeek V4.1-Flash is the smallest model in the new architecture family: a 552B MoE with 8B active params for input and 16B for output. It uses a Causal Encoder-Decoder design with native vision. KV cache drops to 1/4 of HBM and 1/8 of SSD storage vs the previous generation, and API pricing is lower. The post doesn't disclose exact pricing or vision benchmarks.

  5. AI HOT (Curated Pool)Pick
    DeepSeek V4.1-Flash: 1M context, FP4 KV cache, and cross-layer attention reuse

    DeepSeek released V4.1-Flash, targeting long-context efficiency. It supports a 1M-token context window, uses FP4 KV cache to cut memory, and reuses attention across layers to reduce compute. The post does not disclose benchmark scores, parameter count, license, or API pricing—only the technical features are described.

  6. r/LocalLLaMAPick
    DeepSeek V4.1 Flash: beats V4 Pro on benchmarks, cuts API price, and goes open source

    DeepSeek released V4.1 Flash, a 552B MoE model that activates only 8B params on input and 16B on output. It uses a new asymmetric Causal-Encoder-Decoder architecture and scores above DeepSeek V4 Pro on benchmarks. KV cache size drops to 1/4 HBM and 1/8 SSD vs the previous gen, cutting agent-scenario cache costs. The API is live under model name deepseek-flash; V4 Pro will be routed to V4.1 Flash from Sep 14 noon Beijing time and billed at Flash pricing. New peak/off-peak prices start Sep 10 noon, with off-peak at half rate. Weights and a tech report are open on HuggingFace; DeepSeek invites contact for large-scale deployments needing a 2k-GPU cluster.

  7. Hacker News front pagePick
    DeepSeek releases V4.1 Flash model on HuggingFace

    DeepSeek published a new model, V4.1 Flash, on HuggingFace. The post doesn't disclose parameters, benchmarks, or architecture details. HN discussion is active at 853 points and 481 comments, but most are speculating based on the name—Flash usually signals a faster, lighter variant. I'd wait for a technical note before drawing conclusions.

  8. AI HOT (Curated Pool)Pick
    DeepSeek releases V4.1-Flash, API pricing cut alongside

    DeepSeek launched V4.1-Flash today, the smallest model in a new architecture family with native multimodal vision. The new design targets higher ceiling, faster inference, and larger throughput, and is meant to scale to bigger models. V4.1-Flash scores 90.9 on GPQA Diamond, 3471 Codeforces rating, and 36.8 on HLE. Set model name to deepseek-flash in the API; old V4 Flash and V4 Flash Vision Exp are offline and requests are temporarily routed to V4.1-Flash. DeepSeek also claims V4.1-Flash beats V4 Pro on performance, cost, and speed, so V4 Pro requests will be routed to V4.1-Flash starting Sep 14 and billed at Flash rates. API pricing is cut, but the post doesn't list the new numbers—check the pricing page.