DeepSeek released a V4 preview with Pro and Flash tiers. My read is simple: this will not hit like R1, but it will move the cost floor for agentic coding. V4-Pro costs $1.74 per million input tokens and $3.48 per million output tokens. V4-Flash is about $0.14 and $0.28. Both support a 1M-token context window. That package matters more than another benchmark table, because coding agents burn tokens on repos, logs, tool traces, tests, and patch iterations.
MIT Technology Review frames the release around open weights, long-context efficiency, and Chinese chips. The article discloses pricing, API access, web and app availability, and the split between V4-Pro for coding and complex agent tasks and V4-Flash for speed and price. The body we have cuts off during the benchmark section. It does not give complete scores, test sets, comparison models, or degradation curves at long context. So I would not buy the “rivals the best models” line yet. DeepSeek earned attention with R1, but company-run benchmarks are still company-run benchmarks. SWE-bench, Aider polyglot, LiveCodeBench, long-context retrieval, and repo-level bug fixing need independent reruns.
The pricing is the hard part already on the table. I remember Anthropic Claude Sonnet 4.5 landing around $3 per million input tokens and $15 per million output tokens, though I have not rechecked that exact figure. OpenAI’s higher-end lines have also kept long context and tool use in premium territory. V4-Pro at $1.74 and $3.48 is aggressive, especially on output. Coding agents produce a lot of output: plans, diffs, test logs, explanations, retries. Dropping output from the $15 range to $3.48 changes the economics of batch repo work. Flash at $0.14 and $0.28 reads like a sub-agent model for summarization, retrieval compression, test explanation, and first-pass review. That pattern already shows up in Cursor-style and Devin-style stacks: a stronger model handles hard decisions, cheaper models chew through scaffolding.
The 1M-token context window also needs a colder read. The hard part is not accepting 1M tokens. The hard part is using them. Gemini 1.5 Pro made million-token context a public product story long before this, and OpenAI and Anthropic have pushed high-context tiers too. Developers pay for effective context, not advertised context. If V4 lowers attention cost but loses cross-file constraints after 600K tokens, its coding-agent value drops fast. The article says DeepSeek uses a new design for efficient large-text handling. It does not disclose whether that means MLA, sparse attention, block memory, or another mechanism. That missing detail matters because open weights become much more valuable when the community can optimize the inference stack around the architecture.
I am also careful with the “open source” wording. MIT says the model is open source and available to download, use, and modify. In AI, that phrase often blurs four separate things: open weights, open training code, open data, and permissive commercial licensing. R1’s impact came from weight availability and method diffusion, not a fully open training pipeline. If V4 only ships weights and a technical report, enterprises still need to inspect the license, compliance posture, data provenance, and distillation restrictions. The article mentions scrutiny from both the US and Chinese governments. Regulated buyers will not move sensitive workflows because input tokens cost fourteen cents per million.
The Chinese-chip angle needs even more caution. The headline says V4 is a win for Chinese chipmakers. The visible body only says R1 was trained with limited compute and that V4 is more efficient. It does not disclose V4’s training hardware, inference hardware, cluster size, HBM configuration, or domestic accelerator share. If the gain comes from architecture and inference engineering, it does help Huawei Ascend, Cambricon, Moore Threads, and similar ecosystems. Lower attention cost matters more when bandwidth and interconnect are constrained. But “helps the ecosystem” is not the same as “runs production 1M-context workloads on domestic chips today.” We are missing reproducible conditions: batch size, KV-cache strategy, throughput, latency, memory footprint. None of those are in the article body we received.
The part I care about is that DeepSeek is shifting the open-model fight from “can it reason?” to “can it remember a lot cheaply?” R1 pressured the reasoning premium. V4, if independent tests hold up, pressures the default use of closed APIs inside agent workflows. Internal codebases are the obvious wedge. Plenty of teams would rather self-host open weights than send entire repos, tickets, logs, and design docs to an external API.
Still, I would not declare another DeepSeek shock yet. R1 landed because two things were true at once: capability was close enough, and cost was dramatically lower. V4 has clearly shown the second condition. The first still depends on third-party evals and real agent traces. My test would be brutal but fair: give V4-Pro a million-token monorepo, a bug spanning 20 files, failing tests, and noisy logs. Make it patch, rerun, diagnose failures, and preserve style. If it gets near Claude Sonnet 4.5 while costing under one-third as much, closed coding-agent pricing has a real problem.