DeepSeek V4 cut prices twice in two days, with 75% off input/output and another 90% off cached inputs. The article body is blocked by WeChat verification, so the original pricing table, billing rules, context length, and V4 versus V4-Pro differences are not disclosed. I would not treat this as a full launch readout. Still, the disclosed numbers point to a clear move: DeepSeek is pricing for long-context, repeated-call, high-cache coding workloads.
QbitAI says its coding test used about 35 million tokens. The bill dropped from 31.73 yuan to 5.34 yuan, an 83% reduction. If that test is clean, the number matters. Coding agents do not spend like chatbots. They reread the same repo files, dependency maps, tool schemas, test logs, and error traces across many attempts. A cache hit rate moving toward 95% changes the cost of retries. The summary says V4-Pro reached roughly 95–96% cache hits, which sits near the ideal zone for prompt caching on stable codebase context.
My read is that DeepSeek is using price to force a product architecture choice. Teams building coding agents often still resend repo summaries, file chunks, tool definitions, and logs on every loop. A 90% cached-input discount tells them to stop treating context layout as plumbing. Stable system prompts, stable tool schemas, pinned repo indexes, dependency graphs, and deterministic file ordering now affect gross margin. For agent infrastructure teams, cache-key design and context segmentation are no longer backend niceties. They are pricing mechanics.
The competitive angle is sharp. Anthropic’s Claude Sonnet line has owned a lot of serious coding-agent mindshare, and I remember Sonnet 4.5 pricing sitting around $3 per million input tokens and $15 per million output tokens, though I have not rechecked the latest table. OpenAI’s GPT-5 family also leans on mini and nano tiers for cheaper volume calls. DeepSeek’s move feels more like the Chinese cloud playbook: do not win the story first, win the workload spreadsheet. The question it asks customers is not whether one benchmark moves by two points. It asks who can run the same code task 100 times without finance killing the rollout.
I have real reservations about the 83% figure. The disclosed test comes from QbitAI’s summary, not a public DeepSeek invoice in the body we can inspect. We do not get repo size, task type, number of turns, retry count, cache warmup method, or whether the same files were reused heavily. If the workload was built around stable repeated context, a 95% hit rate says less about messy daily development. In real enterprise repos, branches change, CI logs refresh, generated files move, and agents reorder snippets. Multi-agent systems make this worse. Tiny changes in tool schema versions or context assembly can break cache reuse. Since the article body does not disclose those conditions, I would treat the 83% drop as an upper-bound case for cache-friendly workloads.
The word “permanent” also deserves skepticism. China’s model vendors already ran brutal price cuts, and many promo prices later became normal prices. Sustaining them requires actual inference-cost improvements. If DeepSeek can keep cached-input pricing at another 90% discount, it either has strong confidence in KV-cache reuse, batching, routing, and memory economics, or it is accepting margin pressure to pull developers over. Those are very different stories. The available text does not give enough evidence to choose between them.
For practitioners, the action item is plain: replay your own traces. If you run code review, test generation, migration tooling, documentation sync, or knowledge-base agents with stable repeated context, DeepSeek V4’s new pricing can change unit economics now. If your workload is one-off reasoning, fresh long-form queries, or low-reuse tool calls, the headline discount will not show up the same way. Do not benchmark the discount from the marketing number. Measure stable cache hit rate across real retries. If it cannot stay above 90%, the bill will not fall by 83%.