DeepSeek disclosed two V4 variants with unusually explicit MoE sizing: Pro at 1.6T total parameters with 49B active, Flash at 284B with 13B active, both pretrained on 32T tokens. My read is that the important move here is not raw scale. It is product segmentation. DeepSeek is translating backend MoE routing into front-end UX: Expert mode maps to Pro, Fast mode maps to Flash. That tells practitioners more than the headline parameter count, because it speaks to serving economics, latency tiers, and where they think user demand actually sits.
I’m not buying the benchmark line at face value. The post says V4 is on par with “Opus 4.6” on multiple evaluations, with stronger agent ability and world knowledge. But the snippet gives no benchmark names, no dates, no context length, no tool-use setup, no sampling policy, no price, and no throughput. Without those, “parity” is not reproducible. We have seen this pattern all year: vendors pick a favorable mix of coding, Chinese, agent, or long-context tests, then the deployed experience diverges on reliability, refusal behavior, tool-call failure rate, and cost per successful task. If DeepSeek’s full announcement includes a proper eval appendix or system card, that changes the picture. This snippet alone does not carry the claim.
The sizing is still revealing. A 49B-active Pro is not a tiny active model pretending to be frontier. A 13B-active Flash is clearly tuned for cheaper high-volume entry traffic. That mirrors where the market has already gone. Anthropic has long used Sonnet as the volume workhorse and Opus as the halo product. OpenAI has repeatedly pushed mini tiers to absorb broad API demand. DeepSeek making this split explicit suggests it wants more than leaderboard credibility. It wants a cleaner traffic architecture across free consumer usage, app sessions, and API workloads.
The “32T pretraining data” number also needs restraint. Big token counts stopped being self-explanatory a while ago. The post does not disclose deduping, synthetic data share, code mix, multilingual composition, or data cutoff. Over the last year, scaling results have depended more on data quality and curriculum than on raw token volume alone. If DeepSeek wants to claim stronger world knowledge, I want to see freshness, tail-domain coverage, and factual QA behavior, not just a 32T figure.
Same pushback on the “new attention mechanism.” The claim is that it reduces compute and memory demand. Fine, but under what condition? Training or inference? At 8k, 32k, or 128k context? On which hardware? Is it replacing standard attention or only applied in selected layers? This space is full of memory-saving claims that only hold at specific sequence lengths or batch sizes. Until DeepSeek publishes complexity details or deployment measurements, I’d treat this as a promising engineering hint, not a bankable advantage.
So my stance is pretty simple: this announcement matters because DeepSeek is turning model architecture into product routing. That is a mature move. But without eval details, pricing, and latency curves, V4 is still a strong statement, not a settled result.