DeepSeek disclosed three things in one breath: 1.6T total parameters, 49B active parameters, and an MIT license. That is enough to take the release seriously. The “close to Opus-4.6 Max and GPT-5.4 xHigh” line, though, is still headline theater until they publish benchmark names, scores, dates, and weights.
My first read is that the licensing claim matters more than the model-vs-model chest beating. If MIT is real and the weights actually ship under that license, DeepSeek is attacking a bottleneck that has blocked a lot of enterprise adoption over the last year: legal clarity. Plenty of “open” model families still come with custom terms, usage carve-outs, or enough ambiguity to slow procurement. A clean permissive license changes the conversation for distillation, fine-tuning, on-prem deployment, and internal derivatives. A lot of teams do not need the absolute best model. They need one legal can approve, infra can host, and finance can predict.
The 49B active figure also tells you this is still a MoE efficiency play, not just a giant-parameter flex. That part tracks with where the field has gone. Over the last year, several strong open families leaned into huge total parameter counts while keeping per-token compute in a deployable band. DeepSeek has been on that trajectory already. Qwen’s MoE line moved in a similar direction. I have not seen context length, throughput, KV-cache behavior, or price here, so I cannot say where V4 lands in actual serving economics. But 49B active is not cosmetic. It maps directly to inference cost.
I do have a pushback here. “Approximately equal” to Opus-4.6 Max and GPT-5.4 xHigh is only meaningful if the benchmark conditions match. Who ran the evals? Were these public benchmarks, internal suites, or curated tasks? Was tool use enabled? Single-shot or pass@k? Temperature and sampling settings alone can move results enough to make this comparison squishy. I’ve seen the same pattern in hardware claims for years: vendors advertise 10x, production users get 3x or 4x under real constraints. Model marketing works the same way. Without reproducible conditions, “close to” carries low information density.
So my stance is simple: high-interest release, low-confidence capability claim. MIT plus 49B active is the part that deserves attention. The top-line comparison does not, at least not yet. DeepSeek needs to publish the weights, the eval table, and the serving assumptions. Until then, this is a promising distribution move with unverified performance attached.