Anthropic kept Opus 4.7 at $5 per million input tokens and $25 per million output tokens, but it shifted migration cost into tokenization and longer completions. That is a disciplined launch, and a very Anthropic one: keep procurement friction low, keep the SKU simple, ship everywhere at once, and let developers discover the real bill in replay tests instead of in the pricing table.
I’m not fully buying the “same price” framing. The concrete changes disclosed here are narrow but consequential: vision now accepts images up to 2576 pixels on the long edge, and the new tokenizer can turn the same text into 1.0x to 1.35x as many tokens as before. The first change is easy to value. UI agents, chart extraction, screenshot-heavy workflows, and desktop automation all benefit immediately. The second one is where the launch gets slippery. Opus output is already expensive at $25 per million output tokens. If input tokenization expands and xhigh reasoning plus multi-turn use also lengthen outputs, “price unchanged” stops being a useful operating metric. Procurement sees list price. Finance sees monthly burn. Those are often different stories.
There’s a broader pattern here that isn’t in the snippet. Over the last year, Anthropic has generally sold predictability more than price aggression. Sonnet pricing stayed relatively stable from what I remember, and Opus has not chased the kind of hard repricing or routing churn that other vendors have used to redirect demand. OpenAI has repeatedly changed model aliases, defaults, and usage tiers. Google has leaned on Vertex bundling and context-window positioning. Anthropic’s pitch has been steadier: safer behavior, clearer enterprise posture, fewer surprises. A tokenizer change that introduces up to 35% variance for the same text dents that pitch, even if the API card still says $5/$25.
The Mythos Preview detail is the most revealing part. Anthropic is saying, in public, that it has a stronger model gated by cyber risk, and that Opus 4.7 is a reduced-risk version with added controls plus a Cyber Verification Program. That matters more than the benchmark teaser. It suggests Anthropic is no longer shipping purely by capability band; it is shipping by saleable risk profile. In other words, the frontier model is being split into productized slices based on what the company thinks it can safely expose. That is consistent with where agentic coding has gone: the same model improvements that make a system better at multi-step software work also make it better at navigating offensive cyber chains. Anthropic is drawing that line more explicitly than most peers.
I also have pushback on the coding claims as presented. “State-of-the-art” on a third-party eval does not tell me enough. The snippet does not disclose the benchmark protocol, tool permissions, sample size, or whether the task setup mirrors real software maintenance. We’ve already seen enough coding-model hype cycles to know that strong eval numbers do not automatically produce strong production ROI. SWE-bench gains, terminal-agent gains, and actual repair throughput inside a dev org are not interchangeable. The Claude Code updates are more concrete than the vague performance language. /ultrareview and wider access to auto mode change how teams use Claude: less as a copilot for line edits, more as a bounded worker for longer review loops.
The vision bump is the easiest improvement to take seriously. A 2576-pixel long edge, roughly 3.75 megapixels, is material for screen parsing. A lot of teams still crop, tile, or downsample screenshots before sending them to multimodal models. Higher native resolution can remove brittle preprocessing, which often matters more than small benchmark gains. But there is a cost side the article does not disclose: visual token accounting and latency. Bigger images are useful only if the completion economics still work. That part remains unspecified here.
My read is that Opus 4.7 looks less like a fresh capability leap and more like a low-friction renewal package for enterprise buyers. Stable list price, broad cloud availability, safer messaging, and a clean version bump all support that goal. Engineers should treat it more cautiously. Replay real traffic first: long code review sessions, screenshot-agent workflows, and multi-turn reasoning chains. Measure token expansion, average output length, latency, and task completion rate. If you look only at the unchanged $5/$25 number, you will probably understate the real migration cost.