Anthropic kept Opus 4.7 at $5/$25 per million tokens and pushed Claude Code’s default effort to xhigh. My read is blunt: this is less a glory launch than a product-line cleanup. They are trying to make the default Claude experience feel stronger and more reliable without forcing customers to reprice their mental model of Opus.
The article points to three concrete signals. First, 4.7-low through 4.7-high each beat the corresponding higher 4.6 tiers. Second, SWE-Bench Pro goes up by 11 points. Third, vision input jumps to 2,576 pixels on the long edge, about 3.75 MP. The most important part is not “better everywhere.” It is that Anthropic did this without raising list price. When a frontier vendor improves a top-tier model and leaves price untouched, that usually means one of two things: inference efficiency improved enough to absorb the cost, or the company cares more about adoption and retention than immediate ASP expansion. I lean toward the second. Setting xhigh as the Claude Code default says they want better out-of-the-box outcomes more than they want to meter every extra unit of deliberation.
This fits Anthropic’s pattern over the last year. Their strongest product moves have usually not been about the loudest benchmark win. They have been about reliability in actual workflows: instruction following, lower weirdness over long sessions, steadier tool use, less babysitting. I have not checked the current full pricing table, but Sonnet has generally sat far below Opus for cost, and Anthropic has long understood the trade: developers stay when the default mode is dependable, not when the leaderboard gain is two extra points. Making xhigh the default in Claude Code is product strategy, not just model strategy. It turns “more thinking budget” into the normal path instead of a hidden power-user toggle.
I do have some doubts about the tokenizer story in the piece. The claim is that the same input can consume up to 35% more tokens, yet total token use can still fall by up to 50% because reasoning is more efficient. That is not internally inconsistent. If the model takes fewer retries, fewer steps, or fewer back-and-forth loops, total cost can go down even if raw input tokenization gets worse. The problem is that the article does not give the reproduction conditions. Was this measured on single-shot coding tasks or multi-step agent loops? Inside Claude Code with tool calls, or plain API usage? Were cache hits included? Those details change the economics completely. Every major lab has spent the last year selling some version of “more expensive tokens, lower total cost.” Sometimes that is true. Often it is only true on the lab’s preferred workload distribution.
The SWE-Bench Pro jump also needs context. An 11-point increase is substantial for a mature frontier model. Still, benchmark gains do not automatically translate into “better for your daily repo.” Claude Code lives or dies on long-horizon stability. Does it start flailing 20 minutes into a task? Does it patch tests into false greens? Does it quietly lose context inside tool chains? Anthropic’s launch framing around long-running tasks and self-verification is telling. They know the code-agent market has moved past “writes good functions.” Buyers care about unattended runtime now.
The vision bump may matter more than the headline suggests. A 2,576-pixel long edge and roughly 3.75 MP is a real step up for computer-use agents. The hard part in screen understanding is rarely broad semantic recognition. It is reading tiny text, resolving dense UI regions, and acting on pixel-level distinctions without forcing extra crop-and-zoom logic. Better native resolution reduces workflow scaffolding. OpenAI and Google have both spent the last year pushing vision toward agentic desktop use. Anthropic catching up here tells me they do not want Claude to remain “the coding model.” They want the higher-value mix of coding, desktop automation, and document-heavy knowledge work.
I also want to push back on the article’s title line, “literally one step better than 4.6 in every dimension.” That is good publishing language, not good technical language. The body itself admits that some key questions are still speculation: new tokenizer implications, whether this is a new base model, and how close it is to the rumored Mythos line. The title asserts a full-dimensional upgrade. The body does not disclose architecture changes, context-window changes, latency distributions, or evaluation breadth needed to support that claim. Without those, “better in every dimension” really means “better on the dimensions Anthropic chose to show.”
My broader take is that Opus 4.7 looks like a convergence release. Anthropic is aligning capability, default configuration, and cost so Claude Code, API users, and multimodal agent flows can sit on a steadier top-end model. That is a pragmatic move, and honestly a smart one. If 4.7 were the company’s true big swing, I would expect a louder pricing change, a sharper naming break, or a more dramatic benchmark spread. The restraint is the signal. It smells like Anthropic is improving the model people already depend on while keeping runway open for something larger later. I have not verified the Mythos-adjacent chatter, so I would not overstate it. Still, this launch feels like staging, not finale.