Anthropic shipped Claude Opus 4.7 to the top of one leaderboard within 48 hours, and straight into 400 errors for existing workflows. My read is blunt: this was not a simple capability upgrade. It was a rewrite of the product’s objective function. Anthropic appears to have pushed Opus away from “high-end assistant that feels smooth” and toward “more literal, more refusal-prone, more optimized for knowledge-work evaluation.” The numbers in the article support that interpretation. GDPval-AA is cited at 1753 Elo, 79 points above No.2, and hallucination rate is said to drop to 36%. The backlash also has hard evidence behind it. NYT Connections Extended falls from 94.7% on 4.6 to 41.0% on 4.7. MRCR v2 reportedly drops from 78.3% to 32.2%. Same text can consume up to 1.35x tokens. Old thinking parameters can throw 400s. That is not a minor “vibes changed” story. That is a product contract changing under load.
What bothers me most is not that some scores went up and others went down. It’s that Anthropic bundled several painful changes into one release. The article names three of them clearly: tokenizer inflation of 1.0-1.35x on the same text, breaking changes from thinking={enabled,budget_tokens} to adaptive/effort, and hidden thinking by default while billing still applies. You can defend each change in isolation. Tokenization changes happen. Reasoning controls evolve. Logging defaults get tightened. But power users do not buy abstract intelligence. They buy predictability. If the unit price stays flat while the bill rises, if the model name changes by one decimal while behavior boundaries shift, and if observability gets worse while spend stays the same, trust takes the hit.
There’s industry context the piece doesn’t really unpack. OpenAI already ran into a version of this during the GPT-4 Turbo “faster but dumber” episode. One lesson many applied teams took from that mess was not “never upgrade,” but “never let the default model roll into production silently.” I’m going from memory here, but over the last year a lot of coding-agent teams started pinning model versions and running regression suites before routing traffic. Products like Cursor, Windsurf, and Aider are sensitive to model drift for a simple reason: code is not chat. Small shifts in refusal policy, planning depth, or literalness can wreck patch quality, test pass rates, and edit locality. Anthropic’s problem here does not look like a model suddenly got dumb across the board. It looks more like the reward was pulled hard toward enterprise-safe and benchmark-safe behavior, then released as if existing users would absorb it.
I also have some doubts about how much weight to put on GDPval-AA in this particular debate. A 1753 Elo and a 79-point lead are impressive on paper. But that benchmark measures knowledge work across 44 occupations and 9 industries. That matters, especially for a vendor selling expensive API access into enterprise workflows. It does not automatically map to your refactor task, your debugging loop, your long-chain reasoning prompt, or your agent harness. The article says hallucination improved by 25 points largely because the model declines to answer more often. That tends to help in enterprise QA and compliance-heavy settings. In day-to-day collaboration, users often experience the same behavior as stubbornness, extra turns, or brittle literalism. That does not mean users are irrational. It means the benchmark is not capturing workflow friction.
I’m even less convinced by the article’s softer explanation for the “combative” reports. Saying Opus 4.7 is just more literal only explains part of it. Literal execution can change tone and refusal patterns. It does not naturally explain a collapse from 94.7% to 41.0% on NYT Connections Extended. That scale of drop smells like a deeper issue in reasoning-budget allocation, search strategy, or tokenizer interaction with the reasoning stack. The body does not provide the system-card ablations needed to sort that out. I haven’t verified whether Anthropic has published a detailed breakdown of how high-reasoning mode differs internally between 4.6 and 4.7. Without that, blaming prompt ambiguity is too convenient.
I also don’t fully buy the “reputation collapsed in 48 hours” framing. Reddit backlash, a few prominent reversions, and angry power users tell you something real about the top end of the user base. They do not prove the market has rejected the model outright. Anthropic has recovered from rough releases before by fixing bugs, improving docs, and adjusting rate limits. But the tension exposed here is real and bigger than one release. Labs increasingly ship each model version as a multi-objective optimization package. Users increasingly treat those models as stable dependencies. Those two clocks are no longer aligned.
So for me, the important question is not whether Opus 4.7 is “better” or “worse” in some universal sense. It’s that Anthropic has moved Claude’s product definition. It seems more willing to trade away smoothness on some developer tasks for stronger scores in broader knowledge-work evaluation and safer non-answer behavior, while pushing migration cost onto users. If Anthropic does not add a better compatibility layer after this—longer support for 4.6, tokenizer cost estimators, clearer reasoning controls, stronger migration tooling—developers will learn the obvious lesson: never trust the default upgrade, always rerun the regression set. For a company that wants enterprise and agent workloads, that lesson is expensive.