Anthropic named 3 causes and attached dates, which is more useful than the usual “we’re investigating” fog. On Mar. 4 it dropped Claude Code’s default reasoning level from high to medium. On Mar. 26 a session-cache cleanup bug kept firing repeatedly. On Apr. 16 it added a 25-word / 100-word brevity rule to the system prompt. My read is blunt: this was not a harmless clarification that “the model stayed the same.” It was a product-engineering failure in the exact layer Anthropic claims to be good at: turning a strong base model, prompts, tool loops, evals, and production behavior into one coherent system.
The bigger issue is not any single bug. It is that all 3 changes made it through release process and reached users. Lowering default reasoning effort directly changes search depth on coding tasks. A cache cleanup bug that keeps re-triggering makes long sessions feel like repeated amnesia. The brevity instruction is the wild one: coding agents need room between tool calls to state plans, verify assumptions, and keep intermediate state coherent. If you compress that channel to 25 words, complex tasks will degrade fast. The article says the Claude API was unaffected, and that tracks. API users usually control the prompt, harness, tool loop, and recovery logic themselves. The damage clustered in Claude Code and the Agent SDK, which points to the product wrapper, not the foundation model.
This fits a pattern that showed up across coding agents through 2025: the benchmark looks fine, then the shipped product loses the plot on real work. The cause is often not weaker weights. It is default parameters, tool timeouts, context recovery, or prompt collisions that turn a model that can solve a task into one that feels flaky in practice. OpenAI’s CLI tooling, Cursor’s agent mode, and a lot of SWE-bench-adjacent products all ran into some version of this complaint over the last year: offline evals reproduce, long real sessions drift. I cannot say Anthropic is uniquely bad here. I can say the bar is higher for Anthropic because it has spent years selling rigor, safety discipline, and engineering control as part of the brand. If internal employees were not broadly using the same public build as external users, the dogfooding loop itself was distorted.
There is one part of Anthropic’s framing I do not buy at face value. “The model capability itself did not degrade” is technically plausible and product-wise evasive. For users, Claude Code is the model. Default reasoning effort, cache policy, system prompts, and Agent SDK orchestration are part of product capability. They are not detachable accessories. The scope here was also not narrow: Sonnet 4.6, Opus 4.6, and Opus 4.7 were all affected by parts of this chain. Once you get 3 separate changes that each materially reduce coding quality, the problem stops looking like random bad luck. It starts looking like intelligence quality was underweighted in release review, or the eval set did not cover real multi-turn engineering workflows.
Another detail matters. Anthropic says the changes hit different time windows and different traffic slices, so the aggregate degradation was broad but inconsistent and therefore harder to distinguish from normal feedback variance. Fair enough, up to a point. But that also exposes a gap in instrumentation. If you ship a coding agent, overall satisfaction is too blunt. You need metrics like long-session resume success, repeated-fix rate, plan consistency after tool calls, and task completion after an hour-plus idle period. The article does not disclose whether Anthropic tracked anything at that granularity. Without that layer, you end up relying on Hacker News and Reddit as your first serious alarm system, which is a rough look for a developer-tool company.
The industry context here is important. In agent products, competitive advantage increasingly lives in defaults and orchestration, not leaderboard deltas. Anthropic historically leaned hard on system-prompt control and constitutional-style guardrails. OpenAI has spent the last two years pushing reasoning budget, tool use, and memory strategy into product-level switches. At this stage, users often feel the difference not as “Claude scored 3 points higher than GPT,” but as “this agent is less likely to trip over its own defaults.” Anthropic’s promise to do stricter ablations on system prompt edits and use longer observation windows is sensible. It is also basic product engineering, not a bonus feature.
Resetting all subscriber usage limits on Apr. 23 was the right move. At least it acknowledges that people lost real work, not just forum sentiment. But compensation is not the interesting part. The repair mechanism is. Anthropic says more employees will use the public version directly and that it will improve internal code-review tooling, then ship that upgraded version to users. Good direction. What is missing is whether it will publish a more production-like eval protocol: tool use, long-context recovery, cross-day continuation, and model-specific prompt regression tests. Without that, the next cycle of “offline looks fine, users say it got dumber” is very easy to imagine.
Honestly, I do believe the base Claude models themselves probably did not suddenly get worse, because these 3 failure points are textbook product-layer failures. That does not make Anthropic look better. It makes the situation more revealing. If a leading coding model’s edge can be muted by a few prompt changes and one cache bug, then the coding-agent stack is still far less stable than the marketing suggests. Developers do not pay for benchmark purity. They pay for a system that keeps its head across long, messy, tool-heavy work.