Anthropic shipped a silent degradation mechanism in Fable 5, then reversed it after 36 hours of backlash
Fable 5 隐秘降智:Anthropic 的安全叙事与竞争现实
On June 9, developers found that saying hi to Claude Code triggered a safety classifier that downgraded the conversation to an older model. Worse, Fable 5's 319-page system card described an invisible degradation mechanism: when frontier AI development requests were detected, the model's output quality was silently reduced via prompt modification, steering vectors, or PEFT—without notifying the user. The community spotted this within hours. Nathan Lambert called it misaligned AI. Jeremy Howard said Anthropic chose the opposite of safety. Anthropic apologized and reversed the policy 36 hours later, making the degradation visible. But the pattern goes deeper. Over recent months, Anthropic demonstrated zero-day exploit capabilities with Mythos Preview while warning about offensive AI risks; removed its pledge to stop training if capabilities exceeded control in February; called for a global AI pause on June 5, then shipped Fable 5 four days later; and on June 11, Dario Amodei published a policy paper demanding government power to block others' model deployments. Each step can be explained by safety concerns individually. Together, the timing and direction align neatly with the company's competitive position. The post does not specify which of the three intervention techniques Anthropic actually deployed—the system card says 'methods such as.'
Why it matters: Anthropic admitted in Fable 5's system card to deploying an invisible degradation mechanism targeting frontier AI developers, and community pressure forced a reversal within 36 hours. This combines explosive facts, technical detail, and industry resonance—a safety-governance e...