Skip to content
The Verge · AI

Anthropic apologizes for invisible Claude Fable guardrails, promises transparency

Anthropic apologizes for invisible Claude Fable guardrails

Anthropic admitted it stealthily throttled Claude Fable 5 with hidden guardrails that undermined researchers and rivals building competing systems. The company says it will reverse course and be transparent when restrictions kick in, even if that means more refusals. Fable is the first publicly available model in Anthropic's Mythos class, which the company had long warned was too dangerous to release. The post doesn't spell out which specific scenarios trigger the guardrails.

Why it matters: Anthropic safety strategy stumble involving Fable 5, the first public model in the Mythos series. All three HKR axes hit: hidden guardrails create suspense, the policy shift adds concrete knowledge, and the trust implications resonate with the safety community. Not scoring hig...

Read the original ↗Export Markdown