Skip to content
QbitAI · WeChat

Fable 5's safety guardrails and anti-distillation triggers are far more aggressive than Anthropic's claimed 5% rate

Fable 5自带反蒸馏机制!检测到就降智,误触率高到离谱

Anthropic's newly released Fable 5 includes a safety classifier that silently switches sessions to the older Opus 4.8 model when it detects cybersecurity, bio, chem, or distillation-related prompts. Anthropic claims a sub-5% trigger rate, but users report false positives on routine coding, security audits, and even greetings. A separate anti-distillation mechanism degrades response quality without any notification when it suspects model-training intent—documented on pages 12 and 58–59 of the system card. Boris acknowledged the issue in comments and said the team is working on it. Fable 5 is free until June 22; token cost is roughly double that of Opus.

Why it matters: Anthropic's new model safety mechanism is backfiring with heavy false positives — a product incident with high discussion value. HKR all hit, but the source aggregates user reports rather than an official response, so it stays below 85.

Read the original ↗Export Markdown