Skip to content
AI Chat-Group Daily (群聊日报)

Anthropic drops Claude Fable 5 / Mythos 5, hits 80.3% on SWE-bench Pro, but safety classifier misfires badly

2026-06-09 群聊日报

Anthropic launched two models: Fable 5 for everyone and the full Mythos 5 for trusted partners only. SWE-bench Pro hit 80.3%, well above Opus 4.8's 69.2% and GPT 5.5's 58.6%. It beat Pokémon FireRed using only screenshots. Pricing is double Opus 4.8 at $10/M input and $50/M output. Early testers burned through quota 2–3x faster than Opus; one user drained 73% of a 5-hour allowance in under two hours. The safety classifier became the day's biggest complaint—asking '9.9−9.11=?' triggered a downgrade, and writing an analysis of Anthropic's own safety report got the request blocked entirely. The article had to be finished by DeepSeek V4 Pro. One member pegged the $200 Coding Plan as roughly $5K–10K in API value, calling it a short-lived arbitrage. GitHub Copilot added Fable 5 the same day but requires dropping zero data retention, a dealbreaker for some enterprises. Anthropic's April advisor tool—where a cheap model calls an expensive one for advice—turns out to be the right cost fix for Fable 5. A rice-blast experiment in the safety report also surfaced a shift: AI is flattening domain expertise, but the people who can spot when its answers are wrong are becoming more valuable.

Why it matters: Anthropic flagship model launch with SWE-bench Pro at 80.3%, far ahead of GPT 5.5's 58.6%. Pricing doubled but the Coding Plan may offer a short-term cost arbitrage. Cross-source cluster confirmed, all three HKR axes hit. Minus 1 point because the post doesn't disclose Mythos ...

Read the original ↗Export Markdown