Anthropic's Claude Sonnet 5.5 nearly matches Opus 5.5 on benchmarks while costing up to 30% less per task
Anthropic 发布 Claude Sonnet 5.5,基准接近 Opus 5.5 且单任务成本最多降 30%
Anthropic released Claude Sonnet 5.5, aimed at everyday tasks like bug fixes and doc writing. It generates output over 30% faster and costs up to 30% less per task—not by lowering token price, but by using fewer tokens per task. Coding gains are the headline: Terminal-Bench 4.0 jumps from 10.3% (Sonnet 5) to 70.6%, and CursorBench 4.0 hits 55.5%, just 2.3 points below Opus 5.5. On the knowledge-work benchmark GDPval-AA, it scores 1,844 vs. Opus 5.5's 1,846. One oddity: max reasoning effort on FrontierCode scores worse than the second-highest setting; Anthropic says a code-review function caused timeouts or scope drift. The model is live on AWS, Google Cloud, and Azure, with new safeguards against cybersecurity risks and distillation attacks. The post does not disclose Haiku 5.5 specs or a firm launch date, only 'in the coming weeks.'
Why it matters: Anthropic mid-tier update with a big coding leap and 30% lower per-task cost—directly useful signal for Claude users. Score capped below 85 because only one source so far, and the post doesn't disclose full benchmark tables or exact pricing; wait for more hands-on results.