Skip to content
AI HOT (Curated Pool)

Claude Sonnet 5.5 hits #2 on AA Intelligence Index, matching Opus 5.5 by spending ~193k output tokens per task

Claude Sonnet 5.5 达到 Artificial Analysis 智能指数第 2 名

Anthropic released Claude Sonnet 5.5, scoring 56 on the AA Intelligence Index—2 points behind Opus 5.5. Pricing stays at $2/$10 per million input/output tokens, but cost per task hits ~$7.60, about 50% more than Sonnet 5, because it uses ~193k output tokens per task at max effort. That's 60% more than Opus 5.5 and 7x GPT-6 Astra. It matches Opus 5.5 on agentic terminal use and knowledge work: 64% on Terminal-Bench 4.0 vs Opus 5.5's 60%, and near-identical scores on AA-Briefcase, GDPval-AA, and AutomationBench-AA. It lags on factual knowledge (54% vs 66% accuracy on AA-Omniscience, but lower hallucination rate at 47% vs 59%) and scientific reasoning, trailing Opus 5.5 by ~6 points on Humanity's Last Exam and SciCode. Evaluations used a pre-release build with a structured-output bug that is fixed for launch; Anthropic expects performance to be at least as good. Context window remains 1M tokens with image and text input.

Why it matters: Anthropic released Sonnet 5.5, and Artificial Analysis provides hard data: 56 on the Intelligence Index (#2), 64% on Terminal-Bench 4.0 matching Opus 5.5 and GPT-6 Astra, but $7.60 per task—50% pricier than Sonnet 5. All three HKR axes hit: tension, numbers, and cost math for ...

Read the original ↗Export Markdown