Claude Sonnet 5.5 enters Arena's Agent Arena and Battle Mode
Claude Sonnet 5.5 上线 Arena 的 Agent Arena 与 Battle Mode 评测
Anthropic's Claude Sonnet 5.5 is now available for voting in Arena's Agent Arena. The leaderboard evaluates models on millions of real-world long-horizon agent tasks where models can use web search, filesystem, and terminal tools. Rankings use a causal tracking method to measure how much a model outperforms the average.
Why it matters: Claude Sonnet 5.5 hitting Arena's Agent leaderboard is a direct user-facing eval signal, hitting all three HKR axes. Score capped at 74 because the post only describes the methodology — no specific win rates or rankings disclosed. Adjust upward once concrete numbers drop.