Skip to content
Trending storyPast story

Claude Opus 5.5 tops Coding Agent Index, but per-task cost rises to $13.04

1 report1 sourceupdated 5 days ago

What happened

Summary

Artificial Analysis 把 Claude Opus 5.5 放在 Claude Code 的“最大出力”模式下跑了一遍 Coding Agent Index,得分 66,比上一代 Opus 5 的 60 分高出 6 分。三项子测试全线上涨:Terminal-Bench 4.0 正确率 63.1%,DeepSWE v1.1 正确率 68....

Coverage

Follow the reports to see the story from different sides.

Sep 24
  1. AI HOT (Curated Pool)Pick
    Claude Opus 5.5 tops Coding Agent Index, but per-task cost rises to $13.04

    Artificial Analysis tested Claude Opus 5.5 under Claude Code max effort and it scored 66 on the Coding Agent Index, up from Opus 5's 60. All three subtests improved: Terminal-Bench 4.0 63.1%, DeepSWE v1.1 68.4%, SWE-Atlas-QnA 66.4%. The trade-off: per-task cost jumped from $3 to $13.04. The post doesn't break down how max effort drove the cost increase.