Cursor releases CursorBench 3.1, scoring coding agents on real multi-file tasks from user sessions
Cursor built CursorBench 3.1 from real user sessions—ambiguous, multi-file coding tasks—and added codebase understanding, bugfinding, planning, and code review problems. Fable 5 Max leads at 72.9%, averaging $18.02, 63,842 tokens, and 76 steps per task. GPT-5.5 High scores 62.6% at just $3.59 per task, a much better cost-performance ratio. Composer 2.5 hits 63.2% for only $0.55, the cheapest on the board. The post doesn't disclose the total number of tasks or grading details; small score differences may not be statistically meaningful.
Why it matters: Cursor drops CursorBench 3.1, a benchmark built from real user sessions. Fable 5 Max leads at 72.9% but costs $18/task and 76 steps; GPT-5.5 High scores 64.3%. Concrete numbers and head-to-head comparison make this directly useful for Cursor users. Not scoring higher because i...