Skip to content
AI HOT (Curated Pool)

Cursor's eval lead confirms Claude Fable 5 hits 72.9% on CursorBench, targeting the hardest 1% of coding tasks

Cursor 评估负责人确认 Claude Fable 5 在 CursorBench 达 72.9% 新高

Cursor's eval lead Nate Schmidt explains on Anthropic's blog how they determined Claude Fable 5 was ready for the hardest 1% of real-world coding problems. The headline number is 72.9% on CursorBench, a significant jump over the prior generation. The post stresses this isn't a generic benchmark grind—it targets long-tail tasks that actually stump developers. The article doesn't disclose the baseline score, test set size, or sample problems, so treat the 72.9% as a directional signal rather than a cross-benchmark comparison point.

Why it matters: Cursor's eval lead publishes on Anthropic's blog with a concrete 72.9% CursorBench score — a substantive first-party eval. The post doesn't disclose the previous-gen baseline or test set size, so score lands at 82 rather than higher.

Read the original ↗Export Markdown