Skip to content
Hacker News front page

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

This ICML 2026 paper defines and measures benchmark saturation across 60 LM benchmarks. Nearly half are already saturated, and older benchmarks saturate faster. Expert curation helps resist saturation; keeping test data private does not. The post doesn't name the specific benchmarks but identifies 14 properties linked to saturation, offering a framework for building longer-lasting evaluations.

Why it matters: ICML 2026 paper with a systematic audit of 60 benchmarks and counterintuitive findings (hidden test sets don't delay saturation). Solid eval-infra research, but no named benchmarks limits immediate impact, capping at 78.

Read the original ↗Export Markdown