A leaderboard tracking 30 model cards to see which benchmarks frontier labs actually report
I checked 30 frontier model cards. Here are the benchmarks labs report
This project scanned 30 model cards from 11 orgs and counted how often 79 benchmarks are mentioned—it measures vendor attention, not benchmark quality. MATH-500 and Arena-Hard are near ceiling, losing discriminative power. DeepSeek's own models gained 40.6 points on AIME and 25.4 on LiveCodeBench in 26 days. Six benchmarks, including BrowseComp and SWE-bench Pro, are reported by at least 4 orgs but have no readable scores. The newer APEX-Agents already appears in 3 independent cards, though scores couldn't be read either.
Why it matters: Scans 30 model cards from 11 orgs, measuring vendor attention rather than benchmark quality — a useful lens. Concrete numbers like MATH-500 near-saturation and DeepSeek's 40.6-point AIME jump in 26 days will spark discussion. Docked because it's a personal project with limited...