Skip to content
Hacker News front page

AGI Ranker fixed its ceiling logic—every model dropped 6–15 points

Show HN: I audited my AI leaderboard scale – every score dropped 6-15 points

AGI Ranker aggregates 10 public benchmarks into a single 0–100 score, with 100 marking the AGI threshold. A v2.0.0 audit found that nearly all benchmark ceilings were incorrectly labeled as best-human; only GPQA Diamond has a published human study (0.81). The other nine ceilings are just the benchmark maximum, not human parity. Fixing this dropped every model by 6–15 points. The score is now a mixed scale, not a pure human comparison. They also removed Artificial Analysis's Intelligence Index after its v4.1 redesign caused double-counting with existing benchmarks. The leaderboard relies on published data, doesn't run evals itself, and weights sources by a four-tier credibility system—self-reports get discounted.

Why it matters: AGI Ranker publicly audited its own methodology and corrected the 'human baseline' on 9 of 10 benchmarks, dropping every model's score by 6-15 points. Hits all three HKR axes for anyone working on model evaluation. Score stays at 72 because this is a methodology fix, not a new...

Read the original ↗Export Markdown