Artificial Analysis launches Intelligence Index v4.2 with private test sets to prevent gaming
Artificial Analysis Intelligence Index v4.2
Artificial Analysis updated its model benchmark to v4.2, adding two new evaluations: AA-Briefcase and GDP.pdf. AA-Briefcase uses a private test set to simulate multi-week knowledge work projects and assess holistic agentic capability. GDP.pdf requires models to synthesize evidence across 4,592 pages of professional documents, graded against 1,275 atomic criteria where a task passes only if every criterion is met. Claude Fable 5.1 leads the index, followed by GPT-6 Astra, which shows an ~85 Elo gain over GPT-5.6 Sol. Private test sets now account for 40% of the weighting, double the v4.1 figure, specifically to reduce gaming by labs.
Why it matters: AA's leaderboard refresh matters for model selection workflows — the private test sets and 4,592-page document eval are more grounded than saturated public benchmarks. Not scoring higher because this is methodology iteration, not a capability breakthrough, and the post only gi...