Skip to content
MIT Technology Review · AI

AI benchmarks are broken. Here’s what we need instead.

The author proposes HAIC benchmarks that evaluate AI over longer periods inside teams and workflows, not on isolated tasks alone. The post lists four shifts and cites a UK hospital study from 2021–2024 plus an 18-month humanitarian case; the key signal is coordination, error detectability, and downstream effects, not a 98% accuracy headline.

Why it matters: This hits all three HKR axes: a contrarian headline, a concrete 4-part framework with two field cases, and a strong resonance with the industry's eval-vs-production debate. It is a strong commentary piece, not a model release, benchmark launch, or research drop, so it lands in `f

Read the original ↗Export Markdown