Introducing SimpleQA
OpenAI open-sourced SimpleQA, a 4,326-question benchmark for factual short-answer QA and model calibration. Two independent AI trainers verified each item; a 1,000-question audit showed 94.4% agreement and an estimated inherent error rate near 3%. The key signal: it is built to challenge frontier models, and the post says GPT-4o scores below 40%.
Why it matters: This is not a routine paper post. HKR-H comes from the inversion that a 'simple' benchmark stumps frontier models; HKR-K comes from the dataset size, agreement rate, and irreducible-error estimate; HKR-R comes from the ongoing industry fixation on hallucination and calibration,so