Why language models hallucinate
OpenAI says language models hallucinate because standard training and evals reward guessing instead of admitting uncertainty. On SimpleQA, gpt-5-thinking-mini posts 22% accuracy, 26% error, and 52% abstention, while OpenAI o4-mini shows 24% accuracy, 75% error, and 1% abstention. The key issue is scoring design, not accuracy-only leaderboards.
Why it matters: Strong HKR-H/K/R: the post reframes hallucination as an eval-objective problem and includes testable SimpleQA numbers. Featured, not p1, because this is a research/explainer release rather than a major model, product, funding, or personnel event.