OpenAI audits SWE-Bench Pro, finds ~30% of tasks are broken
Separating signal from noise in coding evaluations
OpenAI audited SWE-Bench Pro and estimates ~30% of its tasks are broken. An automated pipeline flagged 286 suspicious tasks; Codex-based investigator agents and five experienced engineers then reviewed them. Engineers identified 249 (34.1%) flawed tasks, mostly due to overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts. OpenAI advises model developers to scrutinize results rather than trust leaderboard scores. The post does not disclose a fix timeline or a revised dataset release.
Why it matters: OpenAI audited SWE-Bench Pro and found ~34% of tasks defective — a ratio that forces the industry to re-examine coding benchmark reliability. The post provides concrete defect categories and a human review pipeline. Not scored higher because this is a benchmark quality report,...