Stanford launches Terminal-Bench-Science: scientists set the bar for AI agents on real research workflows
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
Stanford researchers released Terminal-Bench-Science 0.1, a benchmark built from real scientific workflows contributed by practicing scientists. The first release has 70 tasks across life, physical, Earth, mathematical, and engineering sciences. Claude Opus 5 tops the board at 30% resolution rate; GPT-5.6 Sol hits 22.4% and Claude Fable 5 reaches 21.4%. Only 70 tasks made the cut from 920 proposals, with 376 contributors across 22 countries. The benchmark is designed to evolve continuously, giving the scientific community a direct voice in setting the bar for AI capability.
Why it matters: Stanford-led Terminal-Bench-Science 0.1 evaluates AI agents on real research workflows curated by domain scientists — 70 tasks from 920 proposals, Claude Opus 5 at 30% resolution. Hits all three HKR axes: novel setup, concrete numbers, resonates with agent builders and science...