Skip to content
Hacker News front page

Stanford launches Terminal-Bench-Science: scientists set the bar for AI agents on real research workflows

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Stanford researchers released Terminal-Bench-Science 0.1, a benchmark built from real scientific workflows contributed by practicing scientists. The first release has 70 tasks across life, physical, Earth, mathematical, and engineering sciences. Claude Opus 5 tops the board at 30% resolution rate; GPT-5.6 Sol hits 22.4% and Claude Fable 5 reaches 21.4%. Only 70 tasks made the cut from 920 proposals, with 376 contributors across 22 countries. The benchmark is designed to evolve continuously, giving the scientific community a direct voice in setting the bar for AI capability.

Why it matters: Stanford-led Terminal-Bench-Science 0.1 evaluates AI agents on real research workflows curated by domain scientists — 70 tasks from 920 proposals, Claude Opus 5 at 30% resolution. Hits all three HKR axes: novel setup, concrete numbers, resonates with agent builders and science...

Read the original ↗Export Markdown