Skip to content
OpenAI News

OpenAI releases LifeSciBench: a benchmark built by PhD scientists for real research tasks

Introducing LifeSciBench

OpenAI released LifeSciBench, a 750-task benchmark authored and reviewed by PhD scientists with biotech/pharma experience. It tests real research workflows—interpreting conflicting evidence, designing experiments, assessing translational risk—not fact recall. 53% of tasks require processing attached artifacts like figures or sequence files, averaging four reasoning steps per task. Grading uses 25 rubric criteria per task on average, checking scientific validity and operational usefulness, not just final answers. The post does not disclose model scores.

Why it matters: OpenAI released a PhD-scientist-written benchmark with 750 questions testing experimental design, conflicting-evidence interpretation, and translational risk assessment — closer to real research workflows than existing benchmarks. Score capped here because only a preprint and ...

Read the original ↗Export Markdown