NatureBench: AI coding agents beat published SOTA on only 17.8% of Nature-family tasks
NatureBench:AI编码智能体能否匹配Nature系列论文已发表SOTA?
NatureBench pulls 90 cross-discipline tasks from published Nature-family papers to test whether AI coding agents can beat the original SOTA. Under a no-web-search protocol, the strongest agent wins on only 17.8% of tasks (g>0.1). Agents succeed mostly by translating scientific problems into familiar supervised-learning setups, not through genuine invention. Failures are dominated by wrong method choice and insufficient compute, not task misunderstanding. Code, the NatureGym pipeline, and a public leaderboard are released.
Why it matters: NatureBench pulls 90 cross-disciplinary tasks from Nature-family papers to test AI coding agents; the top config beats original SOTA on only 17.8% of tasks. The benchmark design is provocative and the numbers are concrete, but the paper just hit arXiv without peer review — I'm...