WhatWorkedBench exposes a gap in AI agents' grasp of experiments
What happened
Reported on October 2, a team from Carnegie Mellon University and Tsinghua University proposed WhatWorkedBench, which tests whether AI research agents can predict untested configurations from limited experiments. The benchmark covers 30 data sources, 8 workflow types and 1,248 configurations. The study finds a gap between picking the best configuration and understanding experimental effects. On the same observations, agents recovered effects at 0.632, below a Gaussian process at 0.698. Fixing code constraints the agents violated raised average recovery from 0.338 to 0.507 with no new experiments.
Written by AI from the coverage · updated 47 minutes ago
Coverage
Follow the reports to see the story from different sides.
- Computing Life · Share · YageWhatWorkedBench 研究发现 AI 研究代理选中最优配置仍难准确预测未测实验
卡耐基梅隆大学和清华大学团队提出 WhatWorkedBench,评测 AI 研究代理能否利用有限实验预测未测配置,发现选中最优配置与理解实验效应存在差距。基准涵盖 30 个数据源、8 类工作流和 1248 个配置,同批观测下代理的效应还原度为 0.632,高斯过程为 0.698;修复代理违反的代码约束后,平均还原度从 0.338 提升至 0.507,无需新增实验。
Heat over time
Not enough continuous observations to draw a trend yet.