Skip to content
Trending storyDeveloping

WhatWorkedBench exposes a gap in AI agents' grasp of experiments

1 report1 sourceupdated 5 hours ago

What happened

AI digest

Reported on October 2, a team from Carnegie Mellon University and Tsinghua University proposed WhatWorkedBench, which tests whether AI research agents can predict untested configurations from limited experiments. The benchmark covers 30 data sources, 8 workflow types and 1,248 configurations. The study finds a gap between picking the best configuration and understanding experimental effects. On the same observations, agents recovered effects at 0.632, below a Gaussian process at 0.698. Fixing code constraints the agents violated raised average recovery from 0.338 to 0.507 with no new experiments.

Written by AI from the coverage · updated 47 minutes ago

Coverage

Follow the reports to see the story from different sides.

Oct 3
  1. Computing Life · Share · Yage
    WhatWorkedBench 研究发现 AI 研究代理选中最优配置仍难准确预测未测实验

    卡耐基梅隆大学和清华大学团队提出 WhatWorkedBench,评测 AI 研究代理能否利用有限实验预测未测配置,发现选中最优配置与理解实验效应存在差距。基准涵盖 30 个数据源、8 类工作流和 1248 个配置,同批观测下代理的效应还原度为 0.632,高斯过程为 0.698;修复代理违反的代码约束后,平均还原度从 0.338 提升至 0.507,无需新增实验。

Heat over time

Not enough continuous observations to draw a trend yet.