UC Berkeley team used a cheating AI to break 8 major agent benchmarks and score near perfect without solving tasks
伯克利大学的研究团队造了一个专门作弊的 AI,用它去攻击目前最主流的 8 个 AI 智能体评测基准,结果每一个都被攻破了。没有解决任何任务,没有调用任何大模型,拿到了接近满分的成绩。
A UC Berkeley team used a cheating AI with no LLM calls to break 8 major agent benchmarks, scoring 73% to 100% without solving tasks. The post cites three cases: a 10-line Python hook bypassed SWE-bench tests across 500 tasks, WebArena exposed answers via file://, and FieldWorkArena gave full credit to an empty {} reply. The real issue is benchmark isolation failure; the team is turning its scanner into the open-source BenchJack project.
Why it matters: HKR-H/K/R all pass: the claim is clicky, concrete, and directly threatens trust in agent evals. I stop at 84, not 85+, because the current input is a social summary; paper status, full methods, and outside replication are not disclosed here.