The Second Half of Agent Evaluation: Why a Live Benchmark Is Needed
Agent评测的下半场:为什么需要一个「活的」Benchmark?
Claw-Eval-Live evaluates 13 frontier models on 105 tasks, and the top model stays below a 70% pass rate, while HR tasks average only 6.8% pass rate.
Why it matters: HKR-H/K/R all pass: the live benchmark hook is specific, and the post gives 105 tasks, 13 models, HR at 6.8%. Claw-Eval-Live still lacks proven field impact, so this sits in the lower featured band.