Skip to content
Xinzhiyuan · WeChat

The Second Half of Agent Evaluation: Why a Live Benchmark Is Needed

Agent评测的下半场:为什么需要一个「活的」Benchmark?

Claw-Eval-Live evaluates 13 frontier models on 105 tasks, and the top model stays below a 70% pass rate, while HR tasks average only 6.8% pass rate.

Why it matters: HKR-H/K/R all pass: the live benchmark hook is specific, and the post gives 105 tasks, 13 models, HR at 6.8%. Claw-Eval-Live still lacks proven field impact, so this sits in the lower featured band.

Read the original ↗Export Markdown