Skip to content
Trending storyDeveloping

DoGBench releases a benchmark and evaluation for user document generation

1 report1 sourceupdated 2 hours ago

What happened

AI digest

On October 2, Hacker News featured DoGBench, a benchmark for how well AI agents generate user documentation. It uses real open-source project tasks. The paper evaluation covers seven systems and tops out at a combined score of 47.3. A separate leaderboard adds three cloud agents, with a top score of 54.8. The two highs come from different evaluation scopes, not from the same systems improving over time.

Written by AI from the coverage · updated 37 minutes ago

Coverage

Follow the reports to see the story from different sides.

Oct 2
  1. Hacker News front page
    DoGBench: The first user-facing docs generation benchmark. No model scores >50%

    DoGBench 用真实开源项目任务评估 AI 智能体的用户文档生成能力,论文七组系统最高综合得分为 47.3,榜单另含三组云端智能体,最高综合得分为 54.8。

Heat over time

Not enough continuous observations to draw a trend yet.