Trending storyDeveloping
DoGBench releases a benchmark and evaluation for user document generation
1 report1 sourceupdated 2 hours ago
What happened
AI digest
On October 2, Hacker News featured DoGBench, a benchmark for how well AI agents generate user documentation. It uses real open-source project tasks. The paper evaluation covers seven systems and tops out at a combined score of 47.3. A separate leaderboard adds three cloud agents, with a top score of 54.8. The two highs come from different evaluation scopes, not from the same systems improving over time.
Written by AI from the coverage · updated 37 minutes ago
Coverage
Follow the reports to see the story from different sides.
Oct 2
- Hacker News front pageDoGBench: The first user-facing docs generation benchmark. No model scores >50%
DoGBench 用真实开源项目任务评估 AI 智能体的用户文档生成能力,论文七组系统最高综合得分为 47.3,榜单另含三组云端智能体,最高综合得分为 54.8。
Heat over time
Not enough continuous observations to draw a trend yet.