Skip to content
GitHub Blog · AI & MLMichelle Zhou

GitHub releases ReviewBench, an open benchmark for AI code review

ReviewBench: An open benchmark for AI code review

GitHub has put out a research preview of ReviewBench, an open benchmark for AI code review, with a public dataset, evaluation method and self-serve tooling. The benchmark mirrors the distribution of 103.9 million GitHub pull requests and covers 219 pull requests from 187 repositories across 19 languages. Senior engineers independently reviewed the labels, agreeing 96.6% of the time.

Why it matters: ReviewBench scores fixed labels and newly found valid issues separately, so review agents are compared on the same footing without being penalized for catching real problems the labels missed.

Read the original ↗Export Markdown