GitHub releases ReviewBench, an open benchmark for AI code review
ReviewBench: An open benchmark for AI code review
GitHub has put out a research preview of ReviewBench, an open benchmark for AI code review, with a public dataset, evaluation method and self-serve tooling. The benchmark mirrors the distribution of 103.9 million GitHub pull requests and covers 219 pull requests from 187 repositories across 19 languages. Senior engineers independently reviewed the labels, agreeing 96.6% of the time.
Why it matters: ReviewBench scores fixed labels and newly found valid issues separately, so review agents are compared on the same footing without being penalized for catching real problems the labels missed.