GitHub releases ReviewBench, an open benchmark for AI code review
What happened
On October 5 GitHub introduced ReviewBench, an open benchmark for AI code review, with a research preview that publishes the dataset, the evaluation method and a tool for running it yourself. The benchmark mirrors the distribution of 103.9 million GitHub pull requests and contains 219 pull requests from 187 repositories across 19 programming languages. Senior engineers independently rechecked the labels, with 96.6% agreement. The research preview already offers that data and those tools.
Written by AI from the coverage · updated 2 hours ago
Coverage
Follow the reports to see the story from different sides.
- GitHub Blog · AI & MLPickGitHub releases ReviewBench, an open benchmark for AI code review
GitHub has put out a research preview of ReviewBench, an open benchmark for AI code review, with a public dataset, evaluation method and self-serve tooling. The benchmark mirrors the distribution of 103.9 million GitHub pull requests and covers 219 pull requests from 187 repositories across 19 languages. Senior engineers independently reviewed the labels, agreeing 96.6% of the time.
Heat over time
Not enough continuous observations to draw a trend yet.