GitHub announced ReviewBench on 5 October: a public research-preview benchmark for comparing AI code-review agents. Its evaluation set contains 219 public pull requests across 19 programming languages.
What it measures
GitHub says the benchmark’s reference findings combine human reviews, static analysis and multiple AI models. Findings are checked against a shared rubric and labelled by severity and category.
Two questions sit behind the scoring: how many valid issues does a reviewer find, and how many of its comments are valid? These are recall and precision. ReviewBench also evaluates newly identified issues that were absent from the original reference set.
GitHub provides the dataset, evaluation method and runner so developers can inspect and evaluate their own review systems. The company reports that benchmark changes have anticipated the direction of its production experiments. That is GitHub’s reported result, not an independent finding by this publication.
Our assessment
This gives teams a concrete starting point for asking whether a review agent is useful. A reviewer that produces many comments can still waste time if developers have to dismiss most of them.
For your own evaluation, select recent changes with known defects and changes that should pass without comments. Record important problems found, incorrect warnings and review time. Keep security defects separate from minor suggestions so one average score does not hide a serious weakness.
A public benchmark should complement that exercise. Your repository’s dependencies, conventions and failure patterns may differ from the evaluation set. We have not run ReviewBench or compared agents ourselves.
Source: GitHub’s ReviewBench announcement and methodology.
Read the original material
GitHub (opens in a new tab)The article date belongs to Agentic Horizon. The source publication date is listed separately. Vendor claims remain attributed; we have not independently tested the reported capability.

