ReviewBench: An open benchmark for AI code review
We’re launching ReviewBench, a benchmark for code review agents built on representative GitHub pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics.
Alejandro’s work focuses on evaluating coding agents, improving how they reason over code changes and repository context, and optimizing model quality, coverage, latency, and token cost.
We’re launching ReviewBench, a benchmark for code review agents built on representative GitHub pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics.