On this page
- Google developers spent about three hours a week on review
- Published AI reviewer figures are graded differently and mostly self-reported
- Beko's review bot added comments on top of human ones, and closure slowed
- False positives teach engineers to ignore the reviewer
- GitHub ships silence as a feature and reports how often it fires
- Studies find RLHF and human votes both favor longer answers
- Claude Code publishes the exception list a length control needs
- Set a comment cap per review and carry the exceptions through it
- curl ended its bug bounty after the confirmed rate fell below 5%
- Where the published record is thin
Output Discipline: Cap AI Code Review Comments and Enforce a Precision Bar
An AI reviewer that comments too much spends the one input review runs on: a human's attention. Google's Tricorder puts any code review analyzer with a not-useful rate at or above 10% on probation and reserves the right to disable it when "the analyzer results are annoying developers" (Tricorder, ICSE 2015).
- Cap comments per review, never the exceptions. Give the agent a hard comment budget, rank findings by severity, and drop the tail. Error reports, failing test output, security warnings, and destructive-action confirmations stay complete under any cap or length setting.
- Track a not-useful rate per reviewer. Tricorder scored each analyzer by developer clicks; pick your own threshold and act on it.
- Let silence be a result. When nothing clears the bar, the reviewer posts nothing, as GitHub's Copilot code review does on 29% of reviews.
- Set length once, not per prompt. Use a verbosity parameter or a session-wide output style instead of rewriting the prompt; OpenAI's GPT-5 guidance says to keep prompts stable and use the parameter.
- Read count next to address rate. A falling comment count shows the cap is working only if the share of comments developers act on holds or rises.
How much the agent says is a control of the same shape as scope discipline for diff size and approval fatigue for prompts.
Google developers spent about three hours a week on review
Google's developers spent an average of 3.2 hours a week reviewing changes, measured over five weeks starting October 2016, and the median change had one reviewer (Sadowski et al., ICSE-SEIP 2018). Every review comment draws on that.
No company publishes reviewer throughput or time-in-review from before and after agent adoption (see review capacity). In repositories with at least 100 stars, 61.38% of agent-authored PRs have no recorded review activity, and agents write 71.58% of the review comments on agent-authored PRs (Duma et al., EASE 2026). The authors caution that a missing record does not prove nobody looked.
Published AI reviewer figures are graded differently and mostly self-reported
| Reviewer | Figure | Graded by |
|---|---|---|
| Google Tricorder, analyzers | Not-useful rate around 5% (2014); probation at 10% | Developer clicks |
| CodeRabbit, in the wild | Of comments that drew a reply: 56.3% rejected, 36.4% accepted | Developer replies, LLM-labeled; independent |
| Greptile v5 | Comments addressed by the author: 52% (v4) to 66% (v5) | Vendor A/B; grading method not stated |
| Cursor Bugbot | 78.13% of flagged bugs resolved by merge, public repos | LLM judge, vendor-reported |
| Uber uReview | 75% of comments rated useful, over 65% addressed | Engineer ratings and automated check, self-reported |
| GitHub Copilot | Actionable feedback on 71% of reviews, silent on 29% | Vendor-reported |
Tricorder's 10% applies to "Not useful" clicks on static-analyzer findings. It shows that an enforced bar can exist; it is not a number to copy onto an LLM.
Weight the AI rows by who graded them and what they count. The CodeRabbit study covers only the 31,073 of 332,693 comments that drew a developer reply. Uber's ratings come from engineers who use the tool; Cursor's judge scores Cursor and its competitors.
Beko's review bot added comments on top of human ones, and closure slowed
Appliance maker Beko deployed a reviewer built on Qodo's open-source PR-Agent; the study covers three projects (Cihan et al., ICSE-SEIP 2025). The bot averaged 3.65 comments per PR. Human comments went from 0.31 to 0.28 per PR, a drop the authors found not significant, so the bot's comments came on top.
Average PR closure time rose from 5h52m to 8h20m, though one of the three projects got faster. Developers labeled 73.8% of the bot's comments resolved.
False positives teach engineers to ignore the reviewer
Uber's uReview grades, dedupes, and suppresses low-value categories before anything reaches a human (AI reviewing AI). Uber's stated reason: when engineers "encounter many false-positive comments, they start to tune out and ignore them" (Uber).
GitHub ships silence as a feature and reports how often it fires
"Silence is better than noise. In 71% of the reviews, Copilot code review surfaces actionable feedback. In the remaining 29%, the agent says nothing at all." In March 2026 it averaged about 5.1 comments per review (GitHub, March 2026).
The 29% shows a threshold at work; apply the same test to a gate that never returns a clean pass (review as a gate).
Studies find RLHF and human votes both favor longer answers
Singhal et al. found RLHF reward gains in three settings driven largely by response length, with reward models the main source of that bias: "a purely length-based reward reproduces most downstream RLHF improvements over supervised fine-tuned models" (COLM 2024). In human preference votes, LMArena's observational style-control analysis found that "length was the dominant style factor" (LMArena).
No source here measures whether a brevity instruction, in a prompt or an output style, holds over a long session. Prose rules live in agent writing style.
Claude Code publishes the exception list a length control needs
GPT-5's verbosity parameter takes low, medium, or high, and OpenAI's guidance is to "Keep prompts stable and use the param rather than re-writing" (OpenAI).
Claude Code's Concise output style leads with the result and leaves out preamble, narration, and recaps. Safety output stays whole: "error reports, failing test output, security warnings, and confirmations for destructive actions keep their complete content" (Claude Code docs).
Set a comment cap per review and carry the exceptions through it
Most controls in these sources decide per comment: confidence thresholds and category suppression (Uber), category filtering and majority voting across parallel passes in Cursor's launch Bugbot (Cursor), and silence when nothing clears the bar (GitHub). A total cap is rarer: PR-Agent's current docs give /review a num_max_findings setting, "Maximum number of returned findings" (docs), with a default of 3 (configuration).
Going past what these sources measure, start with a hard comment budget per review. A per-comment threshold cannot notice that the twelfth true-but-minor finding costs more attention than it returns; a cap forces that ranking. Carry the exception list through the cap so it never displaces an error report, a security finding, or a destructive-action warning.
curl ended its bug bounty after the confirmed rate fell below 5%
curl's program ran from April 2019 until January 31, 2026, and paid over $100,000 for 87 confirmed vulnerabilities. Daniel Stenberg gave three reasons: AI slop, "humans doing worse than ever," and "the apparent will to poke holes rather than to help." He also writes that the confirmed rate fell from "north of 15% of the submissions" to below 5% starting in 2025 (Stenberg, January 2026).
In July 2025 he put AI slop at about 20% of 2025 submissions and said every report ties up three or four of the seven-person security team, sometimes for up to three hours (July 2025).
Where the published record is thin
- Uber aims to keep uReview's usefulness rate above 75%, the only published target among the AI reviewers here, and none describes enforcing one the way Google disabled its HTML linter because "the linter was so rarely helpful" (Google SWE Book, chapter 20).
- None of these sources reports how a comment cap moves address rate, so a starting budget has nothing to calibrate against.