作者开源 Claude Code 技能 pr-proof,将 CodeRabbit 代码评审噪声减少三分之一
I made CodeRabbit's reviews a third less noisy with an open-source Claude Code skill
作者开源 pr-proof,一套 Claude Code 技能,把每条评审意见当作需对照代码验证的主张来核查。
AI code review has a noise problem. On a public benchmark of 50 real pull requests, CodeRabbit raised 300 issues. 77 of them were real bugs on the benchmark's list. Most of the rest were noise: things that look like problems but aren't once you read the code.
Developers learn to skim review bots, and once you skim, the real bugs slip past you too.
I built pr-proof, three Claude Code skills that treat every review comment as a claim that has to be proven against the code before anyone acts on it. Run over CodeRabbit's reviews of those 50 PRs, it:
- kept 72 of the 77 real bugs (93.5%)
- removed 76 of the 223 noise issues (34%)
- raised CodeRabbit's F1 score from 35.2% to 40.4% (+5.2 points, 95% CI +1.9 to +8.3)
What "noise" looks like
Here are real comments CodeRabbit left on a Grafana PR, and what checking them against the code showed:
"
promResponse.data.groups.at(0)?.rules.map()will throw if groups is empty."
It won't. Optional chaining short-circuits the whole rest of the chain: if at(0) returns undefined, .rules.map(...) is never evaluated.
"
useCanSilenceinvokes hooks before its early return when rule is undefined."
That order is required. The Rules of Hooks say hooks must run unconditionally, before any early return.
"
handleDeleteexpectsRulerRuleDTObut call sites passEditableRuleIdentifier."
Both call sites pass zero-argument closures (RuleActionsButtons.V2.tsx:99).
Each one sounds plausible and falls apart once you read the code, and reading the code is exactly the step review bots skip and leave to you.
How pr-proof works
The core skill, pr-comment-validation, takes review comments from anyone (a human, CodeRabbit, Copilot, another agent) and investigates each one:
- Read the whole file the comment refers to, not just the diff hunk, plus its callers, interfaces and siblings.
- Trace the execution path. Does the failure the reviewer describes actually happen?
- Check the claim against the codebase's own conventions and the official docs for any library behaviour it depends on.
- Return a verdict: valid, partly valid (a real problem, but overstated or with the wrong fix), wrong, or style preference. Every verdict cites the file and line that prove it.
Nothing is changed or posted until every comment has a verdict.
Two more skills build on it:
-
pr-validationdoes the whole loop. It checks out the PR in a worktree, validates every comment in parallel, shows you the verdicts, applies the fixes you approve, and replies on each thread. -
pr-reviewwrites its own review, then has independent subagents try to disprove each finding before anything is posted.
Install it in Claude Code:
/plugin marketplace add TanayK07/pr-proof
/plugin install pr-proof@pr-proof
Then, on any PR with review comments: "are the review comments on PR #123 valid?"
How I measured it
I used Code Review Bench, an open benchmark with 50 real PRs from Sentry, Grafana, Keycloak, Discourse and Cal.com. Each PR comes with a human-written list of the real issues. It already holds every major tool's reviews, scored by Claude Opus 4.5 as the judge.
That made the filter test clean:
- The benchmark had already split CodeRabbit's reviews into 300 individual issues and labelled each one as real or noise.
- pr-proof saw each issue's text and the checked-out code, and never the labels.
- I kept the issues it ruled valid or partly valid and scored them with the benchmark's own published labels. No new judging was involved.
Each run was a headless Claude Code session with no network, no gh and no curl, so it couldn't read the original PR discussions the answers came from.
What didn't work
I also benchmarked pr-review as a reviewer in its own right, against plain Claude Code on the same model. It came out statistically level: 29.8% vs 29.1% F1.
Two things I learned along the way:
- My first version of self-validation made things worse. I told the validators that anything uncertain was wrong. They took that literally and threw out real bugs, calling them "pre-existing" or "looks intentional". Switching to the same criteria
pr-comment-validationuses fixed the harm, but only added about 1.6 points. - Run-to-run variance is large. Two runs of the identical drafting step scored 33.5% and 28.2%. A single run of any tool on 50 PRs can move a few points just by chance. Treat any leaderboard gap under ~5 points with suspicion, including mine.
I also caught a bug in my own scoring halfway through. On 37 PRs the benchmark had silently fallen back to judging whole comments instead of individual issues, which made pr-review look like it scored 45%. It didn't. The harness now refuses to print results if that happens.
Try it, and check my numbers
Everything is in the repo: the skills, the benchmark harness, every review and verdict per PR, and the scripts that turn them into these numbers.
- Repo: https://github.com/TanayK07/pr-proof
-
Reproduce:
bench/README.md
If you run it on your own PRs, I'd love to hear where it gets things wrong.
来源:Google AI:DEV 作者专属(RSS) · dev.to