Rejudge 用 3 个独立模型加 1 个裁判替代 AI 代码自审
Rejudge Replaces Self-Review With 3 Independent Models and a Judge
Rejudge 是一个代码评审 CLI 工具,把同一评审请求发给 3 个上下文隔离的 reviewer,再由无工作区访问权限的 judge 比对报告、在分歧时通过 ask_panel 追问并输出最终结论。
作者实测了用三个隔离模型加一个裁判替代单模型自审的代码评审流程,并给出隐私边界和成本权衡的具体提醒。
I have a hard time trusting an AI coding agent to review code it just wrote.
The model already made the decision that the implementation was good enough to produce.
Then we ask it to inspect that same implementation using many of the same learned habits and assumptions.
Sometimes it catches mistakes. Sometimes it revalidates them. A fresh session helps. A different model helps more.
But once different models disagree, somebody still has to decide which review to trust.
Rejudge makes that disagreement-resolution step explicit.
The architecture
The same review request goes to three reviewers:
reviewer A
reviewer B
reviewer C
Each works in an isolated context. They do not see each other's reasoning, tools, or conclusions. After all three finish, a separate judge receives their reports. The judge can ask follow-up questions when the panel disagrees. Then it writes one final answer.
same question
|
+--> reviewer A
+--> reviewer B
+--> reviewer C
|
judge
|
final answer
Independence happens before collaboration.
How reviewers inspect the code
By default, reviewer tools include:
read
grep
find
ls
git_diff
web_search (if the host provides one)
The normal reviewers do not get:
edit
write
bash
That is a sensible default for code review. 🔒
How judge makes a decision
The judge gets no workspace access. It sees the three reports.
If those reports conflict, it can call ask_panel and request clarification.
That means the judge is not a hidden fourth reviewer. Its role is:
compare
challenge
adjudicate
synthesize
Practical usage
Install (it needs Node.js 22.19.0 or newer):
npm install -g rejudge
Review a diff:
git diff | rejudge "review this change"
Ask a targeted question:
rejudge "does this migration need a lock?"
Resume later:
rejudge --resume <run-id> "what about the rollback path?"
The answer goes to stdout, while progress, the config in use and the run ID go to stderr, so redirecting stdout to a file leaves only the answer. A resumed run reopens the same sessions, and the new question goes to the judge first. The reviewers hear it only if the judge calls ask_panel.
Coding-agent integration
Rejudge supports:
- CLI
- native Pi tool
- Agent Skill for coding agents outside Pi
Agent Skill install:
npx skills add syabro/rejudge -g -y
The skills are a separate copy, so the README says to refresh them after each Rejudge release with npx skills update -g -y. Inside Pi, the extension is one more line:
pi install "$(npm root -g)/rejudge"
Rejudge runs on Pi and reads its provider settings, so a key Pi already accepts works here too.
The models are configurable
Rejudge is not tied to one fixed provider combination. The config has a reviewer list and a separate judge model, with two reviewers as the minimum. Each model also gets a reasoning level from:
minimal
low
medium
high
xhigh
The global file lives at ~/.config/rejudge/config.json, and a .rejudge/config.json in the project wins over it. That makes it possible to build a genuinely mixed panel.
There is an unsafe mode
--unsafe / --full gives reviewers:
edit
write
bash
The docs explicitly say this is not a sandbox. The judge still gets only ask_panel. For review-only work, I would stay read-only.
Multiple providers mean multiple privacy boundaries
Every reviewer model sees the request. Anything a reviewer reads becomes part of that provider session.
The judge does not inspect the workspace directly, but reviewer reports may quote code.
Read-only tools stop local changes. They do not keep file contents private, and the README warns that instructions hidden in a request or in a file a reviewer opens can steer what it reads and reports.
Runs also leave records. Sessions are written to ${TMPDIR}/rejudge/runs/<run-id>/ while a run executes, and cleanup after roughly 24 hours is best-effort. The debugLog option is off by default, and turning it on writes full model thinking to .rejudge/logs/.
For proprietary repositories, this needs to be acceptable before running the panel.
The cost
A fresh review starts with:
3 reviewer calls
+ 1 judge call
Then add tool loops, retries, recovery, and judge follow-ups. Rejudge has no spending cap of its own. This is not a free accuracy multiplier. It is a compute tradeoff. 💸
Where I would use it
I would reserve it for changes where mistakes are expensive:
database migrations
locking/concurrency
authentication
authorization
permissions
security-sensitive code
data deletion
rollback logic
complex refactors
Rejudge is a structured independent second opinion. For consequential code, that can be worth the extra cost.
References
- Rejudge on GitHub: README, install, config and the privacy and cost notes
- Rejudge site: overview and demo
Follow me for more on AI and Software Development:
khasky — LinkedIn / Patreon / GitHub / Bluesky / Mastodon
khaskydev — X / Threads / Instagram / Pinterest / Facebook
来源:Google AI:DEV 作者专属(RSS) · dev.to