Case study · 2026
A code-review harness that argues with itself
Single-pass reviews kept missing cross-cutting bugs. A set of specialist reviewers, each scoped to a dimension and forced to fact-check its own findings against the source, catches them without crying wolf.
Context
Single-pass code review, human or model, reliably misses cross-cutting concerns: a field added in one layer but not forwarded in the middleware, a redaction gap, an authorization guard that didn’t grow with a DTO. Each of those is obvious to a reviewer looking for exactly that thing and invisible to one reading the diff top to bottom.
The problem
More reviewers means more noise. A panel that manufactures low-severity findings to look thorough is worse than no panel, because it trains everyone to ignore it. The harness had to catch the cross-layer bugs and stay quiet when there was nothing real to say.
What I did
- Orchestrated specialist reviewers as parallel subagents, each scoped to one dimension — security, data propagation, redaction, boundaries — and to the files the change actually touches. Reviewers whose target paths are untouched simply don’t run.
- Made every reviewer fact-check its own claims against the cited source file and line before a finding is allowed out, and lower or drop anything uncertain.
- Routed findings to inline comments only when they were both high-severity and anchorable to a diff line. Everything else went to a summary, and “nothing material” returned an empty result on purpose.
- Shared large inputs — the diff, the rules — through a file the agents read, instead of stuffing them into every prompt.
The trade-off
False positives are worse than misses. I tuned the whole thing to be quiet and correct rather than exhaustive and noisy, which means it will occasionally stay silent on something a louder setup would have flagged. That is the price of being believed the rest of the time.
Outcome
Reviews that catch the cross-layer bugs a single pass misses, while staying credible because they never post a finding they can’t back with a file and a line.
What I’d do differently
Track precision over time — what fraction of posted findings were actually acted on — and feed that back into the severity thresholds instead of tuning them by feel.