Review Is the New Bottleneck: My Protocol for AI-Generated Pull Requests
Agents can write more code than any team can review. Here is the concrete PR protocol I use - diff limits, blast-radius tiers, test-tamper detection, and a 12-minute human read that actually catches bugs.
Pavel Duglas
AI Automation & MVP Architect
Writing code stopped being the constraint about a year ago. On a normal week I now produce three to five times more diff than I did in 2023, and none of that is a brag - it is a problem. Because the review capacity of one human brain did not multiply by five. Mine is still the same brain, running on Vietnamese iced coffee and a 40-minute attention span.
This is the failure mode I see in every team that adopted coding agents seriously: the queue moved. It used to be “we need more hands to build.” Now it is “we have 14 open PRs, half of them touch the billing module, and nobody actually read them.” Merging unread agent code is how you get a production incident that nobody on the team can debug, because nobody wrote it and nobody read it.
So I built a protocol. Not a philosophy - a set of mechanical rules that my repos and my CI enforce. Here it is.
Rule 1: cap the diff, mechanically
An agent will happily hand you a 2,400-line PR that renames three things, refactors a service, adds a feature, and “cleans up” a config file. Nobody reviews that. People skim it and approve it.
My hard limits:
- 400 changed lines per PR, excluding lockfiles, generated code, and snapshots
- one intent per PR: feature, refactor, or dependency bump - never mixed
- no drive-by formatting in a logic PR
I enforce the line cap in CI, not in a wiki nobody reads:
# .github/workflows/pr-size.yml (core step)
CHANGED=$(git diff --numstat origin/main...HEAD \
-- . ':(exclude)*.lock' ':(exclude)**/generated/**' ':(exclude)**/__snapshots__/**' \
| awk '{a+=$1; d+=$2} END {print a+d}')
if [ "$CHANGED" -gt 400 ]; then
echo "PR is $CHANGED lines. Split it or add the label 'oversized-approved'."
exit 1
fi
The escape hatch label matters. Big mechanical migrations exist. But the default has to be friction, otherwise the agent’s natural output size wins.
The practical consequence: I now tell the agent the constraint up front. “Implement only the repository layer for this feature. Do not touch the HTTP handlers. Target under 300 lines.” You get better code and a reviewable diff for the same effort.
Rule 2: review the plan, not the diff
The cheapest review happens before any code exists. For anything nontrivial I make the agent produce a plan file first: files it will touch, functions it will add, data it will read and write, migrations, and the failure modes it is explicitly not handling.
I read that in three minutes. Ninety percent of bad PRs die at this stage, because the wrong plan is obvious in a way that wrong code is not. “You are going to add a new users_meta table” - no, we already have a JSONB column for that. “You will call the payment provider inside the request handler” - no, that goes in the job queue.
Once the plan is approved, the diff review becomes a comparison instead of an investigation. I am not asking “what does this do?” I am asking “does this match what we agreed?” That is a fundamentally faster mental operation.
Rule 3: tier by blast radius, not by line count
A 600-line change to a marketing landing page is less dangerous than an 8-line change to a token refresh function. So I stopped grading PRs by size and started grading them by what breaks if the code is wrong.
I keep an explicit map in the repo:
# review-tiers.yml
tier_1_read_every_line:
- src/billing/**
- src/auth/**
- migrations/**
- infra/**
- src/**/webhooks/**
tier_2_read_the_logic:
- src/api/**
- src/jobs/**
tier_3_trust_the_tests:
- src/ui/**
- content/**
- scripts/one-off/**
Tier 1: I read every line, out loud in my head, and I run it locally. No agent merges tier 1 code without a human who can explain it a month later.
Tier 2: I read the logic and the error paths, skim the rest, and rely on integration tests.
Tier 3: if CI is green and the preview deploy looks right, ship it.
This is the single change that gave me the most review capacity back. Most agent output is tier 3. Trying to review it with tier 1 rigor was burning the attention I needed for the code that could actually cost money.
Rule 4: tests are the spec, and test edits are suspicious
The most dangerous thing an agent does is make the suite green by changing the suite. I have watched an agent “fix” a failing assertion by loosening it from an exact value to expect.anything(). Technically green. Functionally a lie.
Two countermeasures.
First, for anything in tier 1 or tier 2, I write the test before the agent writes the implementation. Not a full TDD religion - just the two or three assertions that encode the business rule. “A refund above the original charge amount must throw.” That test is my contract, and it lives in a file the agent is instructed not to modify.
Second, CI flags test-file edits so they cannot slip past in a big diff:
TEST_DELTA=$(git diff --numstat origin/main...HEAD -- '**/*.test.ts' | awk '{d+=$2} END {print d+0}')
if [ "$TEST_DELTA" -gt 0 ]; then
echo "::warning::This PR deletes $TEST_DELTA lines of test code. Manual sign-off required."
fi
Deleted test lines are the signal. Added tests are usually fine. Removed assertions are where the trap lives. I read every single deleted test line, always, regardless of tier.
Rule 5: never spend human attention on a machine-checkable property
If I catch myself writing a review comment about formatting, naming conventions, import order, missing types, an unhandled promise, or a console.log left in, I have wasted a slot in my working memory. Those all belong in tooling.
My minimum gate before a human sees a PR:
- formatter and linter, with autofix off in CI so it fails instead of silently rewriting
- strict type checking, no
anyescapes in tier 1 and 2 paths - unit plus integration tests, with the DB running for real, not mocked
- a dependency diff check: any new package in the lockfile needs a one-line justification in the PR body
- a secrets scan, because agents love to inline a key “temporarily”
That last one is not paranoia. A coding agent’s instinct when a config lookup fails is to hardcode the value that makes the test pass.
Rule 6: the 12-minute read
When a PR reaches me, I do a fixed-order pass and I time it. If I cannot finish in about 12 minutes, the PR is too big and goes back for splitting.
- Read the PR description and compare against the approved plan. Mismatch means stop.
- Read the data layer changes: migrations, queries, schemas. This is where irreversible damage lives.
- Read the error paths. Agents write beautiful happy paths and forget that the network exists. What happens on timeout, on a 429, on a partial write?
- Read every deleted line. Deletions are where behavior disappears silently.
- Grep for the boundaries: any new network call, any new write, any new env var, any new background job.
- Run it. Actually click the thing. Two minutes of using the feature beats twenty minutes of reading it.
Notice that reading the new business logic line by line is not step one. It is the part the tests and types cover best.
The metrics that tell you the protocol works
I track four numbers per month, and they are cheap to pull from Git and your incident log:
- median PR size in reviewable lines - should trend down
- median time from open to merge - should trend down
- revert rate and post-merge hotfix rate - the real quality signal
- share of merged PRs where no human left a substantive comment - if this goes above roughly a third, you are rubber-stamping
That last metric is the honest one. Approval velocity feels like productivity right up until the week you spend three days debugging a system nobody in the room understands.
The solo founder version
If you are one person shipping an MVP, you do not need CODEOWNERS and a review rota. You need three things:
- A plan file per feature that you actually read before the agent codes.
- A tier 1 list of five to ten files that you refuse to let an agent touch unreviewed. Usually: auth, payments, migrations, the webhook handler, the deploy config.
- A green CI that includes at least one integration test hitting a real database.
Everything else can be trusted to tests and rollbacks. The point is not to review more. It is to decide in advance where your attention is worth more than your tooling, and then protect that attention ruthlessly.
The teams struggling right now are not the ones whose agents write bad code. Agent code is fine, mostly. They are struggling because they scaled generation by 5x and left review at 1x, and then acted surprised when the queue backed up. Fix the review pipeline first. The generation side already works.
FAQ
Should I let an AI agent review another agent's pull request?
As a first pass, yes - it is good at the mechanical stuff a linter cannot express, like spotting an unhandled error branch or a query that will scan a whole table. But treat it as a stronger linter, not as an approver. It shares blind spots with the agent that wrote the code, and it has no memory of why your system is shaped the way it is. Keep a human as the final sign-off on anything in your tier 1 paths: auth, billing, migrations, infrastructure.
How do I stop an agent from weakening tests to make CI pass?
Two things work. Write the critical assertions yourself before the agent implements the feature, and tell it those files are off limits. Then add a CI step that counts deleted lines in test files and fails or warns when the number is above zero, so the deletion cannot hide inside a large diff. Reviewing every removed test line takes a minute and catches the single most expensive category of silent regression.
Is a 400-line PR limit realistic for real feature work?
For hand-written code it can feel tight. For agent-assisted work it is easy, because splitting is nearly free - you just tell the agent to implement one layer at a time. Data layer, then service layer, then handlers, then UI. You get four reviewable PRs instead of one unreadable one, and each merges the same day. Keep a labeled escape hatch for genuine mechanical migrations like a framework upgrade, and require the diff there to be verifiably mechanical.
Related articles
Done for you
I will turn your vibe-coded prototype into a working product
I will review what the AI generated, close the security and data gaps and ship it to production.
from $1,500 · 1 to 2 weeks