In September we merged about 10,900 pull requests. In January it was about 2,500. Same people: then and now, roughly 97 people merge code and roughly 80 approve it. What changed is that most of them now work with an agent all day, and the agents got good enough that a typical engineer merges 10 PRs a week instead of 4.
Earlier this year I wrote that writing code got cheap and reviewing it didn’t. This post is about what we did once that stopped being a prediction.
The numbers
Human-authored PRs merged across all our repositories, bots excluded, in two 14-day windows:
| Jan 12-25 | Sep 7-20 | Change | |
|---|---|---|---|
| PRs merged | 1,111 | 3,683 | 3.3x |
| People who merged at least one PR | 97 | 97 | same |
| PRs per author per week | 5.7 | 19.0 | 3.3x |
| Median author, PRs per week | 4.0 | 10.5 | 2.6x |
| Median PR size (lines changed) | 52 | 145 | 2.8x |
| Median time to first human review | 25 min | 48 min | 1.9x |
| People approving PRs | 80 | 80 | same |
| Approvals per approver | 14.4 | 41.2 | 2.9x |
| Approvals by the busiest approver | 145 | 263 | 1.8x |

Two things to get out of the way. First, this growth is our engineers and their own agents (Claude Code, Codex, Cursor). Cody, our background agent platform, authored about 2% of September’s merged PRs. Second, the Claude Co-Authored-By trailer is on roughly half of the human commits in our longest-lived repos now, and that’s a lower bound.
The row I keep looking at is the busiest approver: 263 approvals in two weeks, about 26 per working day, on PRs that are almost three times bigger than in January. That’s a lot of careful reading to ask of one person on top of their own work. When volume grows 3x and the number of reviewers stays flat, something has to give.
So the question for this half of the year was: what does a human reviewer need to look at, and what can we verify without them?
Write it down, then enforce it
The first answer was boring and it came from the earlier posts. Everything an agent or a reviewer needs to know has to be in the repo. In August we turned our workspace of about 90 repos into something an agent can navigate: one short map file at the root (a check fails if it grows over 130 lines), a docs folder, and a page of core beliefs, the invariants every repo has to respect. Two lines from that page describe the approach:
Prefer strengthening the enforcer over writing more prose.
Knowledge that only exists in Slack/Docs/heads does not exist for an agent.
The second line I already believed last year. The first one is what this post is about.
Quality walls
Our earlier review skills and CodeRabbit checks were good at conventions: wrong error type, missing validation, a transaction in the wrong layer. What they didn’t do well was the harder stuff: is this change actually tested, can this money calculation lose precision, can a tenant see another tenant’s data.
Quality walls are our answer. A wall is a rule with a structure an agent can investigate. This is one of the 92 generic ones:
{
"id": "QW-GLOBAL-042",
"title": "Not Enough Tests",
"check": "Whether changed behavior, branches, failures, or invariants have evidence in executable tests. Trace the changed outcome through existing tests at the owning boundary ...",
"reject_when": "A concrete new or changed observable outcome can regress without any focused test failing. Identify the changed outcome and a specific regression that existing assertions would miss ...",
"severity": "major",
"remediation": "Add the smallest tests that prove the missing observable outcomes.",
"exclusions": "Generated code, declarative configuration, behavior already proven at a more appropriate boundary ..."
}
There are 222 rules in the central catalog today: generic ones like this, plus modules for authorization, audit logging, Postgres queries, decimal math, data ingestion and so on. Repos can add their own local rules on top. They can’t disable central ones.
The most important part of the design doc is the list of what they are not:
They are not a linter, a replacement for tests, or another PR-review persona. The evaluator uses each rule as an investigation lead, verifies the actual code path and consequence, and reports only concrete violations.
“A rule is a lead, and a finding requires a verified consequence.” A wall can’t fail a PR because a line looks suspicious. It has to show the regression that would get through.
A few other decisions I like:
- A PR can’t change the rules that evaluate it. The policy comes from the merge base, so the head can’t loosen a rule while it’s being reviewed. A rule change only becomes active after it merges.
- The same rules are meant to run on
maintoo. The design adds a weekly repair run that checks everything that landed since the last run and opens at most one fix PR, which a human merges. - Local rules graduate. The ones that prove useful get promoted into the central catalog through a normal PR. Both of these loops are designed, but only partly rolled out so far.
- Waivers are explicit:
/cody waive <rule> <reason>, and only a code owner who isn’t the PR author can do it.
And the most honest paragraph in the whole doc:
Evaluation completeness and judgment correctness are different guarantees. […] It does not prove that an LLM’s judgment was correct. A defect class becomes deterministic only when it has an executable detector.
The agent decides, Cody owns the check
Quality walls and PR review run as Cody lanes, and each lane reports a GitHub Check Run. Early on we had a nasty bug in that, straight from the spec:
A successful Codex turn therefore becomes a successful Check Run even when the PR-review skill found blocking defects
The agent could do its job, find real problems, and the check would still go green, because “the agent run succeeded” and “the review passed” were the same signal. The fix was to split the responsibility:
The agent decides whether the review passed, failed, or was inconclusive. Cody owns the corresponding Check Run from deterministic registration through verified terminal publication.
The agent returns one bounded result tied to the commit it reviewed. Cody maps it to a check conclusion, re-reads the PR head, updates the exact Check Run and reads it back before it calls it done:
| Agent result | Check conclusion |
|---|---|
pass |
success |
fail |
failure |
inconclusive, missing, partial, cancelled |
action required |
| PR head moved on since the review | neutral |
Everything that isn’t a clear pass fails closed. A model that times out doesn’t get to approve anything.
Two smaller rules came out of the same thinking. Lanes that only apply to some paths (a docs review, say) still always publish their check, as a successful “Not applicable” for that exact commit, and if Cody can’t read the changed files, that fails closed too. And a stale blocking review can only be cleared when a person, not a bot, asks for a re-review on the same commit and Cody approves it. In the words of the requirement: “Only Cody’s own verdict can be replaced, and only on the head it reviewed.”
One policy file, about 100 repos
Between September 27 and October 1 we made these checks required on main: Cody PR Review, Cody Quality Walls, and the requirements and specs reviews for PRs that touch those documents. The branch rulesets are generated from one fleet file in our skills repo, which lists every repo and its lanes. New repos get five lanes automatically. Changing the policy for 100 repos is one reviewed PR.
The PR description is evidence
As the diffs got bigger and less hand-written, the description became more important than the diff. Every PR now uses the same body, and a quality wall checks it:
## Summary
<what changed and why, in one or two sentences>
## Evidence
- **Before:** <screenshot/output/failing test run>
- **After:** <screenshot/output/passing test run, or per-acceptance-criterion evidence>
## Merge Danger
**Door:** one-way or two-way, and how to walk it back or recover
**Blast Radius:** <affected services, consumers, or environments, and what could break>
Evidence has to be something that was actually observed on the current head commit; the skill says to never fabricate a failing run. Merge Danger is the part I’d steal for any team. A two-way door (revert and redeploy) deserves a different review than a one-way door (a migration that drops a column). It tells the reviewer where to spend their attention.
Shift left, but CI wins
Waiting for CI to tell you about a quality wall you could have caught locally is a waste. Since October 1 our shared Claude Code plugin ships a local precheck that runs the same PR review rubric and the same quality walls, pinned to the same versions CI uses. A SessionStart hook installs a git pre-push guard, and a PreToolUse hook stops the agent from running git push or gh pr create without a fresh precheck. This is what engineers see when they try anyway:
Cody local precheck: Run /cody-precheck, fix findings and commit, then rerun.
Explicit waiver: /cody-precheck --skip "reason".
It’s deliberately easy to skip. From the skill: “Local records are advisory and editable; CI remains authoritative.” The local check is there to save time. CI makes the final call.
What we don’t know yet
Plenty. A few things I want to be upfront about:
- We can’t yet measure how good the reviewers are. We have labeled eval corpora for quality walls and simulated PR scenarios for PR review, and the eval README says it plainly: “No live baseline exists yet.” My favorite line in that README is “Do not weaken an expectation or relabel a fixture merely to improve a score.”
- Disputes are a signal, not a rate. Every night we collect findings that engineers pushed back on (“inapplicable”, “pre-existing”). The most disputed rule so far is the one about PR descriptions, which tells you something about engineers too.
- The review machinery became its own bottleneck. In September we had reviewer pods waiting for CPU and third-party review rate limits holding up merges. When review is required, review capacity is production capacity, and it has to be run like production.
- Reverts didn’t go up. They stayed at or below about half a percent of merged PRs in every window I sampled this year. That’s good, but it isn’t proof that quality held. It only shows that nothing obviously broke.
Where the bottleneck is now
A year ago the slow part of shipping was writing the code. In the spring it was reviewing it. Now it is verifying it: proving with evidence that a change does what it says and doesn’t do what it shouldn’t. That’s a better place for the bottleneck to be. Verification can be automated piece by piece, written down as rules, measured, and argued about in PRs. “Somebody senior looked at it” can’t.
As always, if you have any questions or remarks, feel free to ping me on twitter @bobby_donchev.