In December I looked at what a year of AI coding assistants did to our pull requests. The conclusion was that authors got faster and reviewers didn’t, and that the review queue is where it would hurt first.
That happened faster than I expected. This quarter Claude Code went from “a few people use it” to the default way a lot of our engineers work, and our merged PRs per month went from about 2,500 in January to 3,900 in February and 4,400 in March. Same number of people.
Same people, more PRs
Two 14-day windows again, human-authored PRs only, across all our repositories:
| Jan 12-25 | Mar 15-28 | Change | |
|---|---|---|---|
| PRs merged | 1,111 | 1,558 | +40% |
| People who merged at least one PR | 97 | 97 | same |
| PRs per author per week | 5.7 | 8.0 | +40% |
| Median time to first human review | 25 min | 34 min | +36% |
| PRs reviewed by a bot | 28% | 65% | |
| People approving PRs | 80 | 80 | same |
| Approvals per approver | 14.4 | 19.3 | +34% |
| Approvals by the busiest approver | 145 | 184 | +27% |
Claude Code adds a Co-Authored-By trailer to commits by default, so I counted those too. The share of human commits on main with a Claude trailer went from 11% in January to 27% in February and 34% in March. That’s a lower bound, since Copilot and Cursor don’t leave a trailer and some people switch it off.
For the first time the review numbers moved the wrong way. Time to first review went from 25 to 34 minutes. That doesn’t sound like much, but it’s a median, and it moved within a single quarter with the same people on both sides.
The reviewer’s question changed
A year ago when I reviewed a PR, I could assume the author had written every line and could explain every line. The question in my head was “is this correct?”
Now a good chunk of the diff was typed by an agent, and the author’s job was to steer it and check it. Most authors do that well. Some days, honestly, nobody has read every line. So the reviewer’s question became “does anybody understand this change?”, and that is a harder question to answer from a diff.
At the same time a lot of review comments are boring and repetitive. Wrong error type. Transaction opened in the repository instead of the service. A DTO mapper where we pass proto types straight through. Missing config validation. A human shouldn’t spend their limited review attention on that, and at today’s volume they can’t.
So this quarter was about splitting review into the part a machine can check and the part only a human can.
Docs where the agent reads them
The first step was the least exciting one. If a convention only exists in someone’s head, the agent can’t follow it and the AI reviewer can’t check it.
In January we added AGENTS.md and CLAUDE.md files to our main repos, and we started a separate docs repo where an agent keeps the code documentation up to date after every merge. The proposal put it like this:
Instead of passing diffs to a single AI call, we use Claude Code CLI as an autonomous agent
The README of that repo opens with a quote I won’t reproduce in full, but it ends with “Mr. Claude, it’s too much.” That’s the spirit.
Review rules as skills
In late January we created a shared skills repo. Every convention we kept repeating in reviews became a small review-* skill: review-domain-errors, review-cls-transactions, review-connectrpc-only, review-no-dto-mappers, review-type-not-interface, review-config-validation, review-unit-tests and a few more. Thirteen at the end of January.
The skills come with two hooks for Claude Code. One runs after every edit and blocks on lint, type and test failures. The other runs when the agent wants to stop, and spawns one reviewer per principle in parallel against the changed files:
#!/bin/bash
# review-principles.sh - Spawns one Claude agent per principle/skill in parallel
#
# This hook runs on Stop (session end) and reviews all changed files
# against each skill defined in skills/review-*/SKILL.md
#
# Each principle gets its own agent for focused, accurate review.
One agent per principle sounds wasteful. The idea is that an agent looking for exactly one thing is more reliable than one agent juggling a list of thirteen rules.
An AI first pass on every PR
On the PR side we rolled out CodeRabbit in February. One of the first things we changed was turning off its auto-approve. From the commit message:
PRs should still require a human approver. CodeRabbit should review and flag issues but not count as an approving reviewer.
In March we moved its configuration into one central repo. Two parts of that config do most of the work. It reads our AGENTS.md and CLAUDE.md files as review guidelines, so the docs from step one double as review context. And our most common review comments are now pre-merge checks in error mode:
knowledge_base:
code_guidelines:
enabled: true
filePatterns:
- "**/AGENTS.md"
- "**/CLAUDE.md"
- "**/.claude/rules/"
- "**/.cursor/rules/"
# ...
reviews:
pre_merge_checks:
title:
mode: "error" # conventional commits with a ticket number
description:
mode: "error"
custom_checks:
- name: "Domain Errors Over Plain Errors"
mode: "error"
instructions: "Business logic must throw domain error classes (...) instead of plain Error ..."
- name: "CLS Transactions Over Manual db.transaction()"
mode: "error"
instructions: "Database transactions must use the CLS-based @Transactional() decorator ..."
- name: "Config Validation With Joi"
mode: "error"
instructions: "Any new or modified NestJS configuration registered via registerAs() must have a corresponding Joi validation schema ..."
(There’s also poem: true in there. We’re keeping it.)
By the end of March, CodeRabbit reviewed about 60% of human PRs.
What it didn’t fix
The numbers above are after all of this. Time to first review still went up, and the busiest reviewers still approve more every month. The AI first pass removes the boring comments. Someone still has to understand the change and own the approval.
What I take from this quarter:
- Generation is solved enough. The next tool worth building is one that helps us verify code.
- Conventions belong in the repo, written for an agent. The humans benefit just as much.
- A rule that a machine can check should never cost a human reviewer a minute.
- The approve button is the scarcest resource we have. We should treat it that way.
Next we want agents doing work in the background, not just in an engineer’s terminal. More on that soon.
As always, if you have any questions or remarks, feel free to ping me on twitter @bobby_donchev.