In 2025 a lot of our engineers started using an AI coding assistant. Copilot in the editor for most people, Cursor for some, and since September Claude Code in our biggest monorepo. Everybody felt faster. I wanted to know what actually changed, so I pulled every merged pull request from GitHub for two comparable two-week windows, one in January and one in October, and compared them.
Short version: the people writing code got a lot faster. The people reviewing it did not.
The numbers
Human-authored PRs merged across all our repositories, bots excluded, in two 14-day windows (January 12-25 and October 12-25):
| Jan 2025 | Oct 2025 | Change | |
|---|---|---|---|
| PRs merged | 375 | 844 | 2.3x |
| People who merged at least one PR | 57 | 85 | 1.5x |
| PRs per author per week | 3.3 | 5.0 | +52% |
| Median PR size (lines changed) | 47 | 82 | +74% |
| Median time to first human review | 29 min | 27 min | same |
| PRs with an AI review comment | 0% | 36% | |
| People approving PRs | 43 | 62 | 1.4x |
| Approvals per approver | 9.4 | 13.5 | +44% |
| Approvals by the busiest approver | 36 | 107 | 3x |
Monthly merged PRs (bots included this time) went from about 1,300 in January to about 2,300 in November.
Part of that is growth. The team grew by roughly half and we started a handful of new services. But per person we merged 50% more PRs a week, and those PRs were bigger. I can’t pin all of that on AI, but it is the shape I expected.
What we set up
Nothing fancy. Most of it was individual engineers picking up tools, and a few things we set up for the whole team.
In September we added Claude Code to our largest Go/Node monorepo. Anyone can mention @claude in a PR comment and it picks up the request:
name: Claude Code
on:
issue_comment:
types: [created]
pull_request_review_comment:
types: [created]
permissions:
contents: read
issues: write
pull-requests: write
jobs:
claude:
runs-on: ubuntu-latest
if: contains(github.event.comment.body, '@claude')
steps:
- uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
Next to it we wrote a CLAUDE.md with the title “Rules for Agents”. It is around 340 lines and reads like this:
**MUST** rules are enforced by CI; **SHOULD** rules are strongly recommended.
- **BP-1 (MUST)** Ask the user clarifying questions.
- **C-1 (MUST)** Follow TDD: scaffold stub -> write failing test -> implement.
- **C-7 (SHOULD NOT)** Add comments except for critical caveats; rely on self-explanatory code.
- **T-3 (MUST)** ALWAYS separate pure-logic unit tests from DB-touching integration tests.
- **D-1 (MUST)** Use SQLC for type-safe database access with proper Go structs.
Our newest service got Cursor rules in its repo. And by October, Copilot’s PR reviewer and Cursor’s review bot were leaving comments on about a third of our human PRs.
Writing rules for agents is writing rules for humans
The most useful thing that came out of this was the rules file, more than any code the assistants wrote.
We had conventions before. They lived in people’s heads, in old Slack threads and in review comments that got repeated every other week (“please don’t open a transaction in the repository”). Writing a CLAUDE.md forced us to put them in one place for the first time. I’d argue that file is as useful to a new engineer on their first day as it is to the agent.
The second lesson is in the first line of that file. A rule in a markdown file is a suggestion. The agent follows it most of the time and so do humans. If a rule really matters, a tool has to enforce it: a linter, a type check, a CI step. Our MUST rules are the ones we can check in CI, and in practice those are the ones that hold.
Review didn’t change, and that’s the problem
Look at the bottom of the table. Time to first review stayed flat at under half an hour, which sounds great. But the number of approvals per approver went up 44%, and the busiest approver went from 36 approvals in two weeks to 107. That’s more than ten a working day on top of their own work.
Authors scale with tools and hiring. Reviewers only scale with hiring. The AI review bots help a bit: they catch typos, a missing null check, the occasional real bug. They don’t approve anything, though, and they shouldn’t. Somebody still has to understand the change and press the button.
PRs also got bigger, 47 lines to 82 at the median. A bigger diff that the author partly didn’t type is also harder to review.
What I’m watching in 2026
If the assistants keep getting better at this pace (and it looks like they will), the volume on the writing side will keep growing and the review queue is where it will hurt first. So that’s where we want to put our effort next year:
- Put our documentation where the agents read it, and keep it up to date with an agent instead of hoping someone does it.
- Turn more of the review comments we keep repeating into rules a machine can check, so human reviewers can spend their time on the questions only a human can answer.
- Keep measuring. It took me an afternoon to pull these numbers and I learned more from them than from a year of “it feels faster”.
I’ll write a follow-up once we have a few months of data.
As always, if you have any questions or remarks, feel free to ping me on twitter @bobby_donchev.