<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>Engineering on donchev.is</title>
        <link>https://donchev.is/tags/engineering/</link>
        <description>Recent content in Engineering on donchev.is</description>
        <generator>Hugo -- gohugo.io</generator>
        <language>en-us</language>
        <lastBuildDate>Thu, 08 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://donchev.is/tags/engineering/index.xml" rel="self" type="application/rss+xml" /><item>
        <title>Quality walls: the bottleneck is verification</title>
        <link>https://donchev.is/post/quality-walls/</link>
        <pubDate>Thu, 08 Oct 2026 00:00:00 +0000</pubDate>
        
        <guid>https://donchev.is/post/quality-walls/</guid>
        <description>&lt;img src="https://donchev.is/post/2026/quality-walls.png" alt="Featured image of post Quality walls: the bottleneck is verification" /&gt;&lt;p&gt;In September we merged about 10,900 pull requests. In January it was about 2,500. Same people: then and now, roughly 97 people merge code and roughly 80 approve it. What changed is that most of them now work with an agent all day, and the agents got good enough that a typical engineer merges 10 PRs a week instead of 4.&lt;/p&gt;
&lt;p&gt;Earlier this year I wrote that &lt;a class=&#34;link&#34; href=&#34;https://donchev.is/post/writing-code-got-cheap-reviewing-it-didnt/&#34; &gt;writing code got cheap and reviewing it didn&amp;rsquo;t&lt;/a&gt;. This post is about what we did once that stopped being a prediction.&lt;/p&gt;
&lt;h2 id=&#34;the-numbers&#34;&gt;The numbers&lt;/h2&gt;
&lt;p&gt;Human-authored PRs merged across all our repositories, bots excluded, in two 14-day windows:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th style=&#34;text-align:right&#34;&gt;Jan 12-25&lt;/th&gt;
&lt;th style=&#34;text-align:right&#34;&gt;Sep 7-20&lt;/th&gt;
&lt;th style=&#34;text-align:right&#34;&gt;Change&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;PRs merged&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;1,111&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;3,683&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;3.3x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;People who merged at least one PR&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;97&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;97&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;same&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PRs per author per week&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;5.7&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;19.0&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;3.3x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median author, PRs per week&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;4.0&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;10.5&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;2.6x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median PR size (lines changed)&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;52&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;145&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;2.8x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median time to first human review&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;25 min&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;48 min&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;1.9x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;People approving PRs&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;80&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;80&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;same&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Approvals per approver&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;14.4&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;41.2&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;2.9x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Approvals by the busiest approver&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;145&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;263&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;1.8x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;img src=&#34;https://donchev.is/post/2026/review-load.png&#34;
	
	
	
	loading=&#34;lazy&#34;
	
		alt=&#34;Output per author grew 3.3x since January while the number of approvers stayed flat&#34;
	
	
&gt;&lt;/p&gt;
&lt;p&gt;Two things to get out of the way. First, this growth is our engineers and their own agents (Claude Code, Codex, Cursor). Cody, our &lt;a class=&#34;link&#34; href=&#34;https://donchev.is/post/meet-cody/&#34; &gt;background agent platform&lt;/a&gt;, authored about 2% of September&amp;rsquo;s merged PRs. Second, the Claude &lt;code&gt;Co-Authored-By&lt;/code&gt; trailer is on roughly half of the human commits in our longest-lived repos now, and that&amp;rsquo;s a lower bound.&lt;/p&gt;
&lt;p&gt;The row I keep looking at is the busiest approver: 263 approvals in two weeks, about 26 per working day, on PRs that are almost three times bigger than in January. That&amp;rsquo;s a lot of careful reading to ask of one person on top of their own work. When volume grows 3x and the number of reviewers stays flat, something has to give.&lt;/p&gt;
&lt;p&gt;So the question for this half of the year was: what does a human reviewer need to look at, and what can we verify without them?&lt;/p&gt;
&lt;h2 id=&#34;write-it-down-then-enforce-it&#34;&gt;Write it down, then enforce it&lt;/h2&gt;
&lt;p&gt;The first answer was boring and it came from the earlier posts. Everything an agent or a reviewer needs to know has to be in the repo. In August we turned our workspace of about 90 repos into something an agent can navigate: one short map file at the root (a check fails if it grows over 130 lines), a docs folder, and a page of core beliefs, the invariants every repo has to respect. Two lines from that page describe the approach:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Prefer strengthening the enforcer over writing more prose.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;Knowledge that only exists in Slack/Docs/heads does not exist for an agent.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The second line I already believed last year. The first one is what this post is about.&lt;/p&gt;
&lt;h2 id=&#34;quality-walls&#34;&gt;Quality walls&lt;/h2&gt;
&lt;p&gt;Our earlier review skills and CodeRabbit checks were good at conventions: wrong error type, missing validation, a transaction in the wrong layer. What they didn&amp;rsquo;t do well was the harder stuff: is this change actually tested, can this money calculation lose precision, can a tenant see another tenant&amp;rsquo;s data.&lt;/p&gt;
&lt;p&gt;Quality walls are our answer. A wall is a rule with a structure an agent can investigate. This is one of the 92 generic ones:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-json&#34; data-lang=&#34;json&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;p&#34;&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nt&#34;&gt;&amp;#34;id&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt; &lt;span class=&#34;s2&#34;&gt;&amp;#34;QW-GLOBAL-042&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nt&#34;&gt;&amp;#34;title&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt; &lt;span class=&#34;s2&#34;&gt;&amp;#34;Not Enough Tests&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nt&#34;&gt;&amp;#34;check&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt; &lt;span class=&#34;s2&#34;&gt;&amp;#34;Whether changed behavior, branches, failures, or invariants have evidence in executable tests. Trace the changed outcome through existing tests at the owning boundary ...&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nt&#34;&gt;&amp;#34;reject_when&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt; &lt;span class=&#34;s2&#34;&gt;&amp;#34;A concrete new or changed observable outcome can regress without any focused test failing. Identify the changed outcome and a specific regression that existing assertions would miss ...&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nt&#34;&gt;&amp;#34;severity&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt; &lt;span class=&#34;s2&#34;&gt;&amp;#34;major&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nt&#34;&gt;&amp;#34;remediation&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt; &lt;span class=&#34;s2&#34;&gt;&amp;#34;Add the smallest tests that prove the missing observable outcomes.&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nt&#34;&gt;&amp;#34;exclusions&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt; &lt;span class=&#34;s2&#34;&gt;&amp;#34;Generated code, declarative configuration, behavior already proven at a more appropriate boundary ...&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;p&#34;&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;There are 222 rules in the central catalog today: generic ones like this, plus modules for authorization, audit logging, Postgres queries, decimal math, data ingestion and so on. Repos can add their own local rules on top. They can&amp;rsquo;t disable central ones.&lt;/p&gt;
&lt;p&gt;The most important part of the design doc is the list of what they are not:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;They are not a linter, a replacement for tests, or another PR-review persona. The evaluator uses each rule as an investigation lead, verifies the actual code path and consequence, and reports only concrete violations.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&amp;ldquo;A rule is a lead, and a finding requires a verified consequence.&amp;rdquo; A wall can&amp;rsquo;t fail a PR because a line &lt;em&gt;looks&lt;/em&gt; suspicious. It has to show the regression that would get through.&lt;/p&gt;
&lt;p&gt;A few other decisions I like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A PR can&amp;rsquo;t change the rules that evaluate it. The policy comes from the merge base, so the head can&amp;rsquo;t loosen a rule while it&amp;rsquo;s being reviewed. A rule change only becomes active after it merges.&lt;/li&gt;
&lt;li&gt;The same rules are meant to run on &lt;code&gt;main&lt;/code&gt; too. The design adds a weekly repair run that checks everything that landed since the last run and opens at most one fix PR, which a human merges.&lt;/li&gt;
&lt;li&gt;Local rules graduate. The ones that prove useful get promoted into the central catalog through a normal PR. Both of these loops are designed, but only partly rolled out so far.&lt;/li&gt;
&lt;li&gt;Waivers are explicit: &lt;code&gt;/cody waive &amp;lt;rule&amp;gt; &amp;lt;reason&amp;gt;&lt;/code&gt;, and only a code owner who isn&amp;rsquo;t the PR author can do it.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And the most honest paragraph in the whole doc:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Evaluation completeness and judgment correctness are different guarantees. [&amp;hellip;] It does not prove that an LLM&amp;rsquo;s judgment was correct. A defect class becomes deterministic only when it has an executable detector.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&#34;the-agent-decides-cody-owns-the-check&#34;&gt;The agent decides, Cody owns the check&lt;/h2&gt;
&lt;p&gt;Quality walls and PR review run as Cody lanes, and each lane reports a GitHub Check Run. Early on we had a nasty bug in that, straight from the spec:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;A successful Codex turn therefore becomes a successful Check Run even when the PR-review skill found blocking defects&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The agent could do its job, find real problems, and the check would still go green, because &amp;ldquo;the agent run succeeded&amp;rdquo; and &amp;ldquo;the review passed&amp;rdquo; were the same signal. The fix was to split the responsibility:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The agent decides whether the review passed, failed, or was inconclusive. Cody owns the corresponding Check Run from deterministic registration through verified terminal publication.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The agent returns one bounded result tied to the commit it reviewed. Cody maps it to a check conclusion, re-reads the PR head, updates the exact Check Run and reads it back before it calls it done:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Agent result&lt;/th&gt;
&lt;th&gt;Check conclusion&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pass&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;success&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;fail&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;failure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;inconclusive&lt;/code&gt;, missing, partial, cancelled&lt;/td&gt;
&lt;td&gt;action required&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PR head moved on since the review&lt;/td&gt;
&lt;td&gt;neutral&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Everything that isn&amp;rsquo;t a clear pass fails closed. A model that times out doesn&amp;rsquo;t get to approve anything.&lt;/p&gt;
&lt;p&gt;Two smaller rules came out of the same thinking. Lanes that only apply to some paths (a docs review, say) still always publish their check, as a successful &amp;ldquo;Not applicable&amp;rdquo; for that exact commit, and if Cody can&amp;rsquo;t read the changed files, that fails closed too. And a stale blocking review can only be cleared when a person, not a bot, asks for a re-review on the same commit and Cody approves it. In the words of the requirement: &amp;ldquo;Only Cody&amp;rsquo;s own verdict can be replaced, and only on the head it reviewed.&amp;rdquo;&lt;/p&gt;
&lt;h2 id=&#34;one-policy-file-about-100-repos&#34;&gt;One policy file, about 100 repos&lt;/h2&gt;
&lt;p&gt;Between September 27 and October 1 we made these checks required on &lt;code&gt;main&lt;/code&gt;: Cody PR Review, Cody Quality Walls, and the requirements and specs reviews for PRs that touch those documents. The branch rulesets are generated from one fleet file in our skills repo, which lists every repo and its lanes. New repos get five lanes automatically. Changing the policy for 100 repos is one reviewed PR.&lt;/p&gt;
&lt;h2 id=&#34;the-pr-description-is-evidence&#34;&gt;The PR description is evidence&lt;/h2&gt;
&lt;p&gt;As the diffs got bigger and less hand-written, the description became more important than the diff. Every PR now uses the same body, and a quality wall checks it:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-markdown&#34; data-lang=&#34;markdown&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;gu&#34;&gt;## Summary
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;gu&#34;&gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;p&#34;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;what&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;changed&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;and&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;why&lt;/span&gt;&lt;span class=&#34;err&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;in&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;one&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;or&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;two&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;sentences&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;gu&#34;&gt;## Evidence
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;gu&#34;&gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;-&lt;/span&gt; **Before:** &lt;span class=&#34;p&#34;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;screenshot&lt;/span&gt;&lt;span class=&#34;err&#34;&gt;/&lt;/span&gt;&lt;span class=&#34;na&#34;&gt;output&lt;/span&gt;&lt;span class=&#34;err&#34;&gt;/&lt;/span&gt;&lt;span class=&#34;na&#34;&gt;failing&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;test&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;run&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;-&lt;/span&gt; **After:** &lt;span class=&#34;p&#34;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;screenshot&lt;/span&gt;&lt;span class=&#34;err&#34;&gt;/&lt;/span&gt;&lt;span class=&#34;na&#34;&gt;output&lt;/span&gt;&lt;span class=&#34;err&#34;&gt;/&lt;/span&gt;&lt;span class=&#34;na&#34;&gt;passing&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;test&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;run&lt;/span&gt;&lt;span class=&#34;err&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;or&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;per-acceptance-criterion&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;evidence&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;gu&#34;&gt;## Merge Danger
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;gu&#34;&gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;gs&#34;&gt;**Door:**&lt;/span&gt; one-way or two-way, and how to walk it back or recover
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;gs&#34;&gt;**Blast Radius:**&lt;/span&gt; &lt;span class=&#34;p&#34;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;affected&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;services&lt;/span&gt;&lt;span class=&#34;err&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;consumers&lt;/span&gt;&lt;span class=&#34;err&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;or&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;environments&lt;/span&gt;&lt;span class=&#34;err&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;and&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;what&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;could&lt;/span&gt; &lt;span class=&#34;na&#34;&gt;break&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Evidence has to be something that was actually observed on the current head commit; the skill says to never fabricate a failing run. Merge Danger is the part I&amp;rsquo;d steal for any team. A two-way door (revert and redeploy) deserves a different review than a one-way door (a migration that drops a column). It tells the reviewer where to spend their attention.&lt;/p&gt;
&lt;h2 id=&#34;shift-left-but-ci-wins&#34;&gt;Shift left, but CI wins&lt;/h2&gt;
&lt;p&gt;Waiting for CI to tell you about a quality wall you could have caught locally is a waste. Since October 1 our shared Claude Code plugin ships a local precheck that runs the same PR review rubric and the same quality walls, pinned to the same versions CI uses. A &lt;code&gt;SessionStart&lt;/code&gt; hook installs a git pre-push guard, and a &lt;code&gt;PreToolUse&lt;/code&gt; hook stops the agent from running &lt;code&gt;git push&lt;/code&gt; or &lt;code&gt;gh pr create&lt;/code&gt; without a fresh precheck. This is what engineers see when they try anyway:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Cody local precheck: Run /cody-precheck, fix findings and commit, then rerun.
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Explicit waiver: /cody-precheck --skip &amp;#34;reason&amp;#34;.
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It&amp;rsquo;s deliberately easy to skip. From the skill: &amp;ldquo;Local records are advisory and editable; CI remains authoritative.&amp;rdquo; The local check is there to save time. CI makes the final call.&lt;/p&gt;
&lt;h2 id=&#34;what-we-dont-know-yet&#34;&gt;What we don&amp;rsquo;t know yet&lt;/h2&gt;
&lt;p&gt;Plenty. A few things I want to be upfront about:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;We can&amp;rsquo;t yet measure how good the reviewers are. We have labeled eval corpora for quality walls and simulated PR scenarios for PR review, and the eval README says it plainly: &amp;ldquo;No live baseline exists yet.&amp;rdquo; My favorite line in that README is &amp;ldquo;Do not weaken an expectation or relabel a fixture merely to improve a score.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Disputes are a signal, not a rate. Every night we collect findings that engineers pushed back on (&amp;ldquo;inapplicable&amp;rdquo;, &amp;ldquo;pre-existing&amp;rdquo;). The most disputed rule so far is the one about PR descriptions, which tells you something about engineers too.&lt;/li&gt;
&lt;li&gt;The review machinery became its own bottleneck. In September we had reviewer pods waiting for CPU and third-party review rate limits holding up merges. When review is required, review capacity is production capacity, and it has to be run like production.&lt;/li&gt;
&lt;li&gt;Reverts didn&amp;rsquo;t go up. They stayed at or below about half a percent of merged PRs in every window I sampled this year. That&amp;rsquo;s good, but it isn&amp;rsquo;t proof that quality held. It only shows that nothing obviously broke.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;where-the-bottleneck-is-now&#34;&gt;Where the bottleneck is now&lt;/h2&gt;
&lt;p&gt;A year ago the slow part of shipping was writing the code. In the spring it was reviewing it. Now it is &lt;em&gt;verifying&lt;/em&gt; it: proving with evidence that a change does what it says and doesn&amp;rsquo;t do what it shouldn&amp;rsquo;t. That&amp;rsquo;s a better place for the bottleneck to be. Verification can be automated piece by piece, written down as rules, measured, and argued about in PRs. &amp;ldquo;Somebody senior looked at it&amp;rdquo; can&amp;rsquo;t.&lt;/p&gt;
&lt;p&gt;As always, if you have any questions or remarks, feel free to ping me on twitter &lt;a class=&#34;link&#34; href=&#34;https://twitter.com/bobby_donchev&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;@bobby_donchev&lt;/a&gt;.&lt;/p&gt;
</description>
        </item>
        <item>
        <title>Writing code got cheap. Reviewing it didn&#39;t.</title>
        <link>https://donchev.is/post/writing-code-got-cheap-reviewing-it-didnt/</link>
        <pubDate>Sun, 05 Apr 2026 00:00:00 +0000</pubDate>
        
        <guid>https://donchev.is/post/writing-code-got-cheap-reviewing-it-didnt/</guid>
        <description>&lt;img src="https://donchev.is/post/2026/writing-code-got-cheap.png" alt="Featured image of post Writing code got cheap. Reviewing it didn&#39;t." /&gt;&lt;p&gt;In December I &lt;a class=&#34;link&#34; href=&#34;https://donchev.is/post/a-year-of-ai-coding-assistants/&#34; &gt;looked at what a year of AI coding assistants did to our pull requests&lt;/a&gt;. The conclusion was that authors got faster and reviewers didn&amp;rsquo;t, and that the review queue is where it would hurt first.&lt;/p&gt;
&lt;p&gt;That happened faster than I expected. This quarter Claude Code went from &amp;ldquo;a few people use it&amp;rdquo; to the default way a lot of our engineers work, and our merged PRs per month went from about 2,500 in January to 3,900 in February and 4,400 in March. Same number of people.&lt;/p&gt;
&lt;h2 id=&#34;same-people-more-prs&#34;&gt;Same people, more PRs&lt;/h2&gt;
&lt;p&gt;Two 14-day windows again, human-authored PRs only, across all our repositories:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th style=&#34;text-align:right&#34;&gt;Jan 12-25&lt;/th&gt;
&lt;th style=&#34;text-align:right&#34;&gt;Mar 15-28&lt;/th&gt;
&lt;th style=&#34;text-align:right&#34;&gt;Change&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;PRs merged&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;1,111&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;1,558&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;+40%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;People who merged at least one PR&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;97&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;97&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;same&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PRs per author per week&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;5.7&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;8.0&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;+40%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median time to first human review&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;25 min&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;34 min&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;+36%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PRs reviewed by a bot&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;28%&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;65%&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;People approving PRs&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;80&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;80&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;same&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Approvals per approver&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;14.4&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;19.3&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;+34%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Approvals by the busiest approver&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;145&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;184&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;+27%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Claude Code adds a &lt;code&gt;Co-Authored-By&lt;/code&gt; trailer to commits by default, so I counted those too. The share of human commits on &lt;code&gt;main&lt;/code&gt; with a Claude trailer went from 11% in January to 27% in February and 34% in March. That&amp;rsquo;s a lower bound, since Copilot and Cursor don&amp;rsquo;t leave a trailer and some people switch it off.&lt;/p&gt;
&lt;p&gt;For the first time the review numbers moved the wrong way. Time to first review went from 25 to 34 minutes. That doesn&amp;rsquo;t sound like much, but it&amp;rsquo;s a median, and it moved within a single quarter with the same people on both sides.&lt;/p&gt;
&lt;h2 id=&#34;the-reviewers-question-changed&#34;&gt;The reviewer&amp;rsquo;s question changed&lt;/h2&gt;
&lt;p&gt;A year ago when I reviewed a PR, I could assume the author had written every line and could explain every line. The question in my head was &amp;ldquo;is this correct?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Now a good chunk of the diff was typed by an agent, and the author&amp;rsquo;s job was to steer it and check it. Most authors do that well. Some days, honestly, nobody has read every line. So the reviewer&amp;rsquo;s question became &amp;ldquo;does anybody understand this change?&amp;rdquo;, and that is a harder question to answer from a diff.&lt;/p&gt;
&lt;p&gt;At the same time a lot of review comments are boring and repetitive. Wrong error type. Transaction opened in the repository instead of the service. A DTO mapper where we pass proto types straight through. Missing config validation. A human shouldn&amp;rsquo;t spend their limited review attention on that, and at today&amp;rsquo;s volume they can&amp;rsquo;t.&lt;/p&gt;
&lt;p&gt;So this quarter was about splitting review into the part a machine can check and the part only a human can.&lt;/p&gt;
&lt;h2 id=&#34;docs-where-the-agent-reads-them&#34;&gt;Docs where the agent reads them&lt;/h2&gt;
&lt;p&gt;The first step was the least exciting one. If a convention only exists in someone&amp;rsquo;s head, the agent can&amp;rsquo;t follow it and the AI reviewer can&amp;rsquo;t check it.&lt;/p&gt;
&lt;p&gt;In January we added &lt;code&gt;AGENTS.md&lt;/code&gt; and &lt;code&gt;CLAUDE.md&lt;/code&gt; files to our main repos, and we started a separate docs repo where an agent keeps the code documentation up to date after every merge. The proposal put it like this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Instead of passing diffs to a single AI call, we use Claude Code CLI as an autonomous agent&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The README of that repo opens with a quote I won&amp;rsquo;t reproduce in full, but it ends with &amp;ldquo;Mr. Claude, it&amp;rsquo;s too much.&amp;rdquo; That&amp;rsquo;s the spirit.&lt;/p&gt;
&lt;h2 id=&#34;review-rules-as-skills&#34;&gt;Review rules as skills&lt;/h2&gt;
&lt;p&gt;In late January we created a shared skills repo. Every convention we kept repeating in reviews became a small &lt;code&gt;review-*&lt;/code&gt; skill: &lt;code&gt;review-domain-errors&lt;/code&gt;, &lt;code&gt;review-cls-transactions&lt;/code&gt;, &lt;code&gt;review-connectrpc-only&lt;/code&gt;, &lt;code&gt;review-no-dto-mappers&lt;/code&gt;, &lt;code&gt;review-type-not-interface&lt;/code&gt;, &lt;code&gt;review-config-validation&lt;/code&gt;, &lt;code&gt;review-unit-tests&lt;/code&gt; and a few more. Thirteen at the end of January.&lt;/p&gt;
&lt;p&gt;The skills come with two hooks for Claude Code. One runs after every edit and blocks on lint, type and test failures. The other runs when the agent wants to stop, and spawns one reviewer per principle in parallel against the changed files:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-bash&#34; data-lang=&#34;bash&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;cp&#34;&gt;#!/bin/bash
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;cp&#34;&gt;&lt;/span&gt;&lt;span class=&#34;c1&#34;&gt;# review-principles.sh - Spawns one Claude agent per principle/skill in parallel&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;#&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# This hook runs on Stop (session end) and reviews all changed files&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# against each skill defined in skills/review-*/SKILL.md&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;#&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Each principle gets its own agent for focused, accurate review.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;One agent per principle sounds wasteful. The idea is that an agent looking for exactly one thing is more reliable than one agent juggling a list of thirteen rules.&lt;/p&gt;
&lt;h2 id=&#34;an-ai-first-pass-on-every-pr&#34;&gt;An AI first pass on every PR&lt;/h2&gt;
&lt;p&gt;On the PR side we rolled out CodeRabbit in February. One of the first things we changed was turning off its auto-approve. From the commit message:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;PRs should still require a human approver. CodeRabbit should review and flag issues but not count as an approving reviewer.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In March we moved its configuration into one central repo. Two parts of that config do most of the work. It reads our &lt;code&gt;AGENTS.md&lt;/code&gt; and &lt;code&gt;CLAUDE.md&lt;/code&gt; files as review guidelines, so the docs from step one double as review context. And our most common review comments are now pre-merge checks in &lt;code&gt;error&lt;/code&gt; mode:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-yaml&#34; data-lang=&#34;yaml&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nt&#34;&gt;knowledge_base&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;  &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;code_guidelines&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;    &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;enabled&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;kc&#34;&gt;true&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;    &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;filePatterns&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;      &lt;/span&gt;- &lt;span class=&#34;s2&#34;&gt;&amp;#34;**/AGENTS.md&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;      &lt;/span&gt;- &lt;span class=&#34;s2&#34;&gt;&amp;#34;**/CLAUDE.md&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;      &lt;/span&gt;- &lt;span class=&#34;s2&#34;&gt;&amp;#34;**/.claude/rules/&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;      &lt;/span&gt;- &lt;span class=&#34;s2&#34;&gt;&amp;#34;**/.cursor/rules/&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;      &lt;/span&gt;&lt;span class=&#34;c&#34;&gt;# ...&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;reviews&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;  &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;pre_merge_checks&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;    &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;title&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;      &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;mode&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;error&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;   &lt;/span&gt;&lt;span class=&#34;c&#34;&gt;# conventional commits with a ticket number&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;    &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;description&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;      &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;mode&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;error&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;    &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;custom_checks&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;      &lt;/span&gt;- &lt;span class=&#34;nt&#34;&gt;name&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;Domain Errors Over Plain Errors&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;        &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;mode&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;error&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;        &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;instructions&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;Business logic must throw domain error classes (...) instead of plain Error ...&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;      &lt;/span&gt;- &lt;span class=&#34;nt&#34;&gt;name&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;CLS Transactions Over Manual db.transaction()&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;        &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;mode&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;error&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;        &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;instructions&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;Database transactions must use the CLS-based @Transactional() decorator ...&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;      &lt;/span&gt;- &lt;span class=&#34;nt&#34;&gt;name&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;Config Validation With Joi&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;        &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;mode&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;error&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;        &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;instructions&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;Any new or modified NestJS configuration registered via registerAs() must have a corresponding Joi validation schema ...&amp;#34;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;(There&amp;rsquo;s also &lt;code&gt;poem: true&lt;/code&gt; in there. We&amp;rsquo;re keeping it.)&lt;/p&gt;
&lt;p&gt;By the end of March, CodeRabbit reviewed about 60% of human PRs.&lt;/p&gt;
&lt;h2 id=&#34;what-it-didnt-fix&#34;&gt;What it didn&amp;rsquo;t fix&lt;/h2&gt;
&lt;p&gt;The numbers above are after all of this. Time to first review still went up, and the busiest reviewers still approve more every month. The AI first pass removes the boring comments. Someone still has to understand the change and own the approval.&lt;/p&gt;
&lt;p&gt;What I take from this quarter:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Generation is solved enough. The next tool worth building is one that helps us verify code.&lt;/li&gt;
&lt;li&gt;Conventions belong in the repo, written for an agent. The humans benefit just as much.&lt;/li&gt;
&lt;li&gt;A rule that a machine can check should never cost a human reviewer a minute.&lt;/li&gt;
&lt;li&gt;The approve button is the scarcest resource we have. We should treat it that way.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Next we want agents doing work in the background, not just in an engineer&amp;rsquo;s terminal. More on that soon.&lt;/p&gt;
&lt;p&gt;As always, if you have any questions or remarks, feel free to ping me on twitter &lt;a class=&#34;link&#34; href=&#34;https://twitter.com/bobby_donchev&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;@bobby_donchev&lt;/a&gt;.&lt;/p&gt;
</description>
        </item>
        <item>
        <title>A year of AI coding assistants: what changed in our pull requests</title>
        <link>https://donchev.is/post/a-year-of-ai-coding-assistants/</link>
        <pubDate>Sun, 14 Dec 2025 00:00:00 +0000</pubDate>
        
        <guid>https://donchev.is/post/a-year-of-ai-coding-assistants/</guid>
        <description>&lt;img src="https://donchev.is/post/2025/ai-coding-assistants-2025.png" alt="Featured image of post A year of AI coding assistants: what changed in our pull requests" /&gt;&lt;p&gt;In 2025 a lot of our engineers started using an AI coding assistant. Copilot in the editor for most people, Cursor for some, and since September Claude Code in our biggest monorepo. Everybody &lt;em&gt;felt&lt;/em&gt; faster. I wanted to know what actually changed, so I pulled every merged pull request from GitHub for two comparable two-week windows, one in January and one in October, and compared them.&lt;/p&gt;
&lt;p&gt;Short version: the people writing code got a lot faster. The people reviewing it did not.&lt;/p&gt;
&lt;h2 id=&#34;the-numbers&#34;&gt;The numbers&lt;/h2&gt;
&lt;p&gt;Human-authored PRs merged across all our repositories, bots excluded, in two 14-day windows (January 12-25 and October 12-25):&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th style=&#34;text-align:right&#34;&gt;Jan 2025&lt;/th&gt;
&lt;th style=&#34;text-align:right&#34;&gt;Oct 2025&lt;/th&gt;
&lt;th style=&#34;text-align:right&#34;&gt;Change&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;PRs merged&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;375&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;844&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;2.3x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;People who merged at least one PR&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;57&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;85&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;1.5x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PRs per author per week&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;3.3&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;5.0&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;+52%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median PR size (lines changed)&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;47&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;82&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;+74%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median time to first human review&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;29 min&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;27 min&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;same&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PRs with an AI review comment&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;0%&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;36%&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;People approving PRs&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;43&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;62&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;1.4x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Approvals per approver&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;9.4&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;13.5&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;+44%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Approvals by the busiest approver&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;36&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;107&lt;/td&gt;
&lt;td style=&#34;text-align:right&#34;&gt;3x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Monthly merged PRs (bots included this time) went from about 1,300 in January to about 2,300 in November.&lt;/p&gt;
&lt;p&gt;Part of that is growth. The team grew by roughly half and we started a handful of new services. But per person we merged 50% more PRs a week, and those PRs were bigger. I can&amp;rsquo;t pin all of that on AI, but it is the shape I expected.&lt;/p&gt;
&lt;h2 id=&#34;what-we-set-up&#34;&gt;What we set up&lt;/h2&gt;
&lt;p&gt;Nothing fancy. Most of it was individual engineers picking up tools, and a few things we set up for the whole team.&lt;/p&gt;
&lt;p&gt;In September we added Claude Code to our largest Go/Node monorepo. Anyone can mention &lt;code&gt;@claude&lt;/code&gt; in a PR comment and it picks up the request:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-yaml&#34; data-lang=&#34;yaml&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nt&#34;&gt;name&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;l&#34;&gt;Claude Code&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;on&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;  &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;issue_comment&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;    &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;types&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;p&#34;&gt;[&lt;/span&gt;&lt;span class=&#34;l&#34;&gt;created]&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;  &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;pull_request_review_comment&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;    &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;types&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;p&#34;&gt;[&lt;/span&gt;&lt;span class=&#34;l&#34;&gt;created]&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;permissions&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;  &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;contents&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;l&#34;&gt;read&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;  &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;issues&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;l&#34;&gt;write&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;  &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;pull-requests&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;l&#34;&gt;write&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;&lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;jobs&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;  &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;claude&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;    &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;runs-on&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;l&#34;&gt;ubuntu-latest&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;    &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;if&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;l&#34;&gt;contains(github.event.comment.body, &amp;#39;@claude&amp;#39;)&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;    &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;steps&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;      &lt;/span&gt;- &lt;span class=&#34;nt&#34;&gt;uses&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;l&#34;&gt;anthropics/claude-code-action@v1&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;        &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;with&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;w&#34;&gt;          &lt;/span&gt;&lt;span class=&#34;nt&#34;&gt;anthropic_api_key&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;l&#34;&gt;${{ secrets.ANTHROPIC_API_KEY }}&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Next to it we wrote a &lt;code&gt;CLAUDE.md&lt;/code&gt; with the title &amp;ldquo;Rules for Agents&amp;rdquo;. It is around 340 lines and reads like this:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-markdown&#34; data-lang=&#34;markdown&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;**MUST** rules are enforced by CI; &lt;span class=&#34;gs&#34;&gt;**SHOULD**&lt;/span&gt; rules are strongly recommended.
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;-&lt;/span&gt; **BP-1 (MUST)** Ask the user clarifying questions.
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;-&lt;/span&gt; **C-1 (MUST)** Follow TDD: scaffold stub -&amp;gt; write failing test -&amp;gt; implement.
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;-&lt;/span&gt; **C-7 (SHOULD NOT)** Add comments except for critical caveats; rely on self-explanatory code.
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;-&lt;/span&gt; **T-3 (MUST)** ALWAYS separate pure-logic unit tests from DB-touching integration tests.
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;-&lt;/span&gt; **D-1 (MUST)** Use SQLC for type-safe database access with proper Go structs.
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Our newest service got Cursor rules in its repo. And by October, Copilot&amp;rsquo;s PR reviewer and Cursor&amp;rsquo;s review bot were leaving comments on about a third of our human PRs.&lt;/p&gt;
&lt;h2 id=&#34;writing-rules-for-agents-is-writing-rules-for-humans&#34;&gt;Writing rules for agents is writing rules for humans&lt;/h2&gt;
&lt;p&gt;The most useful thing that came out of this was the rules file, more than any code the assistants wrote.&lt;/p&gt;
&lt;p&gt;We had conventions before. They lived in people&amp;rsquo;s heads, in old Slack threads and in review comments that got repeated every other week (&amp;ldquo;please don&amp;rsquo;t open a transaction in the repository&amp;rdquo;). Writing a &lt;code&gt;CLAUDE.md&lt;/code&gt; forced us to put them in one place for the first time. I&amp;rsquo;d argue that file is as useful to a new engineer on their first day as it is to the agent.&lt;/p&gt;
&lt;p&gt;The second lesson is in the first line of that file. A rule in a markdown file is a suggestion. The agent follows it most of the time and so do humans. If a rule really matters, a tool has to enforce it: a linter, a type check, a CI step. Our MUST rules are the ones we can check in CI, and in practice those are the ones that hold.&lt;/p&gt;
&lt;h2 id=&#34;review-didnt-change-and-thats-the-problem&#34;&gt;Review didn&amp;rsquo;t change, and that&amp;rsquo;s the problem&lt;/h2&gt;
&lt;p&gt;Look at the bottom of the table. Time to first review stayed flat at under half an hour, which sounds great. But the number of approvals per approver went up 44%, and the busiest approver went from 36 approvals in two weeks to 107. That&amp;rsquo;s more than ten a working day on top of their own work.&lt;/p&gt;
&lt;p&gt;Authors scale with tools and hiring. Reviewers only scale with hiring. The AI review bots help a bit: they catch typos, a missing null check, the occasional real bug. They don&amp;rsquo;t approve anything, though, and they shouldn&amp;rsquo;t. Somebody still has to understand the change and press the button.&lt;/p&gt;
&lt;p&gt;PRs also got bigger, 47 lines to 82 at the median. A bigger diff that the author partly didn&amp;rsquo;t type is also harder to review.&lt;/p&gt;
&lt;h2 id=&#34;what-im-watching-in-2026&#34;&gt;What I&amp;rsquo;m watching in 2026&lt;/h2&gt;
&lt;p&gt;If the assistants keep getting better at this pace (and it looks like they will), the volume on the writing side will keep growing and the review queue is where it will hurt first. So that&amp;rsquo;s where we want to put our effort next year:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Put our documentation where the agents read it, and keep it up to date with an agent instead of hoping someone does it.&lt;/li&gt;
&lt;li&gt;Turn more of the review comments we keep repeating into rules a machine can check, so human reviewers can spend their time on the questions only a human can answer.&lt;/li&gt;
&lt;li&gt;Keep measuring. It took me an afternoon to pull these numbers and I learned more from them than from a year of &amp;ldquo;it feels faster&amp;rdquo;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I&amp;rsquo;ll write a follow-up once we have a few months of data.&lt;/p&gt;
&lt;p&gt;As always, if you have any questions or remarks, feel free to ping me on twitter &lt;a class=&#34;link&#34; href=&#34;https://twitter.com/bobby_donchev&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;@bobby_donchev&lt;/a&gt;.&lt;/p&gt;
</description>
        </item>
        
    </channel>
</rss>
