AI

Six weeks of AI code review: what it caught and what it let through

On this page

The comment started with three words: This breaks prod.

It was a PR of about five hundred lines — in a repo I had just encountered, and I was’t the one who wrote it. An agent did, after opening three files the diff didn’t show.

Weeks later, the same agent let a bug through to staging, and it stayed there for twelve days. Between those two stories sits what an AI can do in a code review, and what it can’t.

In my first weeks on this project, about 25 PRs were created weekly, across two repos I had never opened.

I would open a diff, read it, understand every line, and still not be able to say whether it was right. Because "right" depended on things that were not in the diff: the conventions the team followed, the screen that component drew, the other endpoint that called the same route.

Approving like that is signing off on something you did not check. And a green CI makes that far too comfortable. Your name stays on the PR, and the bug nobody saw ships to production with your approval next to it.

The review I wanted to do was the ideal one. You check out the author’s branch, run what there is to run, and confirm nothing broke. You check whether the new test tests anything, or whether it would pass without the fix. You stress it: log in as the role that should not see that screen, shrink the window until the table overflows, submit the empty form. And you check whether the PR fits the house, whether it uses the helper and the pattern the neighboring module already adopted, or reinvents everything.

Doing that for 25 PRs a week, in code I did not know, did not add up.

And much of that code had been written by an agent. If one agent wrote it, could another one do the review I had no time for? And if it could not do all of it, which part could I hand over?

One number to keep in mind before answering: ask people why they review code and the top answer is finding bugs. Bugs are only 14% of what reviews actually produce.

What we look for in a review

That number comes from Alberto Bacchelli and Christian Bird, who studied code review inside Microsoft in 2013. They observed 17 developers reviewing, interviewed each of them, and surveyed 873 programmers and 165 managers.

Finding bugs is the top motivation for 44% of programmers and 44% of managers. But when the authors classified 570 review comments, defects ranked in fourth, at 14%.

Both lists from the study

One of the questions was: why do you do code review? Each person picked their top three reasons, and the authors ranked them by score:

1. Finding defects
2. Code improvement
3. Alternative solutions
4. Knowledge transfer
5. Team awareness
6. Improving the development process
7. Avoiding build breaks
8. Shared code ownership
9. Tracking rationale
10. Team assessment

And the nine categories of the 570 comments:

| Category | Comments |
|—|—|
| Code improvement | 29% |
| Understanding (questions and doubts) | ~22% |
| Social communication | ~15% |
| Defects | 14% |
| External impact | ~5% |
| Testing | ~4% |
| Review tool | ~3% |
| Knowledge transfer | 2% |
| Misc | ~6% |

> **NOTE:** the values marked ~ are read off the paper’s chart. The exact ones (29%, 14%, and the 12 knowledge transfer comments) are in the text.

The largest slice, 29%, is code improvement, which is exactly style and convention: use the practice the team already uses, remove dead code, make it readable. Second comes understanding, people asking what something does and why. The authors heard the same thing in the interviews, without asking:

the most difficult thing when doing a code review is understanding the reason of the change

To understand, reviewers read the description, try to run the changed code, email the author, and 20% to 40% of the time get up and go talk in person. Owners of the changed files don’t even read the description: they go straight to what they know. Everyone else gets stuck. "Not knowing files (or [dealing with] new ones) is a major reason for not understanding a change."

I did not know a single file.

A decade later, Turzo and Bosu hand-classified 2,500 comments from OpenStack Nova and asked 160 developers to rate how useful each kind was. What developers value most is what shows up least: functional defects, validation, and logic top the ratings and add up to only 19% of comments. Documentation and code organization make up more than 40%.

Turzo and Bosu’s table

| Category | Usefulness rating (1-5) | % of comments |
|—|—|—|
| Functional defect | 4.38 | 0.48% |
| Logical | 4.11 | 2.28% |
| Question | 3.99 | 13.56% |
| Documentation | 3.73 | 33.32% |
| Organization of code | 3.68 | 7.68% |

A reviewer that doesn’t need to be human

I started with the obvious: use AI to gather the context I did not have. It reads the ticket, the CLAUDE.md, the ADRs and the surrounding code before looking at the diff. The first version was a skill I invoked in Claude Code, in the terminal, one PR at a time. Soon after, it became cc-harness: a small program around Claude Code that takes a PR number, picks the right skill for the repo, runs the whole review on its own, and hands me the result in a file, so I read it before posting anything.

Along the way, I went and actually studied code review, and it turned into a study with more than twenty reading notes. The sentence I wrote down early on sums up where I landed:

Code review did not become obsolete. It became more necessary. But what we look for in it changed.

And one decision came out of the reading: I gave up on the idea that an agent that reviews needs to behave like a human. The question became "what should a machine reviewer do that a person wouldn’t?".

Some things are only cheap for a machine. Opening the docstring three files away to see if it is still true. Finding every caller of a route. Booting the app for every single PR, without getting lazy, at four in the afternoon on a Friday.

Collina’s checklist

The review skill did not start from scratch. Matteo Collina publishes the skills he uses to work on Node.js. One of them, nodejs-core, has a rule called reviewing-prs.md: the guide he follows to review core PRs. Commit message, scope, tests, documentation, CI. And a red flags section I really like:

Tests that would pass without the fix: the test doesn’t actually exercise the changed code path. Verify by mentally (or actually) reverting the fix and checking if the test would still pass.

I took that skeleton and adapted it to the project. Collina’s blocks are still there, with nearly the same names: Scope and size, Commit messages, Quality red flags, and the bar for when to request changes and when to approve with comments. On top went the house rules: authorization as its own axis, backend and frontend patterns, environment variables.

But Collina’s checklist delegates execution to CI. It asks whether CI passed and shows how to read the result. Running things locally is something it tells you to suggest to the contributor. For Node core, that makes complete sense: the CI is huge and there is no screen to look at.

On this project, the story is different. In one PR, the members table started overflowing the screen. At 1280×800, the container was 902px and the table was 1443px, and the Email and Edit columns sat behind a horizontal scroll. No test measures width, so CI had no way to catch it, and a diff does not show pixels.

It is like reviewing a kitchen’s floor plan and never opening a drawer. The plan can be right and the drawer can still hit the fridge.

The review starts running things

So the review started running things on its own. For each PR, the agent:

  • creates an isolated git worktree at the PR’s commit, outside my working tree;
  • runs typecheck, lint and tests, with MySQL in a throwaway container;
  • serves the PR’s own frontend on port 3100 (never 3000, which is serving my branch);
  • opens the app in Chrome through the Chrome DevTools MCP, logs in, goes to the changed flow and tests it.

Screenshots are only evidence. What confirms a layout problem is measuring, with evaluate_script, on the real DOM. That table was caught this way, at 1280 and again at 1440:

Inline GitHub review comment anchored to lines 160 to 165 of registered-users-table.tsx. It says this is a real bug: putting the role chips inline inside the fullName cell blows the column past its size; at 1280x800 the members-page container is 902px and table.scrollWidth is 1443px, so 541px of the row, including Email and the Edit action, sits behind a horizontal scroll; at 1440x900 the shortfall is still 381px.

When the backend can’t boot, the frontend runs on the MSW handlers the project already has. The agent turns them all on, seeds a test organization, and marks every tweak with // REVIEW SCRATCH: Not part of the PR. so nobody mistakes mock behavior for the PR’s behavior.

NOTE: MSW hides things too. In one PR, the screen showed Make and Model just fine because the mock served those fields. The backend never sent them. The bug only surfaced by reading the transformer. Running the app does not replace reading the code – it adds to it.

"This breaks prod."

That comment from the opening had no screen at all.

In the legacy app, a PR of about five hundred lines turned on cursor pagination for a transactions route. The author had done their homework: they checked the eight other callers of the function they changed and concluded the change was purely additive.

The AI went after the callers of the route, not the function. It found three screens that still paged by offset, and all of them would start returning the same rows on every request. The comment opened with three words:

Inline GitHub review comment on src/server/api/transactions/get.js, lines 87 to 89. It says the change breaks production: with_cursor true is set unconditionally, so find.js drops the OFFSET clause and the route stops returning total. Three other callers of /api/transactions/get still page on offset, linked as PodTransactions.jsx:232, demo/index.jsx:284 and transaction_details/pod.jsx:215. Below, a suggestion block replaces with_cursor true with with_cursor: before_id !== undefined and with_total: before_id === undefined.

The author replied an hour and a half before the merge: "this was a real prod break", along with the commit that fixed it. It never reached production because something opened three files the diff did not show, and that nobody had a reason to open.

How the harness works

cc-harness has almost no agent code. In the post about the harness via the SDK, I showed the claude-agent-sdk preset that turns on Claude Code’s tools, system prompt and loop in two lines. The project’s AGENTS.md locks that in as a rule: "The point of this project is that the SDK already is the harness."

// trimmed from src/index.ts
const result = query({
  prompt,
  options: {
    tools: { type: "preset", preset: "claude_code" },
    systemPrompt: { type: "preset", preset: "claude_code" },
    settingSources: ["user", "project", "local"],
    skills: [reviewSkill, "humanizer"],
    mcpServers, // read from the harness's own .mcp.json
    permissionMode: "bypassPermissions",
  },
});

settingSources brings in my environment: the same skills and plugins I use in interactive claude, symlinked into ~/.claude/skills. Edit a SKILL.md and it applies to the next review and the next terminal session. humanizer always rides along, to cut filler. At first I wanted the review to read like a person wrote it; later I saw there was no need at all. It is an AI review, and what matters is that it does its job.

What the harness adds sits around that, in a bash wrapper called cc-review. It detects the repo from the git remote, never from the folder name, and picks the skill:

Skill routing diagram: cc-review leads to a 'which repo?' box that splits into three lanes, portal, legacy and other. In two columns, local mode and post mode: portal uses the portal's local skill and the portal's post skill; legacy uses the legacy local skill and github-pr-review; other uses local-pr-review and github-pr-review. A banner at the bottom says humanizer rides along in all of them.

That split came from a mistake. I tried to make one skill serve both repos, and pointed at the legacy app it demanded authorization policies and an MSW layer that repo does not have. The comment left in the script is, to me, the most important sentence in the project:

The wrong one does not merely miss things, it asserts things the repo does not have.

Inside the skill, the review always follows the same ten steps, recorded in the agent’s own task list. That list is what cc-status reads to report which phase a review is in:

Diagram of a cc-harness review: ten boxes in sequence, Target, Spec, Big picture, What changed, Scope check, Find problems, Verify, Write, Humanize and Lint. A bracket over the first five reads context before the diff. Under Verify, a browser icon notes that each finding is tested. A curved arrow goes from Lint back to Write, noting fix and run again. After Lint, an arrow leads to the PR notes document. Above, a robot labeled cc-status watches a progress bar for the current phase.

Two lines from the skill sum it up. In Big picture: "A diff can be locally correct and still wrong for the system." In Verify, where every finding is tested before it becomes a comment: "A killed false positive is the system working."

The output is a notes file with a fixed layout. In local mode, which is what I use most, nothing goes to GitHub: I read it, cut what I disagree with, and only then post. One comment, as it sits in the file:

<!-- anchor path="apps/frontend/src/users/components/member-roles-panel.tsx" start_line=88 line=90 side=RIGHT -->

This is a real bug.

When `holdsNoneHere` is true the field reads `No role in this group` and
nothing else, so the `< 2` guard hides the toggle. [...]

```suggestion
          {everyGroup.length < (holdsNoneHere ? 1 : 2) ? null : (
```

*Verified: read the branches in `member-roles-panel.tsx` for a member with one
grant on another branch [...] `everyGroup.length` is 1, and the button renders nothing.*

The anchor carries the file and lines the GitHub API needs, and a linter checks each one against the diff before anything gets posted. The proof line, *Verified:* when the agent tested it and *Inferred:* when it only read it, never goes to the PR. It exists for me, when I decide whether a finding stays.

Twelve days

In another PR, the AI asked for a change: the documentation said disconnection was event_code=2, and the code used 1.

The author disagreed and changed the documentation to 1. On re-review, the AI read it again, saw code and documentation agreeing, and was satisfied. The PR went in.

Twelve days later, staging showed that disconnection is 2. For those twelve days, the last_disconnected field showed the last connection. Anyone who looked at it saw the wrong information, with all the confidence of a well-named field. It never reached production, but it sat there the whole time.

You can tell this as an AI failure. I tell it as expected behavior.

In the first review, code and documentation disagreed, and that is something it can check. In the second, they agreed. Which number the device actually sends is not in the diff. It is in the vendor’s documentation, in the device itself, in the head of someone who knows the system. A suspicious reviewer would have asked why the documentation changed in the same PR as the code. The agent doesn’t ask, because from inside the PR there was nothing left to check.

The AI checks. A human attests to what is right. That part I don’t delegate.

Two passes

That is why the project’s workflow has two passes:

  1. cc-harness does the first one and always posts the review as COMMENT. The AI does not approve and does not block.
  2. A human does the second one and decides: APPROVE or REQUEST_CHANGES.

Heander, Söderberg and Rydenfält sat next to ten developers through 34 real reviews. Almost half of what goes through a reviewer’s head is trying to understand the context and the rationale of the change, and tools ignore that part: reviewers go looking for it in Jira, in chat, in the docs. The first five steps of the pipeline are that search, done before I get there.

I pulled the project’s PRs from August 14 until today. The AI reviewed 79. Of the 30 findings it marked critical or high, 28 were confirmed. Across all its findings, 54 were correctness bugs, 23 were missing tests or tests that did not test anything, and 45 were code conventions.

NOTE: this is little data: six weeks, one team, two repos, and I classified the severity myself with an agent’s help. And it does not show that the AI finds more than a human, or the opposite, because the two passes almost never landed on the same PR.

A review with the app running takes about 12 minutes and costs a median of US$ 7.57. That is expensive for a one-line PR and cheap next to broken pagination in production.

Build your own

cc-harness and the project’s skills are private. But I wrote a generic version of the skill for this post, review-that-runs, with Collina’s skeleton and the rules that paid off most. It knows nothing about my project, and learns yours as it goes:

mkdir -p ~/.claude/skills/review-that-runs
curl -L -o ~/.claude/skills/review-that-runs/SKILL.md \
  https://gist.githubusercontent.com/geeksilva97/f3b55b015ca0d0c711793c3852d2a502/raw/SKILL.md

It needs two things in your Claude Code. The first is an authenticated gh. Since it only reads the PR and never posts, you can run it with a read-only fine-grained token, scoped to the repos you review, with Contents, Pull requests and Metadata set to Read-only. gh uses whatever token is in GH_TOKEN:

GH_TOKEN=github_pat_... claude

With that token, when you do post the review, you post it with your usual gh. The skill never had permission to.

The second is the Chrome DevTools MCP:

claude mcp add chrome-devtools -- npx chrome-devtools-mcp@latest

Then ask for review PR 123 inside the repo. It creates the worktree, runs the gates, boots the app if the PR touches a screen, and writes pr-123-review-notes.md without posting anything. The rules I consider non-negotiable sit at the top:

  • validation means the real app, running in Chrome, not reading the diff;
  • every finding ends with *Verified:* or *Inferred:*, and no proof means no comment;
  • any mock tweak made to run the app is marked // REVIEW SCRATCH and named in the overview;
  • the review goes out as COMMENT, and approving is a human’s call.

And when you post, the review closes with a single line: First pass by review-that-runs. Approving is a human’s call. Whoever opens the PR knows where it came from, and knows the sign-off still belongs to someone.

If you want to go further, the original reviewing-prs.md rule is in Collina’s skills repository, and the Agent SDK’s claude_code preset, which I showed in the previous post, runs the same skill outside the terminal.

Your team will notice on the first PR where the review arrives with the table width measured in pixels, or with the three callers the diff did not show.

Checking and attesting

I still get about 25 PRs a week. The difference is that when I open one, the context is already on the table: the screen was opened, the tests were questioned, the callers were found, and every finding comes with its proof.

What is left for me is the part that does not fit in the diff. Knowing that the device sends 2.

Every piece of the harness exists so that my sign-off requires less attention. But it is still a sign-off. When I approve, I know what I am signing.

Thanks for reading!

If you want to dig deeper:

We want to work with you. Check out our Services page!

We want to work with you

The engineers who write here are the same ones who join client teams. Let’s build something that grows with your business.

Book a call