← Back to blog

What Is Automated Code Review, and What Does It Miss?

Automated code review spans linters, static analysis, and AI reviewers. What each catches, what none know without a plan, and who owns the merge.

Search “automated code review” and the pages that rank split into two camps. Static analysis vendors say it means rules running in continuous integration. AI review vendors say it means a model reading your pull request. Both are right about the mechanism and quiet about the same two questions: what none of these tools can see, and who owns the result. This post takes the three layers in order, shows what each catches on one change, and then gets to the part the tool pages skip. Software can read a diff. It cannot know what the diff was for, and it cannot be the one who decided it ships.

What is automated code review?

Automated code review is the use of software to read a code change and report on defects, style, security, and maintainability before a person reviews or merges it. Wikipedia’s definition is close to that, and it has been true for twenty years. What changed in the last two is the third layer.

The category now has three layers, and most pages ranking for the term describe only one of them.

Layer What it reads What it produces Examples
Rules Tokens and syntax Pass or fail against a fixed list Linters, formatters, import ordering
Analysis Data flow, types, known patterns Findings matched to a rule id Static analysis, SAST, dependency scanners
A reviewer that reads The diff, the surrounding code, and, if one exists, the plan Findings in prose and a verdict Proctor on a YAGNI Team, and the AI review tools on the market

The first two are deterministic. Give them the same code twice and they say the same thing, which is why they belong in the merge checks and why nobody argues with them. The third is a reader. It can notice that a retry loop has no upper bound or that a new endpoint skips the tenancy filter every other endpoint applies. It can also be wrong in ways a rule cannot. That difference decides where each layer belongs.

What do linters and static analysis actually catch?

Everything that can be written down as a rule ahead of time. Unused imports, inconsistent formatting, a SQL string built by concatenation, a dependency with a published vulnerability, a nullable value dereferenced without a check. Static analysis is good at this because the rule is the specification: the tool knows exactly what it is looking for and reports it the same way every time.

The limit is the same as the strength. A rule catches what someone already knew to write a rule for. Static analysis has no opinion on whether the change is a good idea, whether the approach fits the rest of the system, or whether the tests test what the ticket asked for. It reports that a function is 60 lines long. It cannot say whether the function does too much or the domain is that complicated.

Keep this layer. It costs almost nothing per pull request, and a reviewer that reads should never spend its attention on what a rule catches for free.

What does an AI reviewer catch that rules cannot?

The things a careful engineer notices on a first read and a rule cannot express. A retry with no backoff. A migration that adds a column but no index on the column every query will filter by. A test that asserts the mock was called instead of asserting the behavior. An error path that logs and continues where the caller expects a throw.

None of those has a rule id. Each requires reading the change in context and asking whether it makes sense. That is what a model does well, and why the AI review layer exists. Sourcegraph’s roundup of automated review tools, published in May 2026, draws the same line: rule-based analysis matches code against patterns, while an AI reviewer sends the diff to a model and gets back semantic feedback. CodeRabbit, Copilot code review, Qodo, and Greptile sit in this third layer, and so does Proctor on a YAGNI Team.

The cost of a reader is that it can be wrong in prose. A reader that raises a finding might have misread the intent, applied a convention this codebase does not follow, or flagged something the author already weighed. That is why the third layer needs two things the first two never did: a source of intent to read against, and a person who decides.

What does none of these layers know?

What the change was supposed to do.

Every layer of automated review starts from the diff, and the diff is the answer to a question nobody wrote down. A reviewer that only has the diff is reconstructing the question from the answer: what the ticket was, what approach was chosen, what was rejected, what is deliberately out of scope. A model is good at guessing that. A wrong guess still produces a confident finding about a problem that is not there.

The fix is upstream of the review. If a plan exists before the code does, the reviewer reads the change against the plan instead of against its own reconstruction. “Does this do what the approved plan says, and nothing the plan ruled out” is a question a reader can answer well. “Is this good” is not.

That is why, on a YAGNI Team, the review of a ticket starts before Build. Reeve, the Staff Engineer, plans the ticket against the real repository with the hard-to-reverse risks in hand, and checks its own plan before the Plan gate opens: the approach, the risks, what the plan leaves out. Reeve never builds. A person reads the plan and approves it, or sends it back. Only then does Wright write code. When Proctor reviews the pull request, the approved plan is the first thing it reads, and its business fit check is the question above, asked of every file in the diff.

An AI reviewer on a repository with no plans is still worth having. It is just doing the hardest half of its job, working out intent, with the least information.

Why did automated review make pull requests slower in the one study that measured it?

Because comments are not free, and the tool pages never count them.

An industrial study presented at ICSE 2025, Automated Code Review In Practice, followed a company that rolled an LLM reviewer out across 4,335 pull requests, 1,568 of them reviewed by the tool. The developers resolved 73.8% of the automated comments, which says the findings were mostly worth acting on. The average time to close a pull request went from 5 hours 52 minutes to 8 hours 20 minutes. The reviewer added work to every pull request, and the work landed on the people who had to read, judge, and answer each comment.

Microsoft’s engineering blog reports the opposite outcome at a much larger scale: an AI reviewer on more than 90% of pull requests, over 600,000 a month, with median completion time improving by 10 to 20 percent. Both can be true. The difference is not the model. It is how many findings reach a person, in what form, and who triages them.

This is the number to ask any automated review vendor for, and almost none publish: minutes of human attention per pull request, added or removed. A reviewer that posts fifteen inline comments, three of them wrong, has moved the team from writing code to arguing with a bot. A reviewer that posts one review with a verdict, lists what it checked, and keeps the findings it discarded one click behind the review gives the person merging something to read in minutes. Clearing a pull request review backlog covers the human side of that queue; the automated side is the same arithmetic.

Who owns the review when the reviewer is software?

A person, and the tool has to make that easy rather than merely allowed.

Every vendor page ends on “a human should still review before merging.” It is true and it is not a design. Ownership of a review means three things, and an automated reviewer either supports them or erodes them.

A named person merges. Not “the team”, not the tool’s approve button wired into branch protection. On a YAGNI Team, Proctor never blocks and never merges. It posts one review with a verdict. On a pull request the Team opened, an approving verdict takes the draft out of draft with no click, and then a person merges. The repository’s branch rules stay in charge. What Amazon’s reported sign-off rule got right is exactly this: the decision to ship is a person’s, on the record, every time.

The record shows what was checked. A quiet review is worth nothing if it is indistinguishable from no review. Proctor’s review lists the checks it ran, and the findings it set aside sit behind the review with their reasons, so the person merging can see that the tenancy filter was looked at and found present. Every review is also a row on Work, with what it read, what it found, and what it cost.

Corrections go somewhere. When the person merging overrules a finding, that has to change the next review or the reviewer never gets calibrated. On a Team, an overturned verdict lands on Proctor’s track record, and a correction on the pull request can become a Playbook rule the Team reads next time. Who reviews AI-written code argues that this loop is where trust in an automated reviewer is earned.

What does one pull request look like through all three layers?

Take a ticket on a subscriptions service: cancelling a subscription must also revoke the customer’s API tokens, which today keep working until they expire. The approved plan says: revoke in the cancel handler’s transaction, add the index the revocation query needs, and do not touch the token issuance path. Wright builds it and opens a draft pull request touching five files.

Question the reviewer asks Rules Static analysis Proctor, reading against the approved plan
Is the code formatted, with no unused imports? Yes, two import fixes Not its job Not its job; the rules already ran
Is there a known-bad pattern? No Flags a string-built query in a test fixture, a false positive Reads the fixture, discards the finding with the reason
Does the new query filter by customer? No opinion No rule for tenancy Database check: it filters by customer_id, like every other query on the table
Is the index the plan called for present? No opinion No opinion Database check: the migration adds it, leading with the tenancy column
Does the revocation run in the same transaction as the cancel? No opinion No opinion Deep review: it does not; the revoke call sits after the commit. Finding raised
Did the change stay inside the plan? No opinion No opinion Business fit: one file edits the issuance path, which the plan ruled out. Finding raised
Do the tests test the behavior? No opinion No opinion Tests check: the test asserts the revoke function was called, not that the tokens stop working. Finding raised
Verdict Pass Pass, one false positive Changes requested: three findings, two discarded findings behind the review with reasons

Rules and analysis both pass this pull request, and they are right to: nothing in it violates a rule. The three findings that matter are about intent, and two are only findable because a plan existed to compare against. Wright addresses them, Proctor follows up on the new commit, and when the review is clean an approving verdict marks the pull request ready. The person supervising the Team reads the plan, the review, and the diff, and merges. Three clicks over the life of the ticket: accept the proposal, approve the plan, merge.

Can an AI reviewer review AI-written code?

Only if the reviewer is a different reader than the author.

Sonar’s State of Code survey, published in March 2026, found that 38% of developers say reviewing AI-generated code takes more effort than reviewing a colleague’s, and 61% say AI produces code that looks correct but is not reliable. Generated code is fluent. It passes the linter, the names are good, and the bug is in the assumption behind the approach rather than in any line a rule can see.

A model reviewing its own output mostly re-runs the judgment that produced it. It catches the typo and misses the assumption it was already confident about. Adversarial AI code review makes the case for a separate reviewer whose job is to argue with the change, and covers the calibration problem that follows.

On a YAGNI Team the separation is structural. Wright builds. Proctor reviews as a different Agent with different checks and its own track record, and the plan a person approved sits between them, so the review is a comparison rather than a second opinion from the same mind. Where a Team wants more, Fletcher (beta) walks the change in a browser in a sandbox and attaches video and screenshots, never editing application code.

How does a YAGNI Team run automated code review?

Through the GitHub App the workspace installs, on the repositories a Team attaches. Engaging Proctor on a Team with a repository turns reviews on, and one Team reviews a repository, so no pull request gets two competing reviews.

Proctor reviews every pull request opened there, whoever wrote it, and posts one review through its checks: deep review and business fit always on, database, security, and tests on the line, plus any check the Team adds in plain words, such as “anything touching billing gets the pricing rules read”. It follows up on each new commit until the pull request is clean. It never blocks and never merges. The rules layer keeps running in CI as before: Proctor is additive, and the review it posts assumes the linter ran.

Most engineering orgs start with a Team at the Pull request reviews size, Proctor alone, and read its findings beside their own for a week before adding the build lines. How to get Proctor reviewing your pull requests walks through the setup.

Where to start

Keep the rules. If your linters and static analysis are not in the merge checks already, put them there this week.

Then ask one question of any AI reviewer you evaluate, including Proctor: what does the person merging have in front of them, and how long does it take to read. A wall of inline comments and no verdict is a backlog. One review, the checks it ran, the findings it set aside with reasons, and the plan the change was measured against is review. AI code review tools turns that one question into the full criteria list for the category, with the tools on the market beside Proctor.

Who merges is not a question. It is a person. The only thing a tool should change is how much that person has to reconstruct before they can click.

To see a pull request of yours go through Proctor’s review, see how the engineering org works or book 30 minutes and bring one. The Agents run on vetted US-hosted open-weight models at 60% or more under comparable frontier API rates, usage-based on one rate card.