← Back to blog

GitHub AI Code Review: What It Posts and Who Still Merges

How AI code review works on GitHub pull requests: what a good reviewer posts, why it never blocks, how to set one up, and why a person still merges.

Type “github ai code review” and the results split three ways: GitHub’s own documentation for Copilot code review, vendor pages for the reviewers you can install from the Marketplace, and listicles ranking them. All of them explain how to get comments onto a pull request. Almost none of them answer the questions an engineering lead asks before turning one on: what should the reviewer post, can it stop a merge, what does it cost the people reading it, and who is accountable when it is wrong. This post takes those in order, then shows how to set a reviewer up on a repository in an afternoon, and what one pull request looks like once it is running.

What is AI code review on GitHub?

A reviewer that reads every pull request and posts a review before a person does, through the same review API a human uses.

Mechanically it is a GitHub App. The App is installed on a repository with pull request permissions, GitHub sends it a webhook when a pull request opens or a new commit lands, and the App reads the diff, the surrounding code, and whatever else it has been given, then posts a review: inline comments, a summary, and a verdict of approve, comment, or request changes. From GitHub’s side it is indistinguishable from a person on the team.

There are three ways to get one onto a repository. Copilot code review is built in and is requested like a person. A GitHub App from the Marketplace, which is how CodeRabbit, Greptile, Qodo, and Proctor on a YAGNI Team arrive, is installed once by an admin and reviews everything it is pointed at. A GitHub Action runs a model against the diff inside your own CI, which is the cheapest to start and the one that leaves you maintaining the prompt.

That mechanism is the same for every tool in the category. What differs is everything the mechanism leaves open. Does the reviewer post one review or thirty comments? Does its verdict count against branch protection? Does it read the pull request against anything but the diff? Does it remember being overruled? What automated code review is, and what it misses covers where an AI reviewer sits beside the linters and static analysis already running in CI. This post is about the GitHub side: what shows up on the pull request, and who is in charge of it.

What does GitHub already do on its own?

GitHub ships Copilot code review, and a team should know what it does before adding anything beside it.

GitHub’s documentation for Copilot code review, read in September 2026, describes it this way. Copilot is requested on a pull request the way any reviewer is, by adding it as a reviewer, or automatically through a repository ruleset or an organization setting, so it can run on every pull request opened or marked ready. It reads the change with context gathered from the repository, follows the instructions a team keeps in .github/copilot-instructions.md and AGENTS.md, skips dependency files and logs, and posts its review as a comment with an approval assessment. By default that review does not count toward required approvals. A preview setting can let it, at the repository or organization level, and a new commit dismisses it. GitHub prices it per review, as estimated consumption on top of a Copilot plan, as of September 2026.

Two of those details matter more than the rest. Copilot’s review is a comment, not an approval, unless a team opts in, which is the right default and the one this post argues for below. And it is priced per review, which is the honest unit for this job: a reviewer costs what it reads, not a seat.

Everything below is additive to that. A reviewer that reads the pull request against an approved plan, runs the checks a manager would ask a senior engineer to run, and leaves a record of what it looked at is a second reader, not a replacement for the first. On a YAGNI Team, Proctor runs alongside whatever review a repository has today, including Copilot, and the repository’s branch rules stay exactly as they were.

What should an AI reviewer post on a pull request?

One review, with a verdict, the checks it ran, and the findings it set aside readable behind it.

The failure mode of the category is volume. A reviewer that posts fifteen inline comments on a forty-line change has not reviewed the change. It has generated a triage queue, and the person who has to clear it is the engineer who opened the pull request. The one industrial study that measured this, presented at ICSE 2025, found pull requests took longer to close after an LLM reviewer arrived, even though most of its comments were acted on. The findings were fine. The form was the cost.

The form that works is the one a good senior engineer already uses:

  • One review per pull request. Findings grouped, ordered by what matters, with a verdict at the top. Follow-ups on each new commit until the pull request is clean, not a fresh wall of comments each push.
  • The checks it ran, named. A quiet review is only worth something if the reader can see what was looked at. Proctor’s review lists its checks: deep review and business fit always on, database, security, and tests on the line, plus any check the Team adds in plain words, such as “anything touching billing gets the pricing rules read”.
  • The findings it discarded, with reasons. A reviewer weighs more than it posts. The findings it set aside, and why, sit one click behind Proctor’s review, so the person merging can confirm the tenancy filter was checked and found present rather than never considered.
  • The plan it read against, if there is one. A diff is an answer to a question nobody wrote down. When a plan exists, “does this do what the approved plan says, and nothing the plan ruled out” is a question a reviewer can answer well. On a YAGNI Team that plan is the one a person approved before Wright wrote a line.

Should an AI reviewer be allowed to block a merge?

No. Keep it out of required reviewers, and keep branch protection exactly as it is.

It is tempting to wire the reviewer into the merge rules, because then “AI reviewed it” becomes enforceable. Two things go wrong. The practical one: every false positive becomes a fight with a bot, and the engineer either argues in the thread or learns to write code that keeps it quiet. The one that matters to a manager: a blocking reviewer has been handed a decision. Whether a change ships is the one call software should not own, because nobody can be held to account for it afterwards.

Proctor never blocks and never merges. It posts one review with a verdict, and the repository’s branch rules decide what a verdict means, the same as they do for a person’s. On a pull request a YAGNI Team opened, an approving verdict takes the draft out of draft: with no click once the line is autonomous, or after the person named on the line confirms the verdict while it is supervised. Then a person merges, on GitHub or from the case file in the app once the workspace’s Merge in app switch is on. That is the whole design: merge is always a person’s click, and there is no setting that changes it.

What Amazon’s reported sign-off rule got right is the same point from the other side. A person decides what ships, every time, on the record. A reviewer that posts one readable review makes that cheap. A reviewer that blocks makes it someone else’s problem.

How do you set up AI code review on a GitHub repository?

For a YAGNI Team, it is one Team at the Pull request reviews size, and it takes minutes rather than a sprint.

  1. Install the GitHub App on the repositories you want reviewed. The App is the only door YAGNI has into your code: it reads pull requests and posts reviews through it, and nothing else touches the repository.
  2. Create a Team and choose Pull request reviews. Pick the repositories from the ones the installation covers. A repository another Team already reviews shows as taken, with that Team’s name, because one Team reviews a repository and that is what stops two reviewers posting competing reviews on the same pull request.
  3. Let Proctor post on its own. That is the default on a new Team: Proctor posts its review on the pull request by itself. On a person’s pull request, a clean review reads “approval advised”, so the approval your branch rules count is still a person’s. Read its reviews next to your own for a week or two; if you would rather confirm each verdict first, move the line to supervised.
  4. Add the checks that match your repository. Database, security, and tests toggle on Proctor’s line. Add your own in plain words, optionally scoped to paths: “anything under payments/ gets the idempotency rules read”.
  5. Open a pull request. Reviews start on the next one. Nothing else runs, no code changes, and every review is also a row on Work with the case file behind it: what it read, what it found, and what it cost.

How to get Proctor reviewing your pull requests walks through the screens. Growing the Team later is adding a line: Reeve’s plan line and Wright’s build line turn a review-only Team into one that carries tickets through to draft pull requests, on the same repository claim, the same checks, and the same record.

What does one pull request look like with an AI reviewer on it?

Take a service with a review-only Team on it, and a pull request opened by one of the engineers, not by an Agent. The change adds a webhook endpoint that receives payment events and updates the order’s status.

When What happens on GitHub Who acted
10:04 The pull request opens. CI starts; the linter and static analysis run as before. The engineer
10:06 Proctor posts one review. Deep review: the handler retries on failure with no upper bound. Security: the endpoint verifies the signature after parsing the body rather than before. Tests: the new test asserts the update function was called, not that the order’s status changed. Verdict: changes requested. Two findings it weighed and set aside sit behind the review, one of them the string-built query in a test fixture, discarded with the reason. Proctor
10:41 The engineer pushes a commit that bounds the retry and moves the signature check. They reply on the tests finding: the status assertion lives in an integration test the unit test does not duplicate. The engineer
10:43 Proctor follows up on the new commit: two findings resolved, the tests finding withdrawn with the reply recorded. Verdict: approve, posted as “approval advised” because a person wrote the pull request. Proctor
11:10 The tech lead reads the review and the diff, approves, and merges. Branch protection required one approving human review, as it did before the Team existed. The tech lead

Total human attention on the review: a few minutes to read one review, a reply on one finding, and a merge. That is the number to hold any GitHub AI code review tool to, and the one most of them do not publish. Clearing a pull request review backlog is the same arithmetic from the queue’s side: the first read never waits, and Work shows how long every open review has.

The reply on the tests finding is not lost. It lands on Proctor’s record, and if it should outlive the pull request, the Team proposes it as a Playbook rule so the next review reads the convention instead of flagging it. Who reviews AI-written code makes the case that this loop, corrections that change the next review, is where trust in a reviewer is earned.

How do you tell a good GitHub AI code review tool from a noisy one?

By asking the questions a manager asks of a human reviewer, not the ones a demo answers.

Every tool in the category reads the diff and finds real bugs some of the time. The criteria that separate them are about form, accountability, and the record. This is the table to bring to an evaluation, with Proctor filled in and the columns you should fill in for anything else you trial. The longer version, with the tools on the market filled in, is AI code review tools.

Criterion Why it matters to the person merging Proctor on a YAGNI Team
One review with a verdict, or many comments? Minutes of attention per pull request is the real cost One review, follow-ups per commit, verdict at the top
Can it block a merge? A blocking reviewer owns a decision a person should Never blocks, never merges; branch rules stay in charge
What does it read against? A diff without intent is a guess The approved plan when one exists, the connected tools, and the Team’s Playbook
Can you add your own checks? Your conventions are not in any vendor’s training data Checks in plain words on the line, optionally scoped to paths
Does it show what it set aside? A quiet review is only meaningful if you can see it looked Discarded findings with reasons, one click behind the review
Does a correction change the next review? Otherwise you re-argue the same finding every week Replies land on the record; corrections become Playbook rules
Is there a record per review? Someone will ask what was reviewed and what it cost A row on Work with the case file: what it read, found, and cost
Who is accountable? The tool cannot be The Team’s owner, with a track record per line

Where to start

Turn on a reviewer for one repository, supervised, and read its reviews beside your own for a week. Keep it out of required reviewers. Count minutes, not findings: a reviewer that finds twelve things and costs an hour of triage per pull request has moved work, not removed it.

If the answer is one review, the checks it ran, the findings it set aside, and the plan the change was measured against, you have review. If the answer is a comment thread and a bot to argue with, you have a backlog with a new author.

To see a pull request of yours go through Proctor’s review, see how the engineering org works or book 30 minutes and bring one. The Agents run on vetted US-hosted open-weight models at 60% or more under comparable frontier API rates, usage-based on one rate card, and the Usage page shows what each review cost.