← Back to blog

What Amazon Mandating AI Code Review Means for Your Team

Amazon reportedly made senior engineers sign off on AI-assisted changes after outages. The management lesson for engineering leaders, with sources.

On 10 March 2026 the Financial Times reported that Amazon’s retail technology group had a new rule: junior and mid-level engineers would need a more senior engineer to sign off on any AI-assisted change. The report came after a bad week for amazon.com, and it spread the way these stories do, as a headline about AI code breaking one of the largest stores on the internet and the humans being brought back in to guard the gate.

Amazon says that framing is wrong. Its published response says only one of the incidents involved an AI tool at all, that none involved AI-written code, and that reports of new approval requirements for engineers working with AI tools are false.

This post lays out what was reported, what Amazon said, and what is on the record since, every claim linked to a named source. Then it gets to the part the coverage skipped: why a senior sign-off is the right frame for agent work, why it fails as a bare rule, and what makes it affordable.

What did Amazon reportedly mandate?

The Financial Times saw two documents. One was a briefing note for “This Week in Stores Tech”, the weekly availability meeting of Amazon’s stores organization. The note described a “trend of incidents” with a “high blast radius” and listed “Gen-AI assisted changes” among the contributing factors, along with “novel GenAI usage for which best practices and safeguards are not yet fully established.” The other was an email from Dave Treadwell, a senior vice-president in that group, that opened with “Folks, as you likely know, the availability of the site and related infrastructure has not been good recently.” Ars Technica’s account of the FT report carries the wording; the FT’s own article is behind a paywall.

The sign-off rule is one sentence in that reporting: junior and mid-level engineers will now require more senior engineers to sign off any AI-assisted changes. Nothing in the reporting defines “AI-assisted”, says how a change would be tagged as one, or names the teams it covers.

Fortune reported on 12 March that the FT had seen two versions of the briefing note, and that the reference to GenAI-assisted changes was deleted before the meeting took place. Fortune also counted four high-severity incidents on the retail site in a single week, including an outage on 5 March that ran for around six hours.

What does Amazon say happened?

Amazon’s response is worth reading in full rather than through a headline. In a post on its own site, the company said that only one of the recent incidents involved AI-assisted tooling, and that it “related to an engineer following inaccurate advice that an AI tool inferred from an outdated internal wiki, and none involved AI-written code.” On the reports of new approval requirements for engineers using AI tools, the post says: “That is false.” It also says the incidents were limited to the retail store infrastructure and did not involve AWS.

That was the second time in a month Amazon had pushed back on a story linking an AI tool to an outage. In February the FT had reported that Kiro, Amazon’s agentic coding tool, had deleted and recreated an environment during a mid-December incident affecting AWS Cost Explorer in one region. Amazon’s account attributes that incident to misconfigured access controls, describes it as an extremely limited event affecting a single service, and lists mandatory peer review for production access among the safeguards added afterwards. An Amazon spokesperson also told The Register that the company had “not seen compelling evidence that incidents are more common with AI tools.”

Then, in April, an Amazon director said something on the record that reads like the durable version of the policy. Steve Tarcza, who runs the StoreGen group that supports Amazon’s internal retail developers, told The Register: “Right now, every mutating step that an AI might do requires a human to approve it.” He added that this goes “all the way down to publishing a document for someone to read.”

So the record, in order:

Date What happened Source
Mid-December 2025 Outage affecting AWS Cost Explorer in one region. FT later links it to Kiro; Amazon attributes it to misconfigured access controls and adds mandatory peer review for production access. Amazon, The Register
Early March 2026 Four high-severity incidents on the retail site in one week, including a roughly six-hour outage on 5 March. Fortune
10 March 2026 FT reports the briefing note, the Treadwell email, and the senior sign-off rule for AI-assisted changes. Ars Technica
10 to 12 March 2026 Amazon says one incident involved an AI tool giving inaccurate advice, none involved AI-written code, and reports of new approval requirements are false. FT says the GenAI reference was deleted from the note before the meeting. Amazon, Fortune
29 April 2026 Amazon Stores director Steve Tarcza: every mutating step an AI might take requires a human to approve it. The Register

Whether or not a memo said “senior sign-off on AI-assisted changes”, the thing Amazon has confirmed in public is stricter than the headline: a person approves every consequential act an AI takes, and nothing ships without someone looking at it. That is the rule worth studying, because it is the one your team can actually adopt.

Why is a senior sign-off the right frame for agent work?

Strip the Amazon specifics away and the rule is an old one. Every change to a production system has a named person who decided it should ship, and that person has read enough to own the decision. The AI coding wave quietly eroded it, not by anyone deciding to, but because agent output arrives faster, looks finished, and is right often enough that reviewers stop reading. Who reviews AI-written code covers that calibration drift.

A sign-off rule puts the decision back on a person. That is the correct move, for three reasons.

It fixes accountability before the incident, not after. “The AI did it” is not an incident writeup, any more than “the intern did it” was. A named approver on the record gives the postmortem somewhere to start and gives the approver a reason to read.

It matches review effort to blast radius. Amazon’s own vocabulary in the briefing note was “high blast radius”, and that is the right variable. A change to checkout, pricing, or accounts deserves a senior read whoever wrote it. A test scaffold does not need one whether a person or an agent typed it.

It keeps governing judgment with people. An agent can exercise operating judgment inside a scope someone gave it. Whether that scope is right, how much authority it holds, and what ships stay with a person. A sign-off rule is that boundary written down.

Why does the rule usually fail on its own?

Because a rule is not a workflow. The reporting on Amazon, and the commentary after it, treated the sign-off as the whole answer. It is the easy half.

Picture the senior engineer the rule lands on. A pull request arrives: four hundred lines across nine files, a two-sentence description. The author, a mid-level engineer working with a coding agent, is not sure which parts the agent wrote. There is no plan on record, so the reviewer cannot tell whether the approach was chosen or accepted. The tests pass, which says the code does what the tests check, not that it does what the ticket asked. To sign this, the reviewer has to reconstruct the intent, check the diff against it, then decide.

That is an hour, if they are careful, and the rule now charges that hour for every AI-assisted change from every junior and mid-level engineer. Two things happen next, and neither is what the rule wanted. Either the senior engineers become the bottleneck and throughput drops to the rate at which the most experienced people can read, or they start approving on trust and the sign-off becomes a signature with nothing behind it. That is worse than no rule, because it records accountability that was never exercised.

Nobody in the Amazon coverage quantified this. The engineers in the Hacker News thread on the report kept circling the question anyway: who reviews the reviewers’ time?

What does the sign-off cost when the signer only gets a diff?

Here is the same change, a ticket to add a rate limit to a checkout endpoint, arriving at the senior engineer two ways. The first is how it lands in most shops today. The second is how it lands when the review is built to be signed, which is how a YAGNI Team hands it over.

What the signer needs to know Raw pull request Pull request with a plan, a review, and a record
What was this change supposed to do? Two-line description; reconstruct intent from the diff The approved plan is attached: the limit, the key, the fallback, and why
Was this approach chosen, or was it the first thing that worked? Unknown The plan shows the alternatives, and the check it passed before any code was written
Does it touch anything with a high blast radius? Read all nine files to find out The review lists the checks it ran: security, database, tests, business fit, with findings per check
What did the reviewer decide not to flag, and why? Nothing; a quiet review is indistinguishable from no review The discarded findings sit behind the review, each with its reason
Who has already looked at this, and what did they approve? Git history The case file: proposal accepted by whom, plan approved by whom, review verdict, what it cost
Time to sign with confidence About an hour, or a rubber stamp Minutes: read the plan, read the review, spot-check the diff, merge

The right column is not lighter review. The senior engineer still owns the merge and still reads the diff. What changed is that reconstructing intent and hunting for the dangerous line happened before the pull request reached them, done by something whose job was to argue with the change. That is the difference between a sign-off that scales and one that turns senior engineers into a queue.

What makes a sign-off cheap without making it hollow?

Four inputs, all of which can be attached to every pull request regardless of who or what wrote the code.

A plan a person approved, before the code. The cheapest place to catch a wrong approach is before it exists. When the intent, design, and risks are written down and a person says yes, the pull request reviewer checks execution against a decision instead of guessing at the decision.

A review that argues. A reviewer that summarizes the diff and says “looks good” adds nothing the signer can use. One built to find a reason to reject the change, and to say why it did not, does the senior engineer’s reconstruction work for them. Adversarial AI code review is how that reviewer is built, and kept from flagging everything.

The negative space. The most useful line in a review is often the finding that was considered and set aside, with its reason. It tells the signer what was looked at, which is exactly the information a raw diff cannot carry.

A record that survives the incident. Who accepted the proposal, who approved the plan, what the review found, who merged, and what it cost. The Amazon story’s most telling detail was that the reference to GenAI-assisted changes was deleted from the briefing note before the meeting. Whatever the reason, the record moved. A record that cannot move is what a sign-off rule is really trying to create.

How does a YAGNI Team run the same rule?

YAGNI runs agent teams, managed like your engineering team, and the rule Amazon’s director described, a person approving every mutating step, is how a Team is built rather than a policy anyone has to enforce.

A ticket moves through six stages: Proposed, Plan, Build, Review, QA, Done. Reeve plans the ticket before a line of code is written, and a person approves the plan. Wright builds to a draft pull request. Proctor reviews that pull request against the approved plan, runs its checks (deep review, business fit, database, security, tests, and any check the Team adds), and posts one review with a verdict. The findings it weighed and discarded sit one click behind the review, each with its reason. An approving verdict marks the pull request ready. Then a person merges. Always.

That is the right-hand column of the table above, generated by YAGNI as a byproduct of the work rather than assembled by hand for the signer. The senior engineer who owns the repository opens the case file and finds the proposal, the plan they or a colleague approved, the review, and the cost, before they read a line of the diff. How Proctor reviews pull requests walks through what it posts and what it never does.

Two more things map directly onto the Amazon lesson.

First, the person is named. Every supervised line on a YAGNI Team is addressed to a specific member, and the copy reads “supervised by Priya”, not “needs approval”. The Ladder has two rungs, supervised and autonomous, set per line by a person who has read that line’s track record, and merge sits on neither rung because merge is always a person’s click. A supervised ticket takes three human clicks: accept the proposal, approve the plan, merge.

Second, the floor does not move. Deploys, data migrations, and anything irreversible stay supervised however long the track record runs. That is the same shape as “every mutating step requires a human to approve it”, except that it is a property of the line rather than a memo, so it cannot be deleted before the meeting.

Proctor also reviews pull requests a person wrote by hand on the repositories the Team attaches, which is how most teams start: one review on every pull request, whoever the author, so the senior sign-off has something to read. Managing AI coding agents covers the correction loop that follows, where a reviewer’s edits become Playbook rules the Team reads next time.

Where to start

Do not start by trying to tag which changes were AI-assisted. Nobody, including the reporting on Amazon, has described a way to do that reliably, and the distinction matters less every month. Start with the rule Amazon has confirmed and the inputs that make it affordable.

Name the person who signs off on production changes for each repository you care about. Write down what they need in front of them to sign in ten minutes rather than an hour: the intent, the tests, and one review that argues with the diff. Measure how long a sign-off takes today. If the answer is an hour, the inputs are the problem, not the rule. If you are choosing a reviewer to put in front of that sign-off, AI code review tools: what to look for lists the criteria a manager asks of one.

Then decide, with the record in front of you, which lines of work can carry a lighter gate. Keep the irreversible ones where Amazon keeps them and where YAGNI keeps them, with a person, every time.

If you want to see a pull request of yours go through that review, see how the engineering org works or book 30 minutes and bring one. Usage-based on one rate card, no seats, and the Agents run on vetted US-hosted open-weight models at 60% or more under comparable frontier API rates, so the review that makes the sign-off cheap does not bring back the cost on the other side.