← Back to blog

What Should Autonomous Coding Agents Be Allowed to Own?

Autonomous coding agents can own a ticket from an approved plan to a draft pull request. What stays with a person, and the record that makes it hold.

Search “autonomous coding agents” and the pages that rank sell the unattended part. Hand the agent a ticket, go to lunch, come back to a pull request. Devin, OpenAI’s Codex, Claude Code, GitHub Copilot’s coding agent, Cursor’s background agents, and Google’s Jules all run some version of that loop now. Nearly every page that ranks them ends the same way: “of course, a human should still review before merging.” Then it stops.

That sentence is where an engineering leader’s actual question starts. Not “can it run unattended” but “what may it own, end to end, and what does the record have to show me before I let it.” The short answer: an autonomous coding agent should own the middle of a ticket, from a plan a person approved to a draft pull request, and the three decisions around that middle stay with a named person. On a YAGNI Team the Agent that runs that middle is Wright.

What is an autonomous coding agent?

An autonomous coding agent takes a unit of engineering work and runs the whole loop without a person at the keyboard: it reads the codebase, decides what to change, edits files, installs what it needs, runs the tests, and opens a pull request. The difference from a coding assistant is who is driving. With an assistant, a developer accepts each step. With an agent, the developer hands over a ticket and walks away.

“Background coding agent” is the same thing named from the developer’s chair: the work runs in a cloud environment while the developer does something else, and the result arrives as a pull request. The trigger is usually a ticket in Linear or Jira or a message in Slack.

None of that says what the agent is allowed to do. “Autonomous” in the product copy is a claim about capability, that the loop can run to completion. The word an engineering manager needs is narrower: for each act inside that loop, does it happen with no one looking, or does it wait for a named person? That is a property of the line of work the agent is on, not of the agent, and a person sets it. Managing AI coding agents covers the general frame; this post applies it to the one job everybody wants to make autonomous first.

What does “autonomous” actually cover on one ticket?

Take a ticket apart into its acts, and the word stops being a switch and becomes a list.

  1. Propose. Decide this ticket is next.
  2. Plan. Decide how, and what is out of scope.
  3. Build. Edit the code, add tests, install what the change needs.
  4. Verify. Run the tests and walk anything with a surface.
  5. Open the pull request. Write up what changed and why.
  6. Review. Read the change against the plan and argue with it.
  7. Merge. Decide it ships.
  8. Deploy, migrate, or act outside the repository. Run what a revert cannot undo.

An agent can be autonomous on acts 3 through 6 and still be a well-managed agent. An agent that is autonomous on act 7 or 8 is a different product with a different risk profile, and most pages ranking for this keyword blur the two. The blast radius of each act decides the rung: a wrong plan costs one round of edits, a wrong merge costs an incident, and a wrong migration costs a weekend. Who reviews AI-written code makes that argument for the review act.

What should an autonomous coding agent own end to end?

Here is the split a YAGNI Team draws, act by act.

Act on the ticket Who owns it How it runs on a YAGNI Team
Propose Bailey proposes; a person accepts Bailey reads the Team’s Jira or Linear backlog and proposes the next ticket with evidence. A proposal always waits for a person.
Plan Reeve plans; a person approves Reeve, the Staff Engineer, plans against the real repository and checks its own plan before the gate opens. It never builds. A person clicks “Approve plan” or “Send back”.
Build Wright, end to end Wright clones the repository into a sandbox, implements the approved plan, adds tests, and runs them. This is the act that runs unattended.
Verify Wright runs the tests; Fletcher (beta) walks the change Fletcher writes a test plan from the change, walks it in a browser against the repository environment, and attaches video and screenshots. It never edits application code.
Open the pull request Wright Opened as a draft on GitHub through the workspace’s GitHub App, with the plan and the test runs on the ticket’s case file.
Review Proctor One review with a verdict. An approving verdict marks the pull request ready, with no click from anyone.
Merge A person, always A person merges on GitHub or from the Review case file. There is no merge line and no autonomous merge.
Deploy, migrate, act outside the repo A person, always Deploy and data-migration lines cap at supervised on every Team. That floor does not move.

Read down the middle column and the answer is there. The agent owns Build, Verify, and the pull request. A person owns the three decisions around that middle: accept the proposal, approve the plan, merge. Review is an Agent’s act, but it produces a verdict a person reads, not a merge.

Who approves the plan before a line of code is written?

A person does, and this is the edge the ranking pages skip most often. Most autonomous coding agents go straight from ticket to code. The plan, if there is one, is something the agent narrates to itself on the way, and by the time a person sees anything it is a diff.

The plan is the cheapest place in the whole ticket to say no. A wrong plan sent back costs a few minutes and one more draft. The same mistake caught at review costs a build and a salvage conversation. Caught after merge it costs an incident. So a YAGNI Team puts the gate there. Reeve, the Staff Engineer, drafts the plan: the files it expects to touch, the approach, the tests it will add, and what it is leaving alone. It plans the way a Staff Engineer would, with the hard-to-reverse risks in hand, and checks its own plan for the module it missed or the simpler approach it walked past before the gate opens. Reeve never builds. A person reads the plan and clicks “Approve plan” or “Send back” with a note.

Two things come out of that click. Wright has a contract, so Build can run unattended against something a person already agreed to. And Proctor has something to review against.

Why does the merge stay with a person?

Because merge is the act that makes everything before it real, and the company needs a named person who decided it should be real.

This is not a YAGNI opinion. Amazon’s own engineering leadership has said on the record that every mutating step an AI might take requires a person to approve it (what Amazon mandating AI code review means for your team has the sources). The lesson that lasts is not the rule but what makes it affordable: the person merging needs something readable in front of them, not a diff to reconstruct.

On a YAGNI Team a supervised ticket takes three human clicks: accept the proposal, approve the plan, merge. The merge click is the only one that has no rung. Wright opens the pull request as a draft, Proctor posts one review with a verdict, and an approving verdict marks the pull request ready. Then a person merges, on GitHub or from the Review case file, having read the approved plan, the test runs, and Proctor’s review with the findings it discarded and why. Done means merged, on the source’s say-so.

Where should an unattended agent run?

In a sandbox: a temporary computer for that run, wiped when the run is done, which cannot ship anything anywhere on its own. The only thing that leaves it is a branch and a pull request, through the workspace’s GitHub App installation, which is the one door into the customer’s repositories. There is no personal GitHub connection for an Agent to borrow and no shared credential to trace an incident back to.

A Team’s repositories are set up once as a repository environment: the primary repository plus the dependency repositories a build needs, each pinned, with the setup commands, services, and variables the tests require. Wright builds in it and Fletcher verifies in it, which is the difference between an agent that runs the real test suite and one that claims it did.

What record should an unattended run leave?

Enough that a person who did not watch the run can trust the pull request it produced. “Log everything” is the standard advice, and it is useless: a transcript of the agent’s reasoning is a debugging aid, not a record of who owned what. A YAGNI ticket’s record is the case file, and it has a fixed shape:

  • The proposal and who accepted it.
  • Reeve’s plan, its check, and who approved it.
  • Every run in the sandbox, with the tests that ran and what they returned.
  • The pull request, as a draft and then as ready.
  • Proctor’s review, the verdict, and the findings it set aside with the reasons.
  • Fletcher’s test plan, video, and screenshots where the change had a surface.
  • The clicks a person made, by name, and when.
  • What the ticket cost to run.

The last line is the one the vendor pages avoid. A team running background agents usually has no idea what a ticket costs, because the meter is a monthly bill with no line items. On a YAGNI workspace each ticket’s case file shows what its runs were charged, and the Usage page breaks spend out per developer and per Agent. The models under those Agents are vetted US-hosted open-weight models at 60% or more under comparable frontier API rates, on one rate card that developers and Agents share (pricing).

Corrections become Playbook rules: when a person sends a plan back or overturns a verdict, the Team learns the rule. And every result lands on the line’s track record, which is what the next section needs.

How does a line move from supervised to autonomous?

By one explicit act of a person, reading the record, on one line of work. Never by a streak.

The Ladder on a YAGNI Team has two rungs, set per line. Supervised: the act waits as a gate addressed to a named member of the Team, and one click that is still that person’s dispatches it. Autonomous: the act dispatches and leaves a Receipt. On a new Team, Reeve’s plan line is supervised, so by default Wright does not start building until a person approves the plan.

Beside each line sits its track record: accepted as proposed, edited, sent back, auto-shipped, reversed, over time. A person looks at “22 of the last 25 accepted as proposed” and moves Wright’s build line to autonomous with one act, with that record beside it. Nothing moves a rung on its own, and nothing suggests it. A reversal later is one more count on the record, and setting the line back to supervised is following the evidence, not losing an argument. Pause stops every line on the Team at once. The Ladder lesson walks through it line by line.

The permanent floor sits under all of this. Deploy and data-migration lines cap at supervised on every Team, and merge has no line at all, whatever the record says. That is the row most autonomous-agent products leave to policy, and policy is the thing nobody enforces at 6pm on a Friday.

What does one unattended ticket look like, start to finish?

Here is a ticket on a Team whose build line a person has already made autonomous. The Team works on a payments API, reads its Linear project, and posts to a channel in Slack.

  1. Proposed. Bailey reads the Linear backlog and proposes “Return a Retry-After header on 429 responses from the export endpoint,” with the evidence: three support threads in the last month. Priya, who owns the Team, reads it and clicks “Start plan”.
  2. Plan. Reeve drafts a plan: add the header in the rate-limit middleware, compute it from the limiter’s reset time, add two tests, leave the client SDK alone. Reeve’s own check notes that the limiter already exposes a reset timestamp the first draft was about to recompute, and the plan is redrafted to read it. Priya reads the plan and clicks “Approve plan”.
  3. Build. Because the build line is autonomous, Wright starts without another click. It clones the repository into a sandbox, implements the header, adds the tests, and runs the suite. The runs land on the case file with their output.
  4. The pull request. Wright opens a draft pull request on GitHub with the plan linked. Proctor reviews it: one review, an approving verdict, with one finding it set aside (a style note the Team’s Playbook says not to raise) and the reason. The verdict marks the pull request ready.
  5. QA. Fletcher (beta) writes a test plan from the change, walks the endpoint against the repository environment, and attaches the evidence to the ticket.
  6. Done. Priya opens the Review case file, reads the plan, the review, and the runs, and merges. The ticket moves to done in Linear because GitHub says the pull request merged, and that Receipt is what Done means. Harper’s Daily Brief the next morning posts to Slack with the receipts attached.

Two clicks of Priya’s before any code was written, one after. The same ticket with the build line still supervised has one more click, before step 3, and is otherwise identical. Your first ticket walks through the supervised version.

So what should an autonomous coding agent be allowed to own?

The middle of the ticket. From a plan a person approved to a draft pull request, built and tested in a sandbox, with a review a person can read waiting on the other side. The three decisions around it, accept, approve, merge, stay with a named person, and the record of each run is what keeps those decisions cheap. The review waiting on the other side of the build is the half of that split a tool choice decides, and AI code review tools covers what to look for in it.

An agent that owns more than that is not more autonomous. It is unmanaged, and unmanaged agents are the pilots that impress for an afternoon and decay into babysitting. An agent that owns less is an assistant with a longer leash, and the review load never drops. The split above is the one an engineering team already applies to a new hire’s first quarter: you own the build, we approve the plan, we merge.


YAGNI runs agent teams, managed like your engineering team: Bailey proposes, Reeve plans, Wright builds to a draft pull request, Proctor reviews, Fletcher (beta) walks the change, and Harper reports, with merge always a person’s click. See how agent teams work, or book 30 minutes to draw the split for your own repositories.