← Back to blog

What Is Agentic Engineering, and How Is It Not Vibe Coding?

Agentic engineering is coding agents doing the work under a plan a person approves, a review a person can read, and a record. Vibe coding has none.

Andrej Karpathy coined “vibe coding” in February 2025 to describe a way of programming where you say what you want, accept whatever the model gives you, and never read the diff. A year later, in a post on X in early February 2026, he proposed a different name for what the practice had become: agentic engineering, because the new default is that you are not writing the code directly most of the time, you are orchestrating agents who do and acting as oversight. Addy Osmani wrote it up the same week in one line: “Vibe coding = YOLO. Agentic engineering = AI does the implementation, human owns the architecture, quality, and correctness.” Simon Willison started a patterns guide later that month.

Since then the term has acquired an IBM explainer, several glossary entries, and a dozen vendor posts. Nearly all of them define it, contrast it with vibe coding, and stop at the point where it gets hard: what an engineering org actually does differently on Monday morning, what exists on paper after an agent ships something, and who answers for it. This post is about that part.

What is agentic engineering?

Agentic engineering is the practice of building software where coding agents do the implementation and a person owns everything that makes the implementation safe to ship: the specification, the plan, the quality bar, the review, and the decision to merge. The agent runs a loop, reading the repository, editing files, running tests, and iterating on failures. The engineer runs the agent.

The definitions in circulation differ mostly in emphasis. Karpathy’s names the shift in where the engineer’s time goes, from typing to orchestrating and overseeing. Osmani’s names what the human keeps. Willison’s is the narrowest, agents that can both generate and execute code, and he is candid that the interesting questions are what you do around that. The common ground is the one that matters to anyone running a team: the agent’s output is not the deliverable. The deliverable is a change a named person has decided to ship, with enough in front of them to decide well.

It helps to separate the term from a close neighbor. “Agentic coding” is the 2025 term, used in the research literature, for the loop itself: an agent that plans, edits, runs, and corrects without a human in the middle of each step. Agentic engineering is the 2026 term for the discipline around that loop. The first describes what the agent does. The second describes what the engineer, and increasingly the organization, does.

How is agentic engineering different from vibe coding?

Both use the same models and often the same tools. Claude Code, Codex, and YAGNI Code will all vibe code happily if you let them. The difference is not in the agent. It is in what exists before the code, what exists after it, and who reads which.

Vibe coding Agentic engineering
What exists before the code A prompt A plan: scope, approach, the files it expects to touch, the tests it will add, what it is leaving alone
Who reads the code Nobody, by definition A reviewer, and before that a review the reviewer can read in minutes
What exists after the merge The diff The plan, the review, the test runs, the questions asked and answered, who approved what, and what it cost
Who is accountable Unclear, which is the point A named person, for a named decision, at a fixed checkpoint
Where it fits Prototypes, throwaways, weekend projects, anything you can delete Production code, shared repositories, anything a customer or a colleague depends on
What a mistake costs Nothing, if you really can delete it A rejected plan, a sent-back pull request, or an incident, depending on where it was caught

The row that decides the rest is the first one. If the only thing that exists before the code is a prompt, then every decision about the change is being made inside the agent, and the only place a person can disagree is the finished diff. By then the cheapest moment to say no has passed. Agentic engineering moves the first decision to the plan, where saying no costs one more draft.

The second thing the table shows is that vibe coding is not wrong. Karpathy’s original post was about a throwaway project, and it is still the right way to build one. The failure is vibe coding something that is not a throwaway and finding out later. What autonomous coding agents should own draws the same line per act rather than per project.

Is agentic engineering just a marketing term?

A fair question, and the comment threads under every definition ask it. The honest answer is that the term is only as real as the artifacts it produces. If you call your process agentic engineering and it produces nothing a vibe coder’s does not, it is vibe coding with a better name.

So here is a test that does not depend on the name. After an agent ships a change, can a person who was not in the room answer these three questions from the record alone?

  1. What was approved before any code was written, and by whom?
  2. What did the reviewer actually read, and what did they set aside and why?
  3. What did the change cost, and what would it have cost to catch the same problem a stage earlier?

A team that can answer all three is practicing agentic engineering whatever it calls itself. A team that can answer none is vibe coding whatever it calls itself. Most teams rolling out agents today sit in between, and the point of the discipline is to move them.

The skepticism also has a source worth taking seriously. Several large engineering orgs have tightened their rules on AI-assisted code after production incidents, and Amazon’s mandate that senior engineers sign off on it is the best known. Read charitably, that is not a retreat from agents. It is an organization discovering that the review step had quietly become the whole safety story, and putting a name on who owns it.

What does agentic engineering look like for one ticket?

Definitions are cheap. Here is what the discipline produces on one ticket, on a YAGNI Team, because the artifacts are easier to see when the roles have names.

The Team owns a payments service on GitHub, reads a Linear project, and posts to a channel in Slack. Its owner is Maya, an engineering manager. The ticket is PAY-218, “Retry failed webhook deliveries with backoff.”

Proposed. Bailey, the Product Manager, proposes the ticket from the backlog with the evidence: eleven webhook deliveries failed last week, nine of them to one customer whose endpoint was down for four minutes. Bailey never ships anything; a proposal waits for a person. Maya reads it and clicks Start plan.

Plan. Reeve, the Staff Engineer, drafts the plan in a sandbox against the real repository: a retry schedule on the delivery job, an idempotency key on the request so a retry cannot double-post, a dead-letter table after the last attempt, tests for each. One question comes back to Maya: should a delivery that fails on a 4xx retry at all? She answers no. Reeve checks its own plan before the gate opens, and the check catches that the job queue the service already uses has retry semantics the first draft was about to reimplement, so the plan is redrafted to use them. Reeve never builds. Maya reads the plan and its check and clicks Approve plan.

That is the first artifact. Before a line of code exists, there is a written plan, a check of it, a product decision recorded in the Team’s decision ledger, and an approval with a name on it.

Build. Wright implements the approved plan on a branch in the sandbox, writes the tests, runs the suite, and opens a draft pull request through the workspace’s GitHub App. Nothing touches the trunk. On a new Team this line starts supervised, so the build waits for one click from Maya before it begins, until she has read enough of Wright’s track record to move the line to autonomous.

Review. Proctor, the Reviewer, reads the pull request through its checks (deep review, business fit, database, security, tests) and posts one review with a verdict: approved, one finding fixed in a follow-up commit from Wright (the dead-letter table was missing an index on the lookup the admin page would use), and one finding set aside with the reason (a naming convention the Team’s Playbook already covers). An approving verdict marks the pull request ready. Proctor never merges.

That is the second artifact. The review is a page, not a wall of inline comments, and it says what was set aside and why. Who reviews AI-written code is the longer argument for why that shape matters.

QA. Fletcher, in beta, boots the Team’s repository environment at that revision, writes a test plan from the change, walks the admin page in a browser with a deliberately failing endpoint, and attaches the video, the screenshots, and a verdict. Fletcher never edits application code.

Done. Maya opens the case file, reads the plan, the review, the test runs, and Fletcher’s evidence, and merges on GitHub. The merge webhook moves PAY-218 to Done with a Receipt. Linear closes the ticket because GitHub says the pull request merged. Harper, the Chronicler, names it in the next morning’s Daily Brief in Slack, with the pull request number beside it.

That is the third artifact: the case file, which is the ticket’s whole record, including every click Maya made, by name and time, and what each run cost. She made four on this ticket: the three fixed checkpoints, accept the proposal, approve the plan, and merge, plus one to start the build, because the build line on a new Team begins supervised. Move that line to autonomous once its track record earns it and a ticket takes three. The agentic software development lifecycle walks the same six stages in full.

Why does your review bandwidth not parallelize?

The part of agentic engineering that vendor posts skip is the team. One engineer running one agent can read everything it produces. One engineering manager with eight engineers each running three agents cannot, and the agents do not care. Osmani’s version of this is that cognitive debt compounds: code nobody understands accumulates faster than code anyone wrote. The manager’s version is simpler. Agents parallelize. Attention does not.

That is the real problem the discipline exists to solve, and it is why “the human owns quality” is not enough on its own. Owning quality across thirty pull requests a day requires three things that vibe coding has no use for:

  • A plan that is cheap to reject. Reading a plan takes minutes. Reading a finished diff that implements the wrong plan takes an hour and ends in a rewrite. Putting the first checkpoint at the plan is how a manager’s attention stretches across a Team.
  • A review that can be read, not just run. Linters and test suites scale without people. A reviewer that reads intent and writes one page a person can skim is what lets the merge decision stay human without becoming a bottleneck. Managing AI coding agents covers the accountability side of the same point.
  • A record that survives the people. When the engineer who approved a plan leaves, the record of why still has to exist. Decisions, Receipts, and the case file are that record on a YAGNI Team.

This is also why agentic engineering is a capacity story and not a headcount story. A Team of Agents carries work an engineering org would otherwise staff for. The engineers it reports to are the same engineers, spending their attention at the two or three points where it changes the outcome.

What does agentic engineering cost?

Less than vibe coding, measured per change rather than per token, and the distinction is the point.

Token price is the number on a vendor’s pricing page. Task cost is what it takes to get one change merged, including the retries, the plan that was wrong because the agent did not know the codebase, and the review that caught it late. Vibe coding optimizes the first and is indifferent to the second. Agentic engineering lowers the second by spending a little more up front: a plan costs a draft, and a plan rejected at the gate costs one more draft instead of a build.

Two levers move the bill further. The first is grounding: an agent that reads the Team’s decision ledger and the connected tools before it guesses makes fewer wrong plans, so there are fewer retries to pay for. The second is the models. YAGNI routes the Agents, and YAGNI Code for the developer, across vetted US-hosted open-weight models at 60% or more under comparable frontier API rates, with a router that places each step on the cheapest lane that holds quality. Both show up on the Usage page per developer and per Agent, broken out by the job that spent it, which is how a manager already thinks about a team’s time.

How do you start practicing agentic engineering on a team?

Not by buying a definition. The order that works, on a YAGNI Team or on your own tooling, is the order of the artifacts.

Start with the plan. Make the agent write down what it intends to do before it does it, and make a person approve that document. On a YAGNI Team this is the Plan stage: Reeve plans, a person clicks Approve plan, and only then does Wright build. On a single developer’s terminal it is the same discipline with YAGNI Code, which is the CLI door into the same models, meter, and record.

Then make review readable. A review a person can read in five minutes is the difference between a merge decision that stays human and one that quietly becomes a rubber stamp. Proctor’s one review, with its verdict and what it set aside, is built for that reader.

Then keep the record. Every plan, review, run, decision, and click lands on the ticket’s case file, and every correction a person makes teaches the Team: a sent-back plan becomes a Playbook rule, a product call lands in the decision ledger, an accepted proposal lands on the track record.

Then, and only then, decide what runs without a click. The Ladder has two rungs, supervised and autonomous, set per line by a person who has read that line’s track record. Merge is never on it. A supervised ticket takes three clicks, accept the proposal, approve the plan, merge, and nothing moves a line to autonomous except a person’s explicit act with the record beside it.

That is agentic engineering as an operating practice rather than a term: agents do the implementation, people make the decisions, and the record shows both. It is the practice a YAGNI Team runs by default, managed like your engineering team, with Bailey, Reeve, Wright, Proctor, Fletcher (beta), and Harper each on the hook for one stage. Book 30 minutes to walk one of your own tickets through it.