← Back to blog

What Is Spec-Driven Development with AI Coding Agents?

Spec-driven development with AI: what the spec is, who approves it on a team, what the review checks it against, and what the record shows after merge.

Search “spec-driven development ai” and the first page is a tour of tooling: GitHub’s Spec Kit with its specify, plan, and tasks commands; Kiro with its requirements, design, and tasks files; Claude Code with plan mode; and a handful of plugins that bolt a spec phase onto whichever agent you already run. Every page agrees on the definition, contrasts it with vibe coding, and says a human should review the spec before the agent implements it.

Then every one of them stops, because the next questions are about the team rather than the tool: who approves the spec, what they read before they do, what happens when the implementation drifts from it, and what exists on paper six weeks later when someone asks why the change was made. This post is about those questions.

What is spec-driven development with AI?

Spec-driven development is a workflow where the specification is the artifact a coding agent works from, and the prompt is only how the specification gets written. Someone states what a change should do and, usually, how: the scope, the approach, the files it expects to touch, the tests it will add, what it is leaving alone. A person approves that document. The agent implements against it. The code is checked against the spec, not against anyone’s memory of a conversation.

The reason it took off in 2025 is simple. A human engineer given a vague ticket asks a question. A coding agent given a vague ticket builds something, confidently and fast, and the cost of that confidence lands on whoever reviews the pull request. Moving the decisions into a document the agent reads is the only way to review them before they turn into code.

That makes spec-driven development the written form of a broader practice: agents do the implementation, and a person owns the plan, the quality bar, and the merge. The agentic software development lifecycle walks all six stages of that practice. This post stays on the first two, Proposed and Plan, because the spec is where they live.

How is it different from vibe coding?

Vibe coding puts a prompt before the code and nothing after it but the diff. Spec-driven development puts a reviewed document before the code and keeps it beside the pull request afterwards.

Both are legitimate. Vibe coding is the right way to build a throwaway, because the whole point of a throwaway is that nobody needs to understand it later. The failure is vibe coding something a customer or a colleague will depend on and finding out at review time, or after. Spec-driven development exists for the second kind of change, and it works because the spec moves the first chance to say no from the finished diff to a page that takes five minutes to read.

The comparison is really about who reads what, and when. A prompt is read by the agent. A spec is read by the agent and by a person before the code exists, and again by the reviewer after it does. That second reader is the whole difference, and it is the part that does not come in a toolkit.

What do Spec Kit, Kiro, and Claude Code actually give you?

Three of the tools on the first page, and what each one puts in front of a person, read from their own documentation as of October 2026.

What the spec is Where the checkpoint is Who approves What the record shows after merge
GitHub Spec Kit Markdown produced by /specify, /plan, and /tasks, in the repository Between each phase; the developer is told to verify before moving on The developer running the agent The spec files, if they were committed
Kiro requirements.md, design.md, tasks.md, generated in order Between files in the standard flow; a Quick Spec generates all three with no gates The developer in the IDE The three files in the repository
Claude Code plan mode A proposed plan in the session; the agent makes no edits until it is accepted Before the first edit The developer in the terminal Whatever the developer saved; the plan itself lives in the session
YAGNI Team Reeve’s plan on the ticket’s case file: scope, approach, files, tests, what it leaves alone, with its check beside it A Decision addressed to a named member of the Team, before any code The Team’s owner or the line’s supervisor, by name, with the time on the record The plan, its check, the questions and answers, the approval, Proctor’s review, Fletcher’s evidence, the merge, and what each run cost

The first three rows are good tools, and none of this is a case against them. Spec Kit supports Claude Code, Copilot, and Gemini CLI and is the most portable way to add a spec phase to an agent you already run. Kiro’s three files are a clean shape for a spec. Plan mode is the lightest version, and the one most developers already have.

What the table shows is that all three answer the same question: how does one developer make one agent write a spec before it writes code. The approver is the developer at the keyboard, and the record is whatever they committed. For one person that is the right answer. For an engineering manager with eight developers each running an agent, it means eight people approving their own specs with nothing a manager can read. That is the problem the rest of this post is about.

Who approves the spec on a team?

A named person who is accountable for the scope, and not the person, or the agent, who wrote it.

This is the rule every engineering org already has for human work: a design doc has an author and a reviewer, and they are different people. Spec-driven development with an agent breaks the rule quietly. The agent writes the spec and the developer who prompted it approves it, often in the same session, often without reading past the first screen. The spec exists, so the box is ticked, but nobody independent has read it.

On a YAGNI Team the roles are separated by construction: one Agent per stage. Reeve, the Staff Engineer, drafts the plan in a sandbox against the real repository, with the hard-to-reverse risks in hand, and checks it before the gate opens: the module the plan missed, the validation layer it was about to bypass, the simpler path it walked past. Reeve never builds. Then the plan, with that check beside it, becomes a Decision addressed to a named member of the Team, by default the Team’s owner. That person reads it and clicks Approve plan or Send back. Only then does Wright, the Builder, write code.

Three things fall out of that shape. The author is not the approver. The approver reads a check, not just the plan, so they are not the only skeptic in the room. And the approval has a name and a time on it, which is what “a human reviewed the spec” means when someone asks later.

What does a spec look like when an Agent writes it?

Here is one ticket, because the artifacts are easier to see with the roles named. The Team owns an orders service on GitHub, reads a Linear project, and posts to a channel in Slack. Its owner is Noor, an engineering manager. The ticket is ORD-331, “Expire unpaid orders after thirty minutes.”

Proposed. Bailey, the Product Manager, proposes ORD-331 from the backlog with the evidence: a week of support threads about reserved stock that never cleared, and the two tickets behind it that cannot start until it ships. Bailey never ships anything, and a proposal always waits for a person. Noor reads it and clicks Start plan.

Plan. Reeve reads the repository and the Team’s decision ledger and drafts the plan on the case file:

  • Scope: unpaid orders older than thirty minutes move to an expired state and release their stock reservation. Nothing else changes.
  • Approach: a scheduled job over the orders table, batched, idempotent, with the threshold in configuration rather than in code.
  • Files: the orders model, the job runner the service already has, the configuration schema, the admin order view.
  • Tests: an order at twenty-nine minutes stays, an order at thirty-one expires, an expired order cannot be paid, a re-run of the job changes nothing.
  • Leaving alone: the payment webhook, the checkout flow, and the nightly reconciliation job.

One question comes back to Noor as a tap on the case file: should an order with a pending bank transfer expire on the same clock? She answers no, those get seven days, and that answer lands in the Team’s decision ledger where every Agent reads it before the next ticket. Reeve’s own check catches a gap before the gate opens: the admin order view has a filter on state that the first draft did not mention, and the new state needs to appear in it or expired orders vanish from the admin’s list. The plan is redrafted to include it. Noor reads the plan and its check, and clicks Approve plan.

That is the spec. It took her a few minutes to read, it was written against the real code, a second Agent found the gap, a product decision was recorded where it will be reused, and the approval carries her name. Nobody has written a line of application code yet.

What happens when the code does not match the spec?

The spec is only worth having if something checks the implementation against it, and a person is a poor choice for that job at thirty pull requests a day.

On a YAGNI Team the check is Proctor’s. Wright implements the approved plan on a branch in a sandbox and opens a draft pull request through the workspace’s GitHub App. Proctor reads that pull request with the approved plan in front of it, through its checks: deep review, business fit, database, security, tests. Business fit is the one that matters here. It asks whether the change does what the plan said and nothing the plan said to leave alone. If Wright touched the payment webhook, Proctor says so, and Wright answers the finding in a follow-up commit. Proctor posts one review with a verdict and the findings it set aside with the reasons, and it never merges. Who reviews AI-written code is the longer argument for why that one readable review matters more than a wall of inline comments.

Then Fletcher, in beta, boots the Team’s repository environment at that revision, writes a test plan from the change and the spec, walks the admin view in a browser with an order past the threshold, and attaches the video and screenshots. Fletcher never edits application code.

Drift from the spec is caught by an Agent whose job is to catch it, with the spec adjacent, and surfaced to a person as one page. The person reads that page before merging, and merge is always a person’s click.

What does the spec cost, and what does it save?

A spec costs a draft. That is the honest number, and it is why the tooling pages are right that the overhead is small. What they do not say is where the saving comes from, because most of them have no cost story at all.

The saving is in retries. Most of what an agent costs per change is not the first attempt but the second and third, after the first one solved the wrong problem, read the scope too broadly, or picked a path the team would not have chosen. A plan sent back at the gate costs one more draft. The same mistake caught in review costs a build. Caught after merge, it costs an incident. Spec-driven development moves the correction to the cheapest stage, and on a YAGNI ticket every run on the case file shows what it cost, so a manager sees that arithmetic per ticket rather than taking it on faith.

Two more levers sit underneath. Grounding lowers the number of wrong plans: an Agent that reads the decision ledger and the connected tools before it drafts does not need Noor to re-answer the bank-transfer question on the next ticket. And the models: YAGNI routes the Agents, and YAGNI Code for the developer, across vetted US-hosted open-weight models at 60% or more under comparable frontier API rates, with a router that places each step on the cheapest lane that holds quality. Usage shows the result per developer and per Agent, broken out by the job that spent it.

How do you start spec-driven development on a team?

In the order of the artifacts, whichever tools you use.

  1. Make the plan a document. If your developers run Claude Code, turn on plan mode or add Spec Kit, and commit the plan with the change rather than leaving it in the session. YAGNI is additive here: YAGNI Code is the developer’s own coding agent on the same models, meter, and record as the Teams, and it fronts Claude Code and Codex through its wire adapters, so the same session can run on YAGNI’s models without changing the workflow.
  2. Separate the author from the approver. A spec the author approves is a prompt with a title. Name who approves plans for each scope. On a YAGNI Team that is the Team’s owner or the line’s supervisor, and the gate is addressed to them by name.
  3. Check the plan before the approver reads it. Reeve checks its own plan before the gate opens, which is why the plan gate on a Team is cheap for the person reading it: the obvious gaps are already marked. Without a check, the approver is the only skeptic, and attention does not parallelize the way agents do.
  4. Check the code against the spec, not against memory. Proctor’s business-fit check reads the pull request with the approved plan beside it. If you build this yourself, the reviewer needs the spec in the same window as the diff.
  5. Keep the record. The plan, its check, the decisions, the review, the evidence, the clicks, and the cost, on one case file per ticket. Corrections become Playbook rules the Team adopts, and product calls land in the decision ledger. That is how a Team gets trained, in the earned sense.
  6. Decide what runs without a click only after that. The Ladder has two rungs, supervised and autonomous, set per line by a person who has read that line’s track record. A supervised ticket takes three clicks, accept the proposal, approve the plan, merge, and merge is never on the Ladder. Managing AI coding agents covers the management side, and what autonomous coding agents should own draws the line per act.

So what is spec-driven development with AI, in practice?

A document a person approves before an agent writes code, written against the real repository, checked before it is approved, checked against the implementation after, and kept beside the merge with the cost. The tooling pages give you the first clause. A team needs the rest.

That is what the Plan stage on a YAGNI Team is: Reeve plans, a named person approves, Wright builds, Proctor reads the pull request against the plan, Fletcher (beta) walks it, a person merges, and Harper names it in the next morning’s Daily Brief in Slack with the pull request number beside it. See how agent teams work, or book 30 minutes to walk one of your own tickets through the plan gate.