← Back to blog

What Is AI QA Testing, and What Should a Testing Agent Do?

AI QA testing for an engineering team: what a testing agent may do on a pull request, what it never does, when it runs, who owns the verdict, and the cost.

Search “ai qa testing” and the first page is vendors describing their own features: tests generated from plain language, locators that heal themselves when the page changes, failures triaged into buckets, a two-week adoption plan, and a reassurance that AI will not replace your QA engineers. Search “ai testing agent” and half the results are about the opposite thing, how to test an AI agent you built.

None of the pages answer the question an engineering manager is actually asking before they switch one on: what is this agent allowed to do to my pull request, what must it never do, when does it run relative to code review, who owns the verdict, where does the evidence live six weeks later, and what does it cost per change. Those are the questions you would ask about a new QA engineer. This post answers them for an AI testing agent, the way a YAGNI Team runs one.

What is AI QA testing, and what is an AI testing agent?

AI QA testing is quality assurance where an agent does the testing work a person would otherwise do by hand: reads what changed, decides what to check, runs the application, exercises it, and reports what it saw. The ranking guides split that into features (test generation, self-healing, triage), which is how a vendor sees it. An engineering manager sees it as a job on the team, with inputs, an output, and a boundary.

An AI testing agent is the thing that does that job. It is a teammate in the sense that it takes an assignment, a pull request, and returns an artifact, evidence with a verdict. It is not a teammate in the sense that it gets a vote on the merge.

One disambiguation, because Google blends the two: this post is about an agent that tests software. It is not about testing AI agents, the evaluation of a chatbot or model, which shares the words and nothing else.

On a YAGNI Team, the QA Agent is Fletcher, in beta, one of six named Agents with one craft each: Bailey proposes the next ticket, Reeve plans it, Wright builds it to a draft pull request, Proctor reviews, Fletcher tests, and Harper reports what shipped. The agentic software development lifecycle walks all six stages; this post stays on QA.

What should an AI testing agent be allowed to do?

A short list, and each item is something a QA engineer would recognize as the job.

Read the change and the ticket. The pull request diff, the ticket it implements, and the approved plan. A test plan written from the diff alone misses what the change was for; a plan written from the ticket alone misses what the change actually touched. Fletcher reads both, and the ticket it reads is the one Bailey wrote as a spec and a person accepted.

Boot the application at the exact revision. Not a shared staging server where three other branches are also deployed. Fletcher boots the Team’s repository environment in a sandbox at the pull request’s commit, with the dependency repositories pinned so a web app tests against the API version the record says it did.

Write a test plan from the change. Deterministic checks first, then a bounded exploration of the change in a browser, with a step budget a person set on the line. The default check, when an environment names none, is deliberately weak and marked as defaulted: the entry page renders visible text. A person sets stronger ones.

Run the tests and walk the app. Click through the feature the change introduces, with the inputs the ticket implies and a few it does not.

Record what it saw. Screenshots at each step, a video of the walk, the test results, and a verdict: pass or fail, with the reason.

That is the whole mandate. Everything on it produces evidence and changes nothing in the repository.

What should it never do?

This is the list none of the ranking pages have, and it is the one that makes the agent safe to switch on.

  • Never edit application code. The moment the agent that tests a change can also edit it, the evidence stops being independent. A failing walk on a YAGNI Team is addressed to a person, who sends the ticket back to Wright with the video attached. Fletcher never writes a line of application code.
  • Never weaken a test. An agent under pressure to pass will loosen an assertion if it is allowed to. Fletcher is not allowed to.
  • Never merge, and never mark the pull request ready. Merge is always a person’s click. On a YAGNI Team there is no merge line and no autonomous merge; Proctor’s approving review marks a draft ready, and a person merges. Fletcher’s verdict is one more thing on the case file when they do.
  • Never hold a credential it could reuse. Secrets for the environment are held server-side, injected into the sandbox at run time, kept out of the video and the screenshots, and never placed in a model prompt. The sandbox is destroyed when the run ends. Fletcher reaches repositories only through the workspace’s GitHub App installation, scoped to the repositories the Team attaches; there is no personal GitHub connection to expire or revoke.
  • Never run where it has no environment. Fletcher’s line on a Team stays off until the repository environment is Working. An agent that fakes a walk against an app that did not boot produces evidence that looks like a pass.

When should it run: before or after code review?

After the review says the change is ready, and before a person merges.

The argument is cost. A browser walk is the most expensive thing on the ticket after the build itself: an environment to boot, an app to start, a model watching a browser for a bounded number of steps. Spending that on a pull request the reviewer is about to send back for a missing null check wastes it. Review first, and most of what review catches never reaches QA.

On a YAGNI Team that order is wired in. Wright opens a draft pull request. Proctor reviews it and posts one review with a verdict. An approving verdict marks the pull request ready for review, and that event is what starts Fletcher. AI code review for GitHub covers Proctor’s half of that handoff. The person who merges sees Proctor’s review and Fletcher’s evidence side by side on the Review case file, and the stage word on Work reads QA while Fletcher is working, Done only when the merge webhook arrives.

Where does the evidence live, and who reads it?

Two places, for two readers.

On the pull request, for the engineer: a pass or a fail with the evidence attached, where the reviewer and the author already are. A fail with a video is a bug report nobody had to write.

On the ticket’s case file, for the manager: the test plan, the video, the screenshots, the test results, the verdict, beside the proposal, the plan, Proctor’s review, the runs, the clicks a person made, and what each of them cost. Six weeks later, when someone asks why the change shipped, the case file is the answer, and Fletcher’s walk is part of it.

The reader who matters is the one who merges. Fletcher’s verdict is advice to that person. If they disagree with a fail, they can say so and merge; if they disagree with a pass, the video shows what the agent actually checked. A verdict with the video behind it is a report, not a vote.

What does one pull request look like, end to end?

A Team owns a billing service. Its repository environment is Working: the web app, the API it depends on pinned to a fixed commit, Postgres as a managed service, and a private sign-in with a dedicated test user. Fletcher’s verify line is on. The engineering manager, Priya, has not moved the build line to autonomous yet.

Stage What happened What Fletcher does What Priya reads
Plan Reeve planned “credit notes over the invoice total are rejected with a message”, Priya approved it Nothing yet The plan
Build Wright implemented it in a sandbox and opened a draft pull request Nothing yet The runs
Review Proctor posted one review, approving, with one finding set aside and why; the verdict marked the pull request ready The ready event starts the run The review
QA Fletcher boots the environment at the pull request’s commit, reads the ticket and the diff, writes a test plan, runs the deterministic checks, signs in as the test user, walks the credit note form with a value over the total, a value equal to it, and a negative value Attaches the video, the screenshots, the results, and a fail: the negative value was accepted The video, the fail, the reason
Build, again Priya sends the ticket back; Wright adds the lower bound and a test, pushes; Proctor re-reviews and marks it ready Runs again on the new commit, passes The second video
Done Priya merges on GitHub Nothing. The webhook moves the ticket to Done The Receipt

Four clicks of Priya’s: approve the plan, send back, and merge, plus the one her supervised build line asks for. The negative value was in neither the ticket nor the plan. It is the input a QA engineer tries because they have seen a form before, and the walk found it because it is bounded exploration, not a replay of a script. Harper’s Daily Brief in Slack names the pull request the next morning.

What does it need from you to work?

A repository environment, which is the part the ranking pages gloss as “integrates with your CI/CD”. Setup is one action under Connections: open the repository the GitHub App covers and click Begin setup. YAGNI reads the repository, fills in the commands, services, and variables it finds, boots it in a disposable sandbox, and tries up to three repairs inside the same run when the boot fails on something it can fix. A run that stops on something only you have, a secret or a destination to approve, reads Needs you with the fix beside it. You add the dependency repositories with their pins, the secrets, and the sign-in mode. The first configuration that boots clean goes live on its own, and from then on Wright builds in it and Fletcher walks every pull request on the Team’s repositories that is marked ready. A Team can also switch on a second Fletcher line that signs off a staging deploy, with the same kind of evidence.

Two honest limits. Fletcher is in beta: setup, the browser walk, and the evidence on the pull request work today, and you should switch it on deliberately, one Team at a time, rather than everywhere at once. And a change with no surface to walk, a refactor, a background job, gets a short QA stage: the deterministic checks run, there is nothing to click, and the record says so.

How is this different from the test-automation tools that rank?

The ranking pages are mostly test-automation platforms that added AI, and a few newer agents. They are good at what they describe: generating tests from requirements, keeping scripted tests alive when the UI changes, and triaging failures. Momentic’s guide, for example, notes that every change its agent makes lands as a file change in your repository that goes through your pull request process (as of October 2026), and Katalon’s guide says AI proposes while your team reviews and approves (as of October 2026). Those are the right instincts. The table is what a manager should ask any of them to put in writing.

Criterion Test-automation platform with AI features Standalone AI testing agent Fletcher on a YAGNI Team
Who writes the test plan Your QA team, with AI drafting from requirements The agent, from the app or the diff Fletcher, from the diff, the ticket, and the approved plan, with checks a person sets on the line
Can it edit application code Usually no; some write test files to the repo Varies; often yes to fix tests Never. Not application code, not a test assertion
When it runs Scheduled suites, or on every commit in CI On every pull request, or on demand When the pull request is marked ready, after Proctor’s approving review
Who owns the verdict The suite passes or fails; a person interprets The agent reports; ownership varies A person merges; Fletcher’s verdict is evidence on the case file, with the video
Where the evidence lives The platform’s dashboard The agent’s dashboard, sometimes the pull request The pull request and the ticket’s case file, beside the review and the cost
Cost unit Seats, or a flat monthly fee A flat monthly fee or credits Per run, on the ticket’s case file and on Usage per Agent
Accountable owner Whoever administers the platform Whoever installed the agent The Team’s owner, one named person, on the Team page

The right column is not better at generating tests than the left two. What it adds is the management frame: a named Agent with a written boundary, a fixed place in the pipeline, evidence a person reads before a click that stays theirs, and a cost that lands on the ticket it tested. If you are already managing AI coding agents, that frame is the same one, applied to QA.

What does AI QA testing cost?

A QA run has three costs: the environment it boots, the model that plans and watches the walk, and the person who reads the result. The third is small only if the evidence is good enough that reading it is faster than reproducing the bug.

Most tools price the first two as a seat, a flat monthly fee, or a credit pool, which is a fine way to pay for a product and a poor way to answer “what did testing that pull request cost us”. On a YAGNI Team each of Fletcher’s runs lands on the ticket’s case file with what the workspace was charged, and Usage adds it up per Agent, broken out by the line that spent it, on the one rate card that YAGNI Code developers and Agent Teams share. The Agents run on vetted US-hosted open-weight models at 60% or more under comparable frontier API rates, with a router that places each step on the cheapest lane that holds quality. The rate card itself is shared on a call, not published.

The levers a manager controls are the step budget on Fletcher’s line and the order of operations: review before QA, so the walk is spent on changes that are about to merge.

So what should an AI QA testing agent do?

Read the change and the ticket, boot the application at that exact revision, write a test plan, run it, walk the change in a browser, and attach the video, the screenshots, and a verdict to the pull request and the ticket. It should run after review says the change is ready and before a person merges. It should never edit application code, never weaken a test, never merge, and never hold a credential it could reuse. The verdict is evidence for the person who merges, not a vote.

YAGNI runs agent teams, managed like your engineering team: Bailey proposes, Reeve plans, Wright builds to a draft pull request, Proctor reviews, Fletcher (beta) walks the change in a browser with the evidence attached, and Harper reports to Slack, with merge always a person’s click. The security page states what runs where. See how Agent Teams work, or book 30 minutes to put a Team, with Fletcher on it, on one of your repositories.