User acceptance testing when an agent built the feature

Stefan-Iulian Tesoi · · 6 min read

A cloth tailor's dress form seen from behind against a dark wall, the stand-in a garment can fit exactly and still not fit the person it was made for

The person who owns the acceptance criteria, on a build they can actually use, joined where it matters by one or two of the people the feature is for. A coding agent can show that the criteria it was given pass. User acceptance testing asks whether they were the right criteria, and the author of the code cannot answer that, human or agent.

What does user acceptance testing check that tests do not?

Whether the work meets the need rather than the specification. The ISTQB glossary defines user acceptance testing as "a type of acceptance testing performed to determine if intended users accept the system." The words that matter are intended users. A test suite answers to the specification; UAT answers to the people it was written for.

The same glossary separates verification, "the process of confirming that a work product fulfills its specification", from validation, "confirmation by examination that a work product matches a stakeholder's needs." Everything a coding agent hands back is verification, the evidence a definition of done for agent work asks for. None of it is validation, because the agent's only source for the need was the item.

That is also the short answer to UAT vs QA testing. What teams call QA testing is usually closer to the glossary's system testing, which focuses on "verifying that a system as a whole meets specified requirements." It still answers to the requirements.

Agent's own checksQA testingUser acceptance testing
QuestionDoes the change meet its criteria?Does the system meet its requirements?Do the intended users accept it?
Kind of checkVerificationVerificationValidation
Done byThe coding agentTestersThe criteria owner, with users
CatchesCode that misses the criteriaBreakage between partsCriteria that were wrong or missing

Who does acceptance testing when the author is an agent?

The person who wrote or accepted the acceptance criteria, usually the product manager or product owner. They held the conversation the item was distilled from, so only they can notice what was lost on the way.

Three substitutes are tempting, and each fails:

Real users add what the owner cannot: not knowing what the feature was meant to do, they use it the way it will be used. Bring one or two in for daily workflows; a renamed button needs only the owner.

How do you run UAT at agent pace without becoming the bottleneck?

Test only what a user can see, on a build of that one item, with the steps written before the agent starts. Take a hypothetical team of five whose agents finish forty items a week, twenty-six of them refactors, dependency bumps and test fixes with no visible change. The other fourteen, at ten minutes each, cost the owner under two and a half hours a week. All forty at half an hour each, on a shared staging environment, would cost twenty.

Three habits keep the UAT process that size:

  1. Write the UAT steps into the item. Two or three lines in the form of criteria an agent can verify, aimed at a person. Gherkin has the right shape even unautomated: Cucumber's reference wants each outcome to be "something that comes out of the system (report, user interface, message)".
  2. Test one item per build. A shared staging environment collects a day's merges, so a failure there implicates all of them. A per-branch preview isolates one change.
  3. Use a feature flag when production is the only realistic environment. Pete Hodgson's article on feature toggles calls turning new features on "for a set of internal or beta users" a Champagne Brunch. For UAT, that set is the owner and the invited users.

Laimonade's form of the second habit is a hosted preview. A branch pushed for a backlog item is built on the team's own Cloudflare or Fly.io account, and the URL appears on the item, on the pull request and in the In Review notification sent to the item's testers. A coding agent submitting work can also list steps for a person to check, shown on the item with its automated results. The preview is removed when the pull request is merged, so it is for testing before merge.

When UAT fails, was it the criteria or the build?

Run the item's own criteria on the build under test: if they fail there while passing in CI, the build is wrong; if they pass and the feature is still wrong, the criteria are. General triage for work that misses its criteria applies, and UAT adds three outcomes of its own:

What UAT foundWhat it meansWhat to do
A criterion fails on the preview but passed in CIThe tests exercised a mock, a fixture or a missing settingBack to the builder, with the failing step
Every criterion passes; users still cannot do the jobThe criteria were wrongAccept the item, keep the flag off, write a new item
It works but breaks an unwritten expectationA missing criterionFix it in a follow-up; add it to the item template

The second row is where a feature flag earns its place. It separates two decisions that usually travel together: whether the item was built as specified, and whether the feature should reach users. An item can be closed honestly while the feature stays dark.

A coding agent can prove it built what the item said. User acceptance testing is how a team finds out whether the item said the right thing.

What should be recorded when a feature passes UAT?

Enough that someone a month later can tell who accepted what, on which build.

The commit is the line most often missing; an approval recorded against a branch means nothing once the branch moves. In Laimonade each preview on an item lists the commit it was built from, and no tool a connected coding agent can call marks an item Done, so an item is closed by a person, or by Laimon once nothing is left for a person to check, and the board shows who closed what. The sprint workflow sets out who can make each move.

Frequently asked questions

Can an agent run the UAT script itself?

It can perform the steps, and doing so before a person looks is a useful smoke test. It cannot accept the result. Acceptance is a judgement against a need the agent knows only through the item, so record its run as an automated check, not as acceptance.

Should UAT happen before or after merge?

Before, wherever a build of the branch can be produced. A failure found before merge costs one rework cycle; after merge it costs a revert or a hotfix, and other work may already sit on top of it. When production is the only realistic environment, merge behind a feature flag that is off by default and test there with a named group.

Is UAT still needed if the acceptance criteria were thorough?

For anything a user sees, yes. Thorough criteria make it short, sometimes five minutes, but they were written before the feature existed, and using it reveals what nobody thought to write down. For a refactor or a dependency bump with no visible behaviour, verification is enough.