User acceptance testing when an agent built the feature
Stefan-Iulian Tesoi · · 6 min read

The person who owns the acceptance criteria, on a build they can actually use, joined where it matters by one or two of the people the feature is for. A coding agent can show that the criteria it was given pass. User acceptance testing asks whether they were the right criteria, and the author of the code cannot answer that, human or agent.
What does user acceptance testing check that tests do not?
Whether the work meets the need rather than the specification. The ISTQB glossary defines user acceptance testing as "a type of acceptance testing performed to determine if intended users accept the system." The words that matter are intended users. A test suite answers to the specification; UAT answers to the people it was written for.
The same glossary separates verification, "the process of confirming that a work product fulfills its specification", from validation, "confirmation by examination that a work product matches a stakeholder's needs." Everything a coding agent hands back is verification, the evidence a definition of done for agent work asks for. None of it is validation, because the agent's only source for the need was the item.
That is also the short answer to UAT vs QA testing. What teams call QA testing is usually closer to the glossary's system testing, which focuses on "verifying that a system as a whole meets specified requirements." It still answers to the requirements.
| Agent's own checks | QA testing | User acceptance testing | |
|---|---|---|---|
| Question | Does the change meet its criteria? | Does the system meet its requirements? | Do the intended users accept it? |
| Kind of check | Verification | Verification | Validation |
| Done by | The coding agent | Testers | The criteria owner, with users |
| Catches | Code that misses the criteria | Breakage between parts | Criteria that were wrong or missing |
Who does acceptance testing when the author is an agent?
The person who wrote or accepted the acceptance criteria, usually the product manager or product owner. They held the conversation the item was distilled from, so only they can notice what was lost on the way.
Three substitutes are tempting, and each fails:
- The coding agent. It judges the result against the item it built from, so a gap in the item is invisible to it twice.
- The engineer who reviewed the diff. Clicking through afterwards reuses the reading the review already made.
- A QA lead alone. A tester can confirm behaviour against the written criteria, not that they were the right ones.
Real users add what the owner cannot: not knowing what the feature was meant to do, they use it the way it will be used. Bring one or two in for daily workflows; a renamed button needs only the owner.
How do you run UAT at agent pace without becoming the bottleneck?
Test only what a user can see, on a build of that one item, with the steps written before the agent starts. Take a hypothetical team of five whose agents finish forty items a week, twenty-six of them refactors, dependency bumps and test fixes with no visible change. The other fourteen, at ten minutes each, cost the owner under two and a half hours a week. All forty at half an hour each, on a shared staging environment, would cost twenty.
Three habits keep the UAT process that size:
- Write the UAT steps into the item. Two or three lines in the form of criteria an agent can verify, aimed at a person. Gherkin has the right shape even unautomated: Cucumber's reference wants each outcome to be "something that comes out of the system (report, user interface, message)".
- Test one item per build. A shared staging environment collects a day's merges, so a failure there implicates all of them. A per-branch preview isolates one change.
- Use a feature flag when production is the only realistic environment. Pete Hodgson's article on feature toggles calls turning new features on "for a set of internal or beta users" a Champagne Brunch. For UAT, that set is the owner and the invited users.
Laimonade's form of the second habit is a hosted preview. A branch pushed for a backlog item is built on the team's own Cloudflare or Fly.io account, and the URL appears on the item, on the pull request and in the In Review notification sent to the item's testers. A coding agent submitting work can also list steps for a person to check, shown on the item with its automated results. The preview is removed when the pull request is merged, so it is for testing before merge.
When UAT fails, was it the criteria or the build?
Run the item's own criteria on the build under test: if they fail there while passing in CI, the build is wrong; if they pass and the feature is still wrong, the criteria are. General triage for work that misses its criteria applies, and UAT adds three outcomes of its own:
| What UAT found | What it means | What to do |
|---|---|---|
| A criterion fails on the preview but passed in CI | The tests exercised a mock, a fixture or a missing setting | Back to the builder, with the failing step |
| Every criterion passes; users still cannot do the job | The criteria were wrong | Accept the item, keep the flag off, write a new item |
| It works but breaks an unwritten expectation | A missing criterion | Fix it in a follow-up; add it to the item template |
The second row is where a feature flag earns its place. It separates two decisions that usually travel together: whether the item was built as specified, and whether the feature should reach users. An item can be closed honestly while the feature stays dark.
A coding agent can prove it built what the item said. User acceptance testing is how a team finds out whether the item said the right thing.
What should be recorded when a feature passes UAT?
Enough that someone a month later can tell who accepted what, on which build.
- Who tested, by name, and whether real users took part.
- The build: the preview URL or environment, and its commit.
- The steps run, ideally the ones written into the item.
- What was not tested: devices, roles, data volumes.
- The release decision: merged, flagged on for a named group, or held.
The commit is the line most often missing; an approval recorded against a branch means nothing once the branch moves. In Laimonade each preview on an item lists the commit it was built from, and no tool a connected coding agent can call marks an item Done, so an item is closed by a person, or by Laimon once nothing is left for a person to check, and the board shows who closed what. The sprint workflow sets out who can make each move.
Frequently asked questions
Can an agent run the UAT script itself?
It can perform the steps, and doing so before a person looks is a useful smoke test. It cannot accept the result. Acceptance is a judgement against a need the agent knows only through the item, so record its run as an automated check, not as acceptance.
Should UAT happen before or after merge?
Before, wherever a build of the branch can be produced. A failure found before merge costs one rework cycle; after merge it costs a revert or a hotfix, and other work may already sit on top of it. When production is the only realistic environment, merge behind a feature flag that is off by default and test there with a named group.
Is UAT still needed if the acceptance criteria were thorough?
For anything a user sees, yes. Thorough criteria make it short, sometimes five minutes, but they were written before the feature existed, and using it reveals what nobody thought to write down. For a refactor or a dependency bump with no visible behaviour, verification is enough.