AI-generated tests: what they prove and what they do not
Stefan-Iulian Tesoi · · 6 min read

Only the ones that fail without the change. A test written by the agent that wrote the code shares that code's assumptions, so the check worth running is the new tests against the old code: a test that passes there too proves nothing about what changed.
That one check sorts AI-generated tests into evidence and decoration, and it takes a few minutes. Coverage cannot do the same job, because a line can be executed by a test that asserts nothing about it.
Why are AI-generated tests weak evidence on their own?
Because the author's tests encode the author's reading of the requirement. When a person writes both the code and the tests, a misunderstanding goes into both, and the suite agrees with the bug. A coding agent does the same thing faster: if it read "handle the error case" as "log and continue", its tests assert that the log line was written, and they pass.
Three shapes recur:
- The mirror test. It restates the implementation: call the function, assert it returned what the function returns. It passes on any version that does not throw.
- The mocked-away test. The dependency the change was about is replaced by a mock, so the test exercises the mock.
- The test of old behaviour. It passes before and after the change, because it covers a part of the code the change never touched.
None of this is dishonest, and agent-written unit tests are usually competent at whatever they test. The trouble is what they are offered as: proof that the change works. The rest of that hand-back is covered in what evidence a coding agent should return; this is the item in it that most often proves less than it claims.
A test that passes before the change and after it is not evidence about the change. It may still be a good test. It is evidence of something else.
What check separates a test from a decoration?
Run the new tests against the code as it was before the change. Tests that cover the change should fail there; tests that guard behaviour the change must not disturb should pass. Anything else needs an explanation.
Mechanically: keep the new test files, restore the previous version of the source files, run only the new tests, then restore the change. In git that is git show HEAD~1:<path> > <path> for each source file touched, or a worktree at the parent commit with the tests copied in. For a typical change it is five minutes.
The result sorts into four cases:
| On the old code | The test claims to cover | Verdict |
|---|---|---|
| Fails | The change | Evidence. Keep it |
| Passes | Behaviour that must not change | A guard. Keep it, and label it as one |
| Passes | The change | Proves nothing about the change. Rewrite it |
| Fails | Behaviour that must not change | The change broke something, or the test is wrong |
A real run, from a fix to this product's own model client in September 2026: 19 tests in the file, 10 failed against the previous commit and 9 passed on both. The 9 were guards (an ordinary stream still streams; a caller that cancels is not handed to a fallback model) and are meant to pass on both. The count is what mattered. Without it, "19 tests pass" reads as nineteen pieces of evidence for the change, when ten were.
This is the discipline of watching a test fail first, borrowed from test-driven development and applied after the fact. A test that fails without the change is the only kind that says anything about the change.
What should the hand-back include?
Three things beyond "tests pass", and each is cheap for an agent to produce:
- The before-and-after count. How many of the new tests fail on the old code, and which are guards. A number rather than an assurance.
- The command that produced it, so a reviewer can re-run it instead of trusting it.
- A mapping from each acceptance criterion to the test that covers it. A criterion with no test either needs one or needs a person to check it by hand, and the hand-back should say which. How to write criteria that make the mapping possible is set out in acceptance criteria an agent can verify.
With those three, the review shrinks to reading the table and spot-checking one test from the evidence row. Verifying AI-written code this way is fast because the agent did the tedious part and the reviewer only checks that it was done honestly. It belongs in the team's definition of done as a line of its own, not as an interpretation of "tested".
Laimonade's hand-back records each check as a command, a result and a count, so the reviewer reads numbers rather than a summary. A coding agent connected over MCP cannot mark its own work done; the item waits in review for that reading.
Where do coverage and mutation testing fit?
Coverage says which lines ran, not whether anything checked them. A mirror test can take a function to full coverage while asserting nothing useful, so a coverage figure from the agent that wrote the code is the weakest item in the hand-back. It is still worth having for what it does show: code that no test reaches at all.
Mutation testing is the systematic version of the before-and-after check. The tool makes small deliberate changes to the code, such as flipping a comparison or deleting a line, and counts how many of them the tests notice. A surviving mutant is a line the tests execute but do not check. It is slow, minutes to hours on a real suite, so it earns its place on the modules where a wrong answer is expensive rather than on every change.
The before-and-after check is the cheap special case: one mutation, the actual change, reverted. It answers the question a reviewer is really asking.
One more trap sits under all of it. A suite that selects zero tests exits 0 in CI, which is the empty-selector case in quality gates that fail quietly. Ask for the test count in the output, not just the exit code.
Frequently asked questions
Should the agent write tests before the code?
It helps, for the same reason it helps a person: a test written first and seen to fail has already passed the before-and-after check. The catch is that an agent asked to work test-first can write the test and the code in one pass and never run the failing state. The evidence is a recorded failing run, not the order the files were written in.
Is test coverage useless now?
No, but it answers a different question. Coverage finds code that no test touches, which is still worth knowing. It cannot tell a test that checks a line from one that merely runs it, and agent-written suites widen that gap because producing tests has become cheap. Use it to find gaps, not to prove a change.
Who should write the acceptance tests?
Whoever wrote the acceptance criteria, or someone reading them without seeing the implementation. A test derived from the criterion rather than the code does not share the code's assumptions, which is the property the author's own tests lack. For many items, the command the criterion names already is that test.