Agent work rejected: what to do when criteria are missed
Stefan-Iulian Tesoi · · 6 min read

Decide first whether the code is wrong or the acceptance criteria were, because the two need opposite responses. A large share of agent work rejected at review turns out to be a specification defect, and sending such an item back unchanged runs the same input again and expects a different result.
The instinct is to look harder at the diff. The information that settles it is in the item.
Agent work rejected: what does it actually mean?
That returned work and written criteria disagree. It does not say which of them is wrong, and the reflex to assume the code is what makes the second attempt fail like the first.
When a person wrote the code, that assumption was usually safe. A developer who misread a ticket noticed partway through and asked, so the misreading never reached review. What arrived was the item plus a conversation nobody logged, and a disagreement was genuinely about the code most of the time.
A coding agent has no conversation to add. It has the item and the repository, so what comes back is the item's logical consequence. That makes a rejection a measurement of the item, and it is the only direct measurement of specification quality most teams ever get.
Bad code or bad criterion: how do you tell?
Re-read the item without looking at the diff, and ask whether a competent person given only that item would have produced the same thing.
The question is not whether the code is wrong. It is whether someone reading the item alone, with no access to the discussion it came from, would have built what the agent built. If they would, the item is the defect and the code is a symptom.
That test sorts rejections into three kinds, and they do not share a fix.
| What actually went wrong | The tell | What to change |
|---|---|---|
| The item was right; the work missed it | You can state what was wanted without opening the diff | Nothing upstream — send it back naming the criterion |
| The item was ambiguous; the agent picked a reading | Two people reading the item disagree about what it asked for | The item. Fix it, then re-run |
| The item was right, and the requirement was wrong | The work satisfies every criterion and is still not what you want | Neither. Accept it, close it, write a new item |
The third row is where the real damage happens, because it does not feel like new information. It feels like a failure, and the tempting move is to quietly rewrite the criteria to match what was built. That records the work as correct when nobody has established that it was, and it is the one response code review when an agent wrote the code rules out: a criterion edited after the fact can no longer fail.
W. Edwards Deming's third point for management is the general form of all of this — "cease dependence on inspection to achieve quality", and eliminate the need for it by building quality in first. Review is inspection. A team whose rejection rate stays flat is inspecting rather than improving, and the improving happens in the item.
How do you send work back so it fails differently?
Change something in the item before you re-run it. A returned item that is textually identical to the one that failed will fail the same way, and rework with coding agents is where the cost of skipping that step shows up.
A useful returned item carries four things:
- Which criterion failed, quoted in its own words rather than paraphrased. Paraphrase is where the second misreading enters.
- The observed result as a command and its output. "The import silently skips blank rows" is a report;
pnpm test:importwith its failing assertion is evidence. - What changed in the item, or an explicit statement that nothing did because the implementation simply missed a correct criterion.
- The decision, if one was made. If the rejection resolved an ambiguity, the answer belongs in the item where an agent will read it, not in the thread where it will not.
Rejections that skip the fourth recur. The ambiguity is resolved in someone's head, the item goes back with a sharper tone and the same words, and the next attempt lands in the same place. What an item needs to carry to be executable at all is set out in a backlog an agent can read, and the criteria side in acceptance criteria an agent can verify.
When to take it over yourself
When the difficulty is in expressing the requirement rather than in doing the work. That is the whole rule, and it is not about how hard the code looks.
Some requirements cost more to write down than to implement: a layout that has to feel right, a migration whose edge cases only appear against real data, a fix whose behaviour depends on a judgement nobody has made. Writing an executable item for these takes longer than doing them, and an AI agent failed task of this kind fails again after every rewrite, because the specification is the hard part.
Two bounces on the same criterion is a good threshold. One is ordinary; the second says the thing resists being written down, and a third attempt usually costs more than picking it up. The arithmetic is unkind: a clean item costs about an hour of human attention, one bounce takes it to roughly two, and a second past two and a half — at which point most of its cost is review rather than build.
Recording the reason so the pattern shows up
Record the cause against the item, not the count. A rejection tally says throughput dipped; a rejection reason says what to change on Monday.
Laimonade keeps this on the item: a decision taken during implementation is recorded against the backlog item with what was rejected and why, so the next person to open it sees the reasoning rather than only the outcome. The sprint workflow covers where the handover sits, and common failure modes are collected in troubleshooting. An agent can carry an item as far as In Review and no further, so a rejection is always a person's judgement and always has an author.
Three or four weeks of causes is enough to see the shape, and reading them is what the retrospective is for once sprint planning has become a readiness check. Teams often find one phrasing responsible for a third of their rejections — a template fix rather than a discipline problem, and the difference between handling agent errors and reducing them.
Frequently asked questions
Should a rejected item keep its original estimate?
Keep it, and record the bounce separately. Re-estimating after the fact makes the history describe what the work cost rather than what was predicted, destroying the only signal the estimate carried. The bounce count is the more useful number: it measures specification quality, which is the thing you can act on.
How many attempts before a human takes over?
Two on the same criterion. The first failure is ordinary and usually informative. A second against a criterion already rewritten once says the requirement resists being expressed, and a third attempt typically costs more than implementing it directly. Different criteria failing on successive attempts usually means the item is too large.
Does rework count against velocity?
It should not add points twice, and the bounce should be visible somewhere. Counting the item once on completion keeps velocity comparable across weeks, and recording that it took three passes keeps the cost visible — an unrecorded bounce shows up as a mysteriously slow sprint rather than a specification problem with a name.
Is a rejection a sign the agent is underperforming?
Rarely, and treating it that way sends you looking in the wrong place. An agent produces the item's logical consequence, so a rejection mostly reads the item. The useful comparison is not between agents but between weeks: whether the same phrasing keeps producing the same disagreement.