A sprint retrospective with coding agents in the loop
Stefan-Iulian Tesoi · · 5 min read

Where the specification failed, not how fast the code arrived. Sort every reworked item by cause (the item, the code or the review) and spend the hour on the largest pile. Agent speed is not a retrospective topic, because nobody in the room can change it by trying harder.
A sprint retrospective with coding agents keeps its slot in the calendar and changes its subject. The team still asks what to do differently next week; the difference is that most of the answers now live in how work was specified and reviewed, and those are things the people in the room control completely.
What do the old retrospective questions no longer find?
The classic agile retrospective questions assume the team's own execution was the constraint: we underestimated, we got blocked, we switched context too often. With agents doing most of the building, those answers describe a small share of the week, and the hour gets spent on them anyway because they are familiar.
Three staples stop producing anything useful:
- "Were we too slow?" Execution time is mostly the agent's. Nobody can resolve to make it faster.
- "Did we estimate well?" Points stopped measuring effort once implementation got cheap; what they misjudge now is how likely an item is to come back needing a decision.
- "Who was blocked?" Agents are blocked by missing decisions, which is a specification finding under another name.
What replaces them is one question with a countable answer: which items came back not meeting their criteria, and why. The pillar on sprint planning with AI agents makes the same argument for planning; the retrospective is where it gets checked against what happened.
Why bring the rework ledger instead of impressions?
Because a retrospective without data rediscovers the loudest complaint of the week, and the loudest complaint is rarely the most frequent cause. Rework causes have to be recorded during the sprint, one line per rejected item, and only read out at the end.
Each line needs the item and one of three causes:
- Item. The acceptance criteria were ambiguous, missing or wrong.
- Code. The criterion was right and the implementation missed it.
- Review. The work was fine, but review took days or passed something it should not have.
The habit of recording that cause as the item is sent back is set out in what to do when an agent returns work that does not meet criteria. It costs about a minute per rejection, and it is the difference between a retrospective with evidence and one with a mood. Laimonade records the criteria each item was accepted or returned against alongside the work, so the first column of the ledger is already filled in rather than reconstructed from memory.
An illustrative week for a team of five reads like this:
| Cause | Count | Typical example |
|---|---|---|
| Item: no command in the criteria | 4 | "The export should be fast" |
| Item: scope unstated | 2 | The agent tidied a neighbouring module |
| Code: criterion right, missed | 1 | An off-by-one in a date range |
| Review: waited more than two days | 3 | Nobody owned the review queue on Thursday |
Six of ten rework causes sit in the writing. That is the pile, and it is a better use of the hour than the one code miss that everybody remembers because it was annoying.
Bring the monthly numbers too, when the month turns: specification rate, share accepted without a specification fix, and time from ready to accepted, as described in measuring whether coding agents are working.
An agenda for a sprint retrospective with coding agents
One hour, five parts, in this order:
- Ten minutes: read the ledger aloud. Counts only, no discussion yet.
- Twenty minutes: the largest pile. Take the biggest cause group and read two of its items in full: the original text, what came back, and the sentence that allowed the difference.
- Fifteen minutes: one rule. Write a single change to the team's item-writing rules or its ready gate that would have prevented the pile. One, not five.
- Ten minutes: review lag. If review was a pile, decide who reviews what and when. That is a staffing decision, not a writing one.
- Five minutes: last week's rule. Did it hold, and did the pile it targeted shrink?
A retro format for AI teams that runs past an hour tends to drift back into the old questions, because the data is exhausted and the familiar topics are still there.
A retrospective that ends with five resolutions has made none. One rule, checked next week against the pile it was meant to shrink, is a change.
How do you turn a finding into a rule that sticks?
Put it where the next item is written or accepted, not in the notes. A rule in a document is read once. A rule at the ready gate is applied every time an item moves to Ready: "every criterion names a command" becomes something whoever holds the backlog checks, and an item that fails it goes back with the rule named. How the ready gate sits among the columns is described in the sprint workflow.
Then measure it. The pile the rule targeted should be smaller next week. If it is not, either the rule was wrong or nobody applied it, and both are findings worth the five minutes. The Scrum Guide says the most impactful improvements may even be added to the next sprint's backlog; a rule with an owner and a pile to watch is what that looks like here.
Frequently asked questions
Should the agents' output be discussed at all?
Only as evidence of the item that produced it. Discussing an agent's code quality in the retrospective turns into a debate about a tool nobody in the room controls. The useful form is "this output came from this sentence", which points at something the team can change next week, rather than at a model it cannot.
Who runs the retrospective now?
Whoever the team would pick anyway, usually the scrum master. What changes is who brings the data: the person holding the backlog owns the rework ledger and reads it out, because most findings land in their rules. A facilitator who is not the backlog owner keeps the hour from turning into a defence of the items.
How often should it happen with one-week sprints?
Weekly, and shorter than before. Thirty minutes is enough when the ledger is kept during the week, because the data gathering has already happened. Teams that move to fortnightly retrospectives on one-week sprints tend to lose the link between a rule and whether it worked.