When to stop an AI rollout, and what to keep
Stefan-Iulian Tesoi · · 6 min read

When the constraint you set out to move has not moved after a fair trial, and you can name which constraint it was. A rollout that never stated one cannot be stopped honestly, because nothing that happens would count as failure, so it drifts instead of ending.
The decision to stop an AI rollout is mostly made before the rollout starts. Whoever writes the pilot document either writes down the condition that would end it, or leaves the team to argue about it later, when every hour already spent is an argument for carrying on.
What would count as failure?
A number attached to the constraint the rollout was meant to move, a threshold, and a date to read it on, all written before the first coding agent touches the repository. Missing any one, there is nothing to fail against, and a rollout nobody can fail is one nobody can finish.
"Adopt coding agents" is not a constraint. "Cut the time from ready to accepted below two days" is, and so is "clear the forty-item bug backlog without a new hire". Each implies a reading and a line. The date should be at least two sprints out, because the first sprint spends whatever specified work already existed and flatters everything.
Most failed AI adoption is not dramatic. It is a pilot that produced some output, some enthusiasm and some rework, with no agreed line between success and its absence, so it continued because stopping needed an argument and continuing did not.
What does a fair trial look like?
Two sprints, one team, one repository, a measured backlog, and review time that somebody actually budgeted. A trial missing those did not test the agents; it tested the gap they would have filled.
| Condition | Why it matters | If it was missing |
|---|---|---|
| Two sprints minimum | Sprint one spends a stock of old items | You judged the old backlog |
| One team, one repository | Results can be compared | Four conclusions, none actionable |
| A measured backlog | The ceiling is known in advance | Idle agents read as a tool failure |
| Review time budgeted | Returned work gets read promptly | The review queue absorbed the gain |
| A named backlog owner | New items get written | Writing went to whoever was free |
If two of these were missing, the honest conclusion is that the trial was not run, which is a different decision from the trial having failed. How the preconditions are set up in the first place is covered in the coding agent rollout plan.
Three endings, and only one is a failure
A pilot ends in one of three places, and they call for different decisions:
- The constraint moved. Keep going and name the next constraint, which is usually review capacity or the rate at which items get specified.
- The constraint did not move, and the trial was fair. That is the failure, and it is a legitimate result. Stop.
- The constraint did not move, and the trial was not fair. An unfinished experiment. Fix the missing condition and rerun against a date, or stop and record why the trial could not be run.
The third is the common one. Teams reach for rolling back AI tools when the missing piece was a backlog owner, and the tool gets blamed for an empty queue. Why that pattern appears around week five is the subject of why agent adoption stalls after the first month.
A fourth ending deserves a name: the constraint moved, but moving it cost more than it was worth. Review hours rose faster than accepted items, or one senior engineer spent every afternoon writing specifications. The licence is rarely the real cost; the hours of writing and checking are, and what product ownership costs a small engineering team puts numbers on them.
How do you stop an AI rollout without losing the work?
Keep the items, the causes and the evidence; switch off the dispatch. What a rollout produces that outlasts it is rarely the code. It is a backlog written precisely enough to execute, and a record of what went wrong.
Four things are worth keeping deliberately:
- The rewritten items. An item specified for an agent, with named checks, explicit scope and open questions marked, is a better item for a person too.
- The rework causes. One sentence per rejected item saying whether the item or the code caused it. This is the most reusable finding a pilot produces.
- The review timings. They say what a restart would cost the reviewers, before anyone asks them.
- The pilot document, with its constraint and threshold. It is what makes a later restart credible.
Then close the loose ends on the day of the decision, not whenever someone remembers: revoke the agent's credentials and repository access, review or return anything sitting in review, and close the agent branches nobody will merge.
If Laimonade was part of the pilot, the plans are on the pricing page. You can cancel at any time, and cancellation takes effect at the end of the current billing period. You keep ownership of the content you submit — the terms of service say so explicitly. When AI tools do not work for a team, the backlog writing they forced usually still did.
Why is sunk cost louder here than elsewhere?
Because the costs of a coding agent rollout are paid in visible effort, and the benefits are paid in a queue metric few people watch. Whoever argued for the rollout has spent credibility as well as budget, and stopping reads as losing it.
The effect itself is old and well measured. Arkes and Blumer's study of sunk cost in 1985 included a field experiment with theatre season tickets: people who had paid full price attended more plays than people who had been given a discount. Having paid, they went.
Agent rollouts add three amplifiers of their own:
- Security approval took weeks, and nobody wants to go through it twice.
- Output is visible. Pull requests exist, and a merged one looks like success even when the rework arrives a month later.
- The rollout was announced. A pilot mentioned in an all-hands has an audience for its ending.
The time to decide what would end a rollout is before anyone has spent anything on it. Afterwards, every argument for stopping has to outweigh everything already spent.
Frequently asked questions
Is pausing the same as stopping?
No. A pause has a written date and a condition for resuming; a stop has neither. A pause without them is a stop that nobody announced, which is the worst version: credentials stay live, items go stale in Ready, and the team cannot tell whether to keep writing for agents or not.
What do you tell the team?
The constraint, the reading and the decision, in that order. "We set out to cut ready-to-accepted time below two days; after two sprints it was four; we are stopping." A team can agree or disagree with a number. Told "it was not the right fit", people fill the gap with their own theory.
Can you restart later without losing credibility?
Yes, if the restart names what changed. A rollout stopped on a stated condition earns credibility, because it showed the team the pilot was allowed to fail. Restart with a different constraint or a fixed precondition, such as a backlog owner or budgeted review time, and put the new pilot document beside the old one.