Coding agent rollout: a plan for one team

Stefan-Iulian Tesoi · · 7 min read

A surveyor sighting through a levelling instrument on its tripod with a notebook in hand, taking the measurement that has to exist before anyone starts building

One team, one repository, and a backlog you have measured before anything is installed. A coding agent rollout that begins with tooling stalls about a month in; one that begins by counting how many items are executable as written starts with the number that caps everything after it.

That number is the whole plan in miniature. Agents execute what is specified, so the share of your backlog that is specified is the share of your backlog they can touch — and on most first audits it is smaller than anyone expects.

What has to be true before a coding agent rollout?

Four things, and executive enthusiasm is not among them.

One team, not a department. A rollout across four teams produces four different sets of conclusions, none of which can be compared because nothing was held constant. One team gives you a result you can act on.

A repository whose tests run. An agent hands work back with the commands it ran and their exit codes, and that evidence is worthless if the suite was already red or nobody runs it. Fix that first; it is cheaper before agents than after.

Somebody who owns the backlog. Not a committee. The person who will answer "is this item ready" thirty times in the first fortnight, and who has the authority to say no.

A backlog you have measured, not one you believe in. This is the precondition people skip, and skipping it is what makes the first month look good and the second month look like a plateau.

Every rollout that disappointed had the same shape: a fast first fortnight spending a backlog written over the previous year, then a fall back to the rate at which somebody could write new items. The agents were never the variable.

Why is the backlog the first constraint, not the tooling?

Because installing an agent takes an afternoon and specifying work does not, so the thing you can do quickly is not the thing that decides the outcome.

The measurement is cheap. Take twenty items from the top of the backlog and ask, for each, whether somebody with no access to your team could execute it from the text alone. The method and its four questions are set out in the backlog audit to run before you point agents at it; the four defects that keep appearing are in a backlog an agent can read. What matters for planning is the count, because it is a ceiling and not a score.

If six of twenty pass, an agent has six items of work and your team has fourteen items of writing. That is a useful fact on day one and a demoralising discovery in week five, which is the only real argument for doing it first.

Introducing AI agents to a team in this order also changes what the first month is about. It stops being a tool evaluation, which nobody can fail, and becomes a measurement of how specifiable your work is, which is a thing you can improve.

A first month, week by week

Four weeks, with the agents arriving second. Treat this as the agent adoption plan rather than a pilot you run alongside normal work: the shape matters more than the exact days, but the order does not bend.

WeekWhat happensWhat you learn
1Audit twenty items. Specify ten properly. No agents.The ceiling, and how long an item takes to write
2One agent, one repository, those ten itemsWhether specified items survive contact
3Full loop: specify, dispatch, review, record causesWhere the time actually goes
4Read the causes. Decide what changes.Whether the constraint moved

A few notes on the weeks that people compress.

Week one has no agents on purpose. The temptation is to install something so the week feels productive. Resist it: the only output that matters is ten items written well enough to act on, and an honest count of how long each took. Fifteen to forty minutes is normal. If it is taking two hours, the item is probably two items.

Week two is the pilot proper, and deliberately one agent. Not because more would break, but because two agents make it impossible to tell whether a problem came from the work or from the coordination. How the ceiling on agent count actually works is in how many coding agents one team can run, and it is a week-five question rather than a week-two one.

Week four is the one that gets cancelled. It produces nothing shippable and it is the only week that changes the next month. Put it in the calendar before week one, with a named attendee list.

Laimonade exists for the specifying and checking half of that loop — drafting items with criteria and context already in them, then checking returned work against those criteria. A person still decides what is done. What setting it up involves is in getting started, and what the arrangement looks like from the perspective of whoever holds the team is in Laimonade for engineering leaders.

Which signals mean it is working?

Three, and none of them is throughput.

The signal that gets misread is agents sitting idle. It reads as an agent problem and is almost always an empty specification queue, which is a person problem with a person's solution. Teams that respond by adding another agent make the queue emptier per agent.

Industry survey data is worth reading alongside your own numbers rather than instead of them — the Stack Overflow developer survey consistently shows adoption running well ahead of trust, which is a useful check on whether your rollout is producing enthusiasm or evidence.

When should you stop and fix something else first?

When the constraint you are trying to move is not the one agents touch. Three cases come up repeatedly and all three are worth catching in week one.

There is also a case for not rolling out at all, and it is not defeatist. A team whose work is mostly novel — research, prototypes, design-led features where the requirement is discovered by building — gets less from this than a team with a steady supply of well-understood work. Rolling out AI coding tools into genuinely exploratory work tends to produce fast, confident answers to questions nobody had settled.

Why enthusiasm fades even when the tooling is fine is the subject of why agent adoption stalls after the first month, and it is worth reading before week one rather than after week five.

Frequently asked questions

How many people should the first rollout involve?

One team, and within it one person owning the backlog and two or three reviewing. Fewer than that and the result is one person's working style rather than a team's; more and you cannot tell which change produced which effect. A team of four to eight is the range where a month produces a readable answer.

Does every developer need their own agent?

No, and starting that way hides the constraint. One agent working a properly specified queue tells you more than five agents sharing a thin one, because with five you cannot tell whether the limit is specification, review or coordination. Add the second when the first is regularly idle for want of work.

What if the team is already using agents unofficially?

Treat it as data rather than as a policy problem. People adopting something without being asked to have found it useful, and they know where it breaks. The risk worth addressing is not the tool but the invisibility: work produced without an item, evidence nobody recorded, and no attribution. Bring it into the process rather than starting from zero.

How long before a rollout is worth judging?

Two sprints, not one. The first sprint spends whatever specified work already existed and flatters everything; the second runs on items written during the rollout, which is the rate you will actually live with. A judgement made at the end of week two is a judgement about your old backlog.