Agent pilot: what the first two weeks prove
Stefan-Iulian Tesoi · · 6 min read

Two numbers: how many items an agent executed without a person editing them first, and how much of the rework was caused by the item rather than the code. Both are countable in a fortnight. Velocity and satisfaction are not, and at two weeks they measure mostly the circumstances of the fortnight.
An agent pilot is usually described as an evaluation of a tool. It is more useful, and more uncomfortable, to treat it as a measurement of your own backlog — because that is what it measures whether or not you intended it to.
What is an agent pilot actually for?
Finding out how much of your work is specified well enough to act on. The tool question is secondary, and it is mostly answered by the same evidence.
This matters because the two framings produce different designs. A tool evaluation tempts you to prepare: pick clean items, write them carefully, make the demonstration fair. A measurement of the backlog requires the opposite — you must use items exactly as they were, because the thing being measured is what your backlog is like on an ordinary Tuesday.
The mechanics of a two-week trial are set out in what to look for in an agent-native project tracker, and where a pilot sits inside a longer rollout is in coding agent rollout. What follows is the part those leave out: what the fortnight can and cannot establish, and how to avoid concluding something it did not show.
A pilot that rewrote its own items before running them has answered a question nobody asked: whether agents can execute work written specifically for agents. They can. That was never in doubt.
What does a pilot prove that a demo cannot?
That the tool works on your items, which is the only version of the question that has an unknown in it.
A demo runs an example written by somebody who knew exactly what the tool needed. Every ambiguity was resolved before you saw it, which is what makes a demo smooth and also what makes it uninformative. The item was the product of the vendor understanding their own system, and your items are not.
A proof of concept coding agent exercise is therefore only meaningful on unmodified work. Take ten items you already consider ready — not the ten best, the next ten — and hand them over as they are. The gap between "we consider these ready" and "an agent could act on these" is the finding, and on a first pass it is usually wider than the team expects. How to measure that gap deliberately is in the backlog audit to run before you point agents at it.
Which early numbers mislead?
Four, and they mislead in the same direction, which is why a fortnight so often ends in unjustified confidence.
| Number | Why it flatters | What to use instead |
|---|---|---|
| Items completed in week one | Spends a backlog written over the previous year | Items executed with no human edit |
| Self-reported time saved | Measures relief, not hours | Time from ready to accepted |
| Pull requests merged | Rises whatever the quality | Share of rework caused by the item |
| Satisfaction after a fortnight | Nobody dislikes a new tool in week one | The same question at week eight |
The first row is the one that does real damage. Most teams have months of well-understood work that was never urgent enough to schedule, and an AI coding trial consumes it quickly and impressively. The rate you see is the rate at which that stock is spent, not the rate you will live with, and the fall afterwards gets attributed to the tool rather than to the arithmetic.
A two week pilot is short enough that it is tempting to explain the whole pattern as a novelty effect, and worth resisting. The canonical evidence for people behaving differently because they are being watched is thinner than its reputation: the reanalysis of the original Hawthorne illumination experiments found the dramatic version of the story largely unsupported by the underlying data. "Novelty" is a label, not a mechanism. The mechanism here has a name and you can check it — a stock of pre-specified work, whose size you can count before you start.
How do you read the result honestly?
Sort it into one of three outcomes, and notice that only one of them is about the tool.
- The items were executable and the work came back good. The tool question is answered and the constraint is somewhere else. Scale carefully and watch review capacity, because that is where the cost moves next.
- The items were not executable. The agent produced confident work against acceptance criteria nobody had written. This is the most common outcome and the most useful: you have learned, in a fortnight and cheaply, that the backlog is the constraint. It is not a failed pilot.
- You cannot tell. The items were rewritten for the pilot, or a person edited them mid-flight, or nobody recorded which fixes were specification and which were code. This is the only genuine failure, and it is a design failure rather than a result.
The third is avoidable with one discipline: record, per returned item, whether the fix was to the item or to the diff. It takes a sentence per item and it is the difference between a number and an anecdote. Where that record belongs in the week is covered in the sprint workflow.
Laimonade is built around that loop — drafting items with criteria and context in them, then checking returned work against those criteria, with a person deciding what is done. What it costs against the hours a pilot measures is on the pricing page, and the honest comparison is against the time your team currently spends clarifying work rather than against the tool it replaces.
Frequently asked questions
Should a pilot use real work or a sandbox?
Real work, from the actual backlog, unmodified. A sandbox removes the only variable worth measuring — whether your items are executable as written — and replaces it with a test of the tool under ideal conditions, which the vendor has already demonstrated. Pick work that is real but not urgent, so a bad result costs a week rather than a release.
Who should run it?
The person who will own the backlog afterwards, not the person most excited about the tool. Enthusiasm is useful for getting a pilot started and unhelpful for reading it, and the owner is the one who has to live with the specification work the result implies. One reviewer alongside them is enough.
What counts as a failed pilot?
Only the outcome where you cannot tell what happened. A pilot that shows your items are not executable has produced a real and actionable finding for a fortnight's cost; a pilot whose items were rewritten mid-flight has produced nothing, however good the numbers look. Judge it by whether you can state what you learned, not by whether the answer was encouraging.
Is two weeks long enough?
For the two numbers, yes; for anything about sustained throughput, no. A fortnight reaches the second sprint, which is where the first one's shortcuts surface, and that is what makes it the shortest honest window. Questions about capacity, review load and what the team's steady rate looks like need the month.