Coding agent metrics that are worth tracking
Stefan-Iulian Tesoi · · 5 min read

Three: how many items are specified well enough to start, how many come back accepted without a specification fix, and how long an item takes from ready to accepted. Lines of code and merged pull requests rise on day one whatever else is true, which is why those are the ones that get quoted.
The coding agent metrics worth keeping are the ones that can get worse. A number that only ever goes up is not measuring your team, it is measuring that you turned something on.
Which coding agent metrics are worth having?
Three, and each one fails in a different direction, which is what makes the set useful rather than the individual numbers.
| Metric | What it is | What a bad reading means |
|---|---|---|
| Specification rate | Backlog items per week written well enough to execute | The ceiling is dropping; everything downstream will follow |
| Accepted without a specification fix | Share of returned work where the item was right | The writing has degraded, not the model |
| Ready to accepted | Elapsed time per item, including review | Review is absorbing the gain |
The first is the supply. The second is the quality of that supply, measured by something other than the person who produced it. The third is where the benefit either materialises or quietly gets eaten.
A team can be healthy on any two and in trouble on the third, and the combination tells you which. High specification rate with falling acceptance means somebody is writing quickly and badly. Good acceptance with a long ready-to-accepted time means the items are fine and review is the constraint.
Why does throughput flatter a team?
Because agent throughput is downstream of the specification queue, so it reports on the queue rather than on the agents, and it reports late.
This is worth separating from the first-fortnight effect, which is a different mechanism with the same symptom. A pilot's opening surge comes from spending a stock of pre-written work, and that argument is set out in what the first two weeks of an agent pilot prove. The steady-state version is subtler: once the stock is gone, throughput tracks the specification rate almost exactly, which means it tells you nothing the specification rate did not tell you earlier and more clearly.
It also rises when quality falls. An agent producing more, worse work raises every count-based number until the rework arrives weeks later, by which point the graph has already been shown to somebody.
A metric that cannot fall while things get worse is not a metric. It is a reassurance with a number attached.
Measuring AI developer productivity by output volume reproduces the oldest mistake in the discipline, with a faster producer attached to it. Counting merged pull requests was a poor proxy for value when people wrote them; it is a worse one now that the cost of producing another has dropped.
Reading rework as a specification signal
The ratio between rework caused by the item and rework caused by the code is the most useful number in the set, and almost nobody collects it, because it requires one sentence of judgement per returned item rather than a query.
Record, per rejected item, which of the two it was: the acceptance criteria were ambiguous or wrong, or they were right and the implementation missed them. That is the whole instrument. What it then tells you over a quarter:
- Item-caused share rising. Specification quality is falling. Usually one cause: the person who was writing items well got busy, and the writing quietly moved to whoever had time.
- Code-caused share rising. Something changed that the items do not describe — a refactor in flight, a new area of the codebase, a model or tool change.
- Both flat and low. The constraint has moved somewhere else, and the next thing to measure is review time.
- Both rising. Usually a scope problem: items have got larger, and a large item fails in both ways at once.
Why estimates stop measuring effort and start measuring this kind of risk is covered in do story points still work with coding agents. Laimonade records the criteria an item was accepted against alongside the work, which is what makes the distinction cheap to draw rather than a matter of recollection; how that fits together is in how Laimonade works.
What should a monthly review ask?
Four questions, and each has a threshold that turns it from a number into a decision. An hour a month, on the same day.
- Did the specification rate hold? If it fell two months running, the writing has lost its owner. Name a new one rather than exhorting the team.
- Is the item-caused rework share above a third? That is the point where the template, not the people, is usually the problem. Change the template and watch the next month.
- Has ready-to-accepted time grown while acceptance held? Review has become the constraint. Adding agents here makes it worse, and the arithmetic is unforgiving.
- Does the board agree with the repository? The answer drifts silently, and the method for checking it is in an engineering audit of what your agents actually shipped.
The question behind all four is whether anything would change as a result of the answer. Is our AI tool working is not answerable as asked; "has our specification rate fallen for two months" is, and it comes with an action attached.
What a first month should watch, before any of this has enough history to trend, is in coding agent rollout.
Frequently asked questions
Do DORA metrics still apply?
Yes, and they measure the delivery pipeline rather than the specification one. Deployment frequency, lead time, change failure rate and time to restore are defined in DORA's own guide and remain sound. What they do not see is whether the work was worth doing or whether the item was right, which is precisely the half that agents move. Run both sets.
Should individual agent usage be tracked?
No. A per-person usage number turns the tool into an assessment instrument, and people manage assessment instruments rather than use them. It also measures the wrong unit: the useful figures are per item and per team, because the constraint is the supply of specified work rather than anyone's diligence.
How long before the numbers mean anything?
Two months for a level, a quarter for a trend. A single sprint is noise — one large item or one absence moves every figure. The rework ratio stabilises fastest because it is a proportion rather than a count, which is another reason to start collecting it on day one even though it is the one requiring manual judgement.
What if the numbers look good and the team is unhappy?
Believe both. Good figures with unhappy reviewers usually means the gain landed as extra review load rather than as time back, which throughput will never show you. Ask what people are doing more of than they were three months ago; the answer is often reading, and that is a real cost that no output metric records.