Engineering metrics that still work once agents build
Stefan-Iulian Tesoi · · 8 min read

The engineering metrics that still mean something are the flow metrics, read stage by stage, plus a cost and a rework rate. Cycle time, throughput and the DORA four survive, but the time inside them moves from building to waiting for decisions and review. Velocity and lines of code stay as observations, never targets.
The dashboards do not need replacing. They need reading differently, because the numbers on them now describe a different bottleneck than the one they were built to watch.
Which engineering metrics still mean something?
The ones that measure flow and stability, because neither depends on who did the typing. The ones that measure output volume or estimated effort lose their meaning first.
| Metric | Still meaningful? | What changes when a coding agent builds |
|---|---|---|
| Cycle time | Yes, split by stage | Building shrinks; waiting for specification and review dominates |
| Throughput | Yes, with a quality pair | Rises immediately, including when quality falls |
| Deployment frequency | Yes | Rises with smaller changes; says little about value on its own |
| Lead time for changes | Yes | Shortens; the remaining time is mostly review and pipeline |
| Change failure rate | Yes, and more important | The only DORA metric that can catch a faster producer making worse changes |
| Time to restore | Yes | Unchanged; incidents are still diagnosed by people |
| Velocity | As an observation | Measures how much was specified, not how fast it was built |
| Lines of code, PR count | No | Free to produce, so they measure activity and nothing else |
The four in the middle are the DORA metrics, which came out of more than a decade of research into software delivery metrics and were never about who wrote the code. They measure how often change reaches production, how long it takes and how often it breaks. All three questions still matter, and two of them, failure rate and recovery, matter more when the volume of change rises.
What does not survive is the habit of treating any of them as a total. A cycle time of three days is good news or bad news depending on which part of the three days it was.
Why does every total need splitting by stage?
Because a coding agent compresses one stage and leaves the others alone, so a total that improves can hide a stage that got worse. The total is an average over parts that are now moving in opposite directions.
Take an item's life in five stages: waiting to be specified, specified and waiting to start, building, waiting for review, and review to merge. Before agents, building was usually the longest. The arithmetic below is illustrative rather than a measured case, but its shape is what teams report:
| Stage | Before agents | With agents |
|---|---|---|
| Waiting to be specified | 1 day | 1.5 days |
| Ready, waiting to start | 1 day | 2 hours |
| Building | 3 days | 40 minutes |
| Waiting for review | 0.5 days | 1.5 days |
| Review to merge | 0.5 days | 0.5 days |
| Total cycle time | 6 days | about 3.6 days |
The total improved by 40%, and that is the number that reaches the slide. Underneath it, building fell from three days to under an hour while waiting for review tripled and specification got slower. A team reading only the total concludes the agents worked. A team reading the stages sees that two queues are now holding work for longer than the build ever did, and that the next gain is in neither the model nor the tooling.
Flow metrics were designed for exactly this reading: time in each state, work in progress per state, and where items age. The useful rule is to report cycle time as a stacked bar, never as one number, and to put a limit on how many items may sit in review at once. Which of those queues sets the real ceiling is worked through in how many coding agents one team can run.
Which numbers mislead first?
The ones that count output. Lines of code, merged pull requests and points delivered all rise on the first day an agent is switched on, whatever else is true, which is precisely why they get quoted.
- Lines of code cost nothing to produce and something to own. A rising count is a rising maintenance bill.
- Pull request count rises when changes get smaller, which is good practice, and when an agent splits work it should not have, which is not. The count cannot tell the two apart.
- Velocity keeps working as a record of what a team got through, and stops working as a forecast. It now tracks how much work was specified, a point made in detail in whether story points still work with coding agents.
- Throughput without a quality pair rewards more work of any kind, including work that comes back.
This is Goodhart's law with a faster engine attached: when a measure becomes a target, it ceases to be a good measure. The law always applied to engineering dashboards. What changed is the cost of gaming them. A person inflating a commit count had to spend their own time doing it. An agent inflates it as a side effect of doing what it was asked.
A metric that cannot fall while things get worse is not measuring the team. It is measuring that something was switched on.
The defence is pairing. Every volume number is reported next to a number that falls when the volume is bad: throughput next to rework rate, deployment frequency next to change failure rate, items accepted next to items reopened.
What does a delivery measurement stack look like now?
Four layers, each with one or two numbers, and every number able to get worse. That last property is the test for whether a number belongs on the stack at all.
- Flow. Cycle time split by stage, throughput per week, and work in progress in the review state. This layer finds the queue.
- Stability. Change failure rate and time to restore. This layer catches a faster producer making worse changes.
- Quality. Rework rate, the share of accepted items reopened or followed by a fix within fourteen days, and the share of returned work accepted without a specification change. This layer separates bad code from bad items.
- Cost. Spend per accepted item: model usage, compute and review hours, divided by items a person accepted. This layer is new, and it is the one finance will ask about.
The cost layer deserves its denominator. Dividing agent spend by items attempted makes a failing agent look cheap, because every abandoned run lowers the average. Dividing by items accepted charges the failures to the work that succeeded, which is what they actually cost. Laimonade records the model calls, tokens and cost of every coding run against the item it ran for, so the numerator exists per item rather than as a monthly invoice.
The agent-specific half of this, specification rate and acceptance without a fix, is covered in coding agent metrics that are worth tracking. The stack above sits around it: those numbers explain the supply, and these explain whether the supply turned into delivered change.
What belongs in a report to management?
Five numbers, as monthly trends rather than snapshots, each with one sentence on what moved it. More than five and the report becomes a dashboard someone has to interpret; fewer and it cannot show a trade-off.
| Engineering KPI | Why it is on the report |
|---|---|
| Cycle time, split by stage | Shows where work waits, which is where the next gain is |
| Accepted items per week | Delivered change, counted at the point a person accepted it |
| Rework rate | The counterweight that keeps the second line honest |
| Change failure rate | Whether faster change is breaking production |
| Cost per accepted item | Whether the spend is buying delivery or activity |
Engineering KPIs for a board are a different artefact from a team's working dashboard. The team needs the per-stage detail daily; management needs the direction of five numbers and the reason for each. Mixing the two produces either a report nobody reads or a dashboard nobody can act on.
Measuring engineering productivity this way also changes what an audit can check. When every accepted item carries its evidence, the number on the report can be traced back to the items behind it, which is the method in an engineering audit of what your agents actually shipped. In Laimonade, a coding agent hands work back as In Review and no tool lets it move its own item further; an item counts as delivered only once it is closed, by a person or by Laimon with the closure stamped as Laimon's. That rule is what keeps "accepted items" from being a count of claims.
Frequently asked questions
Should metrics be tracked per agent or per team?
Per team, with per-agent figures kept for diagnosis only. Agents run the work they are handed, so their individual numbers mostly reflect which items they were given. A per-agent leaderboard rewards the agent that drew the easiest items. Per-agent cost and failure rates are still worth having when one configuration or model performs worse, but they belong in an investigation, not on a report.
How long does a new baseline take after agents arrive?
About six to eight weeks. The first two or three weeks spend a stock of already-specified work and flatter every number. After that, throughput settles to the rate at which new work is specified and reviewed, and cycle time per stage stabilises. Comparing anything against the pre-agent baseline before then measures the stock running out rather than the new way of working.
Which single number would you keep if you could keep only one?
Cycle time split by stage, because it is the one number that shows where work is waiting. Every other metric describes an outcome; this one points at the queue to fix next. If it has to be a single figure rather than a stacked bar, keep the time items spend waiting for review, which is where most teams running agents find their constraint.