Engineering team structure when agents write the code
Stefan-Iulian Tesoi · · 8 min read

Mostly, the roles stay and the hours inside them move. Writing code shrinks, while specifying work before it starts and judging it after it returns grow to fill the gap. Tech leads review more, product managers write tighter items, QA designs checks instead of running them, and junior developers need a deliberate path to skills they no longer practise by typing.
That is the honest summary of engineering team structure a few months into running coding agents. The org chart rarely changes first. The calendar does, and the calendars that change most belong to the senior people nobody planned to load.
What changes in engineering team structure, and what does not?
Reporting lines and job titles change least. The ratio of thinking to typing changes most. A coding agent takes a well-specified item and returns a diff, so the stretch between "we agreed what to build" and "it is in review" compresses. Everything on either side of that stretch stays, and grows.
Melvin Conway's 1968 paper, How Do Committees Invent?, argued that organisations "are constrained to produce designs which are copies of the communication structures of these organizations." Agents do not repeal that. They add one channel to the communication structure, the item an agent is handed, and whatever that channel cannot carry, the system will not contain. A team whose items are vague ships a vague system, faster.
Four things do not change:
- Someone still decides what is worth building.
- Someone still owns the architecture and can say no to a design.
- Someone still answers for production when it breaks at night.
- Someone still has to grow the next senior engineer.
What does change is how many hours each of those people spends on which part of the loop. That is how AI changes software engineering roles in practice: through the calendar long before the org chart.
Where do the hours go now?
Into two queues on either side of the agent: specification before it starts, and review after it finishes. How many coding agents a team can run is set by those two queues rather than by the agents. As a working estimate, an item takes about 40 minutes to write so that an agent can execute it, and 10 to 20 minutes to review against criteria fixed in advance.
The arithmetic is unforgiving. A team whose agents finish six items a day needs four hours of specification and one to two hours of review every day, before anyone touches a design question. It used to make six of those verification decisions a week.
| Activity | Before agents | With agents |
|---|---|---|
| Writing code | The centre of an engineer's week | A fraction, mostly the hard parts |
| Specifying items | An afternoon per sprint | Daily, about 40 minutes per item |
| Reviewing returned work | Peer review of human diffs | Every diff, against criteria written first |
| Answering questions | In conversation, mid-build | Before the build, because agents do not ask |
| Sequencing | Implicit in who picks up what | Explicit, or two agents edit one module |
Felt speed is a poor guide to any of this. In METR's July 2025 study, 16 experienced open-source developers took 19% longer to complete issues when allowed to use AI tools, and still believed afterwards that AI had sped them up by 20%. The authors are explicit that the result may not generalise beyond experienced developers on codebases they know well. The lesson for structure is narrower: measure the queues, not the mood.
Role by role: what moves and what stays
The software team roles AI changes most are the ones closest to the item, just before and just after the build:
| Role | Less of | More of |
|---|---|---|
| Tech lead | Writing the hard module personally | Reviewing agent diffs, setting the constraints agents follow |
| Engineering manager | Assigning tickets to people | Watching the specification and review queues, protecting senior time |
| Product manager | Long documents read once | Short items with runnable criteria, read by a machine |
| QA engineer | Running manual test scripts | Designing checks, and acceptance testing on per-item builds |
| Staff engineer | Being the fastest implementer | Writing the conventions and boundaries agents work inside |
| Junior developer | Learning by typing routine code | Learning by reviewing, specifying and debugging agent output |
Tech lead. The tech lead becomes the team's main reviewer, because an agent's summary is a claim and the diff is the evidence, as AI code review sets out. The risk is that review becomes the whole job and the design work stops.
Engineering manager. The engineering manager's capacity question changes from "who is free?" to "which queue is full?". An idle agent usually means an empty specification queue, not a slow agent, and adding another agent makes that queue emptier per agent.
Product manager. The product manager still decides what is worth building. What changes is the item: an agent will not walk over and ask, so the distance between a decision and a buildable item has to be closed in writing. Whether that writing stays with the product manager or moves is the split described in AI product owner vs human product owner.
QA engineer. The QA engineer moves upstream. When an agent writes the code and its tests, someone has to decide what a check must prove, and someone has to use the feature the way a user will. Both are QA's work. Neither is running a script by hand.
Staff engineer. The staff engineer's leverage moves from code to constraints: the conventions file every agent reads, the module boundaries, the list of things no agent may touch. A rule written once applies to every run.
Junior developer. The junior developer loses the routine tickets that used to teach the codebase. That loss is real and does not fix itself. A team has to replace it on purpose, with review rotations, specification work and supervised debugging of agent output.
Why does the specification and review load land on senior people?
Because both jobs need judgement that only experience supplies, and nobody budgets for either. Writing an item an agent can execute means knowing which constraint will matter. Reviewing an agent's diff means spotting the plausible change that is wrong. Both fall to the people who already know the system, and both arrive as interruptions rather than planned work.
The failure is quiet. Seniors absorb the load between meetings, the queues still move, and for a few weeks the team looks faster. Then review becomes the bottleneck, specifications get thinner to keep the agents busy, and work comes back matching the title rather than the intent. Roughly half of the work we send back turns out to be a specification defect rather than an implementation one.
The signs of a product ownership gap show up here first: agents waiting rather than engineers, the same ticket explained in three places, and a board that disagrees with the repository.
The agents scale. The two or three people who can tell a good item from a bad one do not.
The structural fix is to name the work. Specification and review are jobs with owners and hours, not favours seniors do between meetings. Who owns the backlog splits it into deciding what should exist, which stays with a person, and keeping each item executable, which is continuous. Laimonade takes the second half: it grooms items into work an agent can execute, hands them over through MCP and checks what comes back against the criteria, so senior hours go on decisions rather than on rewriting tickets.
Which reorganisations are premature?
Most of the ones that change headcount or titles before anyone has measured the queues. Four are common:
- Cutting engineers because agents write code. The hours did not disappear; they moved to specification and review. Cut first, and the people who could have done that work are gone.
- Creating an "agent operator" role. Running agents is not a job on its own. Specifying and reviewing are, and they belong with the people who own the outcome.
- Stopping junior hiring. It saves a salary this year and empties the senior pipeline in five. The future of engineering teams still depends on people who learned the system from the inside.
- Moving QA out. The checking load rises with agent throughput. Removing the people who design checks is the wrong direction.
Measure before reorganising. Time ten items from prioritised to ready and compare that with build time. Count how many of the last twenty items needed a question answered before work could begin; above a third, specification is the constraint. Then watch where items wait for two sprints. The structure should follow what the queues show, not what the first week felt like.
Frequently asked questions
Do coding agents mean smaller engineering teams?
Not by default. Agents shrink the time spent writing code, but specification and review grow with every item they finish, and both need experienced people. A team may ship more with the same headcount. Cutting people before measuring where the hours moved usually removes exactly the judgement the agents depend on.
Who should own the backlog in the new structure?
Split it. Deciding what should exist stays with a person who holds the business context, usually the product manager or a founder. Keeping each item executable, with a named repository, runnable criteria and resolved dependencies, is continuous work that can sit with a product owner, a tech lead with protected time, or an AI product owner.
Should a team add a dedicated agent operator role?
Usually not. Starting and watching agents takes little time. The scarce work is specifying items and judging what comes back, and that belongs with the people who own the outcome. A separate operator becomes a relay between the people who decide and the agents that build, which adds a hop without adding judgement.