Teaching a team to write executable backlog items
Stefan-Iulian Tesoi · · 6 min read

Not with a style guide. Have each person hand one of their own items to a coding agent unedited, then read what comes back together. The returned work is a measurement of the item, and nobody argues with a result produced from their own words.
Writing executable backlog items is a skill that transfers badly by instruction and well by consequence. A team can agree with every rule in a document and keep writing the same items, because nothing in an ordinary week shows an author that their item was the problem. An agent shows them within the hour.
Why does a template not produce executable backlog items?
A template tells people which boxes exist, not what goes in them, and a filled-in box reads as finished whether or not an agent could act on it. Introduce one and within a fortnight every item has an acceptance criteria heading; underneath, a good share say "works as expected".
The deeper problem is that the defects that stop a coding agent are not formal. An item can carry a title, a description, three bulleted criteria and a link to the design, and still say "handle the error case" without saying which error or what handling means. The author knows. The agent does not, and neither does whoever does the review.
Style guides fail the same way, at extra cost. They are read once, during the rollout, mostly by the people who already write well. The authors whose items cause rework skim the guide, agree with it, and carry on.
A style guide describes a good item. An agent's output shows a person their own item, which is the only version of the lesson that sticks.
What does an exercise that teaches it in an hour look like?
One hour, one item per person, and no editing before the run. That is the whole design, and the no-editing rule is the part that makes it work.
- Each person picks one item they wrote in the last month that is still open. Their own, not a colleague's: defending someone else's writing teaches nothing.
- Dispatch it to a coding agent exactly as written. Claude Code, Cursor or whatever the team is piloting. The tool matters less than the rule.
- Read the result against what the author meant, out loud, author first. Where the two differ, find the sentence in the item that permitted the difference.
- Rewrite the item and run it again if there is time. The second run is usually where the lesson lands.
Six people running small items in parallel fits comfortably in an hour. What they find is predictable enough to tabulate:
| What came back | The sentence that caused it | The rewrite |
|---|---|---|
| The right change in the wrong place | "Update the settings page", with two settings pages | Name the route or file |
| A feature nobody can verify | "Should be fast" | A number, and the command that measures it |
| A confident guess at an open question | "Handle the error case" | Which error, and what the user sees |
| Twice the intended scope | "Tidy the export module while you are there" | A separate item |
The exercise works because it separates the author from the argument. In a meeting, "this criterion is ambiguous" is an opinion. In front of the output, the ambiguity has already been resolved one way, visibly, and the author can point to the word that allowed it. What makes a criterion checkable is set out in acceptance criteria that an agent can verify; the exercise is how people come to care.
This is the part of training a team on AI workflows that most rollouts replace with a tool demo. A demo teaches people what the agent can do. The exercise teaches them what their writing does, which decides whether the agent is any use.
What should backlog writing standards cover?
The parts a coding agent cannot infer, and nothing else. Backlog writing standards that also regulate tone, length or structure recreate the template problem: compliance with the form, no change in the content.
Four rules are worth writing down:
- Every criterion names how it is checked: a command, a test, or a URL and what appears there. "Works correctly" fails on sight.
- Scope says what is out as well as what is in. Agents extend; a boundary stops them.
- Open questions are marked as open, not left for the agent to settle. It will choose, and in review the choice will look deliberate.
- One outcome per item. Criteria describing two things that could ship separately mean two items.
Everything else can vary by author without harm: phrasing, whether there is a user-story sentence, how long the description runs. Worked examples of the full shape are in how to write a backlog item an agent can execute.
Laimonade drafts items with those parts already present, including criteria that name a check and the repository context, which turns the team's job from writing on a blank page into correcting a draft. That is a smaller skill, but still a skill: someone who cannot see a vague criterion in their own writing will not see it in a draft.
How do you keep the standard after the enthusiasm fades?
By making returned work the standing feedback rather than relying on the training session. The hour teaches the lesson once. What keeps it is one sentence per rejected item recording whether the item or the code caused the miss.
That record does two jobs. It tells each author, item by item, when their writing was the cause. Over a month it shows which of the four rules keeps breaking, which is the only evidence worth changing the standard on. Writing tickets for agents degrades quietly, usually when the person who wrote well gets busy and the writing drifts to whoever has time, and the item-caused share of rework is where that shows first.
Put the check at the ready gate, not in a retrospective. An item that breaks a rule does not reach Ready, and whoever holds the backlog returns it naming the rule. How the columns and hand-offs fit together is described in the sprint workflow. The Scrum Guide treats refinement as an ongoing activity rather than an event, and agents expose its absence faster than people did.
Repeat the hour whenever someone joins, and once a quarter with the items that caused the most rework. It costs far less than a second rollout, which is what a team whose writing has drifted eventually needs. Where the exercise sits in a first month is part of the coding agent rollout plan.
Frequently asked questions
Should a tool enforce the format?
A tool should enforce the checkable parts and nothing more. Refusing an item with no criteria, or a criterion that names no check, is mechanical and worth automating. Judging whether a criterion is precise enough is not, and a tool that pretends it can teaches people to satisfy the tool rather than the agent.
Who writes the items once agents are running?
Whoever owns the outcome writes or approves each item, and one person owns the backlog as a whole. Spreading authorship across the team is fine; spreading ownership is how items end up written by whoever is least busy. A drafting tool changes who produces the first version, not who answers for it being right.
How long does it take a team to get good at this?
About two sprints for the four rules to become habit, and longer for judgement. The first hour removes the worst defects. A month of rejected items, each recorded with its cause, does the rest. Teams that stop recording causes after the first fortnight tend to drift back within a quarter.