Spec-driven development: does it replace the backlog?

Stefan-Iulian Tesoi · · 6 min read

An archival engraving of one clockwork instrument with an electromagnet, every gear drawn in detail on a blank page: a single change specified in depth, with nothing to say what comes before or after it

No. Spec-driven development replaces the inside of one backlog item, not the backlog. A spec in GitHub Spec Kit or Kiro describes a single change in depth: requirements, design, tasks. It does not say which change comes next, what blocks it, or whether last month's spec was met. Ordering and checking remain a backlog's job.

The confusion is understandable. Both produce a written statement of what an agent should build, both carry acceptance criteria, and both are now read by tools like Claude Code rather than only by people. They answer different questions at different scales, and a team that adopts specs and retires its backlog loses the second set of answers without noticing, usually until two agents build conflicting changes in the same week.

What is spec-driven development?

Spec-driven development is a way of working in which a written specification, not the prompt and not the code, is the primary artefact an agent builds from. A person describes the change, the tooling expands it into requirements, a design and a task list, a coding agent implements the tasks, and the spec stays in the repository.

Two tools set the current shape of the practice:

EARS is older than either tool. It was developed at Rolls-Royce for aero-engine control requirements and published in 2009, and its value to a coding agent is the same as its value to an engine controller: every requirement is one testable sentence with a named trigger.

Where does a spec overlap with a backlog item?

In the acceptance criteria, and almost nowhere else. A well-written backlog item and a well-written spec both state what the change must do in terms a test can check. "WHEN an export is requested with no matching rows THE SYSTEM SHALL return a CSV containing only the header line" is a good criterion in either place.

Past that sentence, the two documents answer different questions:

QuestionSpec (Spec Kit, Kiro)Backlog item
What exactly must this change do?Yes, in depthYes, in three to six criteria
How should it be built?Yes: design, data model, contractsRarely
Which change comes next?NoYes: its position in the backlog
What does it depend on or block?Only inside the featureAcross features and repositories
Who decided it matters?NoYes: priority, owner, epic
Was it met, and on what evidence?Tasks ticked by the agentReview against criteria, recorded

A spec is a deep description of one change. A backlog an agent can read is a shallow description of many changes, ordered, with the relationships between them recorded. Calling both of them specs for AI coding is accurate. Treating them as interchangeable is the mistake.

What does a spec leave out?

Sequence, dependencies between features, and the verdict. Each is a property of the whole set of changes, and a spec only ever sees one change.

Why do specs go stale after the change ships?

Because nothing forces an update, and the next change usually edits the code rather than the spec. The stated ideal of spec-driven development is that maintenance happens by evolving the spec and regenerating. In practice, a bug fix three weeks later is a two-line patch made directly in Claude Code, and spec.md now describes behaviour the system no longer has.

A stale spec does more damage than a missing one, because agents read it. An agent given a new feature in the same area loads the old spec as context and treats it as current. The result is a change built against superseded requirements, which passes its own tests and quietly reverts the patched behaviour.

A spec is accurate on the day it is implemented. After that it is a historical document that looks like a current one.

A backlog has the opposite property. An item is closed once, with a record of what shipped and on what evidence, and nobody expects a closed item to describe today's system. Its history stays honest because it is labelled as history.

How should specs and a backlog work together?

Let the backlog own the list and let the spec own the inside of one item. Four rules make the split hold:

  1. The backlog item comes first. It carries the title, priority, dependencies and three to six acceptance criteria, written the way an item an agent can execute needs them.
  2. A spec expands the item only when the change is large. A change that touches more than one module or needs a design decision earns a spec, generated from the item and linked to it. A one-file change does not.
  3. Verification reads the item's criteria, not the spec's ticked tasks. The criteria were written before the build, by someone other than the agent, which is what makes them evidence.
  4. A spec is marked superseded when a later item changes the same behaviour. Otherwise the next agent reads it as current.

This is the split Laimonade is built around. The backlog item and its acceptance criteria live in Laimonade, and an agent fetches its work order over MCP from Claude Code, Cursor or Codex, as the connection guide describes. Any spec the agent writes stays in the repository beside the code. The agent hands the result back for review against the item's criteria; it cannot close the item itself.

Frequently asked questions

Is a spec just a long user story?

No. A user story states who wants something and why, in one sentence, and leaves the how open on purpose. A spec states what the system must do as testable requirements, how it will be built, and the tasks that build it. A spec usually starts from a story and adds the design and task plan the story deliberately omits.

Should specs live in the repository or the tracker?

In the repository, beside the code they describe, because agents read them as context and they version with the change. The tracker holds what the repository cannot: order, priority, dependencies across repositories, and the record of whether the work was accepted. Copying a spec into the tracker creates two versions that drift apart.

Who reviews the spec before the agent starts?

Someone other than the agent that generated it, ideally the person who wrote the backlog item. Spec Kit and Kiro both separate requirements from design and design from tasks, and the gap between phases is the cheapest review in the process. A wrong requirement caught in requirements.md costs a minute; caught in a pull request, it costs the build.