How to write a bug report for a coding agent

Stefan-Iulian Tesoi · · 6 min read

A pocket watch movement clamped in a steel movement holder beside a screwdriver: the fault held still on the bench so it can be examined, not described

A check it can run that fails today and should pass after the fix. A bug report for a coding agent needs the steps to reproduce, the expected and actual result, and the last version where it worked, and its acceptance criterion should be the reproduction itself, run again and passing.

Everything else on a bug report is context. An agent given context without a reproduction fixes the first plausible cause it finds and reports success, because nothing on the item can tell it otherwise.

Why is a symptom not enough for a coding agent?

Because a symptom has many causes, and the agent will pick one. "The export button does nothing on large projects" is a fine thing for a user to say and an incomplete thing to hand to an agent. A person triaging it would ask the reporter a question. An agent works from what is on the item.

Given only the symptom, it reads the export code, finds something that could plausibly fail on large inputs (a timeout, a missing page size, an unawaited promise), fixes that, and returns a green suite. The fix may be real. Whether it is the bug is a separate question, and nothing on the item answers it.

The failure has a recognisable shape: the item comes back done, the reporter tries again a week later, and it still fails. The work was not wrong. It answered a different bug.

An agent given a symptom fixes a cause. An agent given a reproduction fixes the bug, and can prove it.

What does a bug report for a coding agent need?

Five fields, and each one replaces a question a person would otherwise have asked the reporter:

FieldWhat it holdsThe question it replaces
Steps to reproduceCommands or clicks, in order, from a stated starting stateHow do I make it happen?
Expected vs actual behaviourOne line each, observable, no interpretationWhat is wrong with what it does?
Last known goodA version, commit or date where it workedDid this ever work?
ScopeThe repository, and the suspected area marked as a guessWhere do I start?
Acceptance criterionThe reproduction, run again, producing the expected resultHow do I know it is fixed?

The steps to reproduce matter most and are written worst. "Open a big project and export" is not a step. "Seed a project with 2,000 items, call POST /export, wait 30 seconds" is. The starting state is the part people leave out: which account, which data, which flag.

Expected vs actual behaviour should be written as observations. "Expected: a CSV with 2,000 rows. Actual: the request returns 504 after 30 seconds." Not "the export is broken", and not "the query is too slow", which is a diagnosis dressed as an observation.

The last known good version turns an open-ended hunt into a bounded one. With one good commit and one bad, git bisect finds the change that introduced the bug in about log2(n) steps: ten steps across a thousand commits. An agent can run that loop mechanically when the reproduction is a command. Without the good version it has to reason from the code, which is where plausible but wrong fixes come from.

Scope should be labelled as a guess when it is one. Naming the repository works as it does for any backlog item an agent can execute; the point specific to bugs is that a suspected file stated as fact anchors the agent on it.

How does the reproduction become the acceptance criterion?

Write the criterion as the reproduction plus the expected result, so it fails today and passes once the bug is fixed. For the export bug: "POST /export on the 2,000-item fixture returns 200 with 2,000 rows within 10 seconds." The agent can check that criterion before it starts, which confirms it is looking at the same bug, and after it finishes, which confirms the bug is gone.

That gives the hand-back a natural order:

  1. Run the reproduction on the unchanged code and record that it fails, with the output.
  2. Write a regression test that encodes the reproduction, and record that it fails too.
  3. Make the fix.
  4. Run both again and record that they pass.
  5. Run the rest of the suite to show nothing else moved.

Step 2 is the one that keeps the bug fixed. A regression test written from the reproduction rather than from the fix is the strongest test an agent writes, for the reason set out in what AI-generated tests prove: it fails without the change by construction. The wording of the criterion itself follows the rules in acceptance criteria an agent can verify.

In Laimonade, create_bug over MCP takes acceptance criteria on the item from the start, and a P1 bug goes straight onto the current sprint's Ready column while lower priorities wait on the backlog. The agent that fixes it hands it back with the before-and-after runs as evidence. It cannot mark its own fix done.

What happens when the report names the wrong cause?

The agent fixes the named cause, and the bug survives. A diagnosis is the most persuasive sentence on a bug report and the least checked, so an agent tends to treat it as the specification.

A real case, from this product's own development setup in September 2026: a local database tool started failing with a connection error. Everything about the symptom said network. The message read "connection closed", and the obvious report would have been "database connection drops on startup". The cause was a directory renamed days earlier. The tool's launcher read a configuration file from the old path, found nothing and exited, and the client reported the dead process as a closed connection. A report that named the network would have sent an agent to retry logic and timeouts, all of which were fine.

Two habits stop a wrong diagnosis from steering the fix:

The same defect makes bug items a common source of the status drift in a backlog an agent can actually read: an item closed against the wrong cause reads as done while the bug is still live.

Frequently asked questions

What if the bug cannot be reproduced?

Then it is not ready for an agent, and the item should say so rather than be dispatched in hope. The useful work is getting a reproduction: more logs, the reporter's exact data, a stack trace from the failing environment. An agent can help with that as a separate investigative item whose output is a reproduction, not a fix.

Should the agent write the regression test first?

Yes. A regression test written from the reproduction and seen to fail on the unchanged code proves the agent found the bug the reporter saw. Written after the fix, it tends to encode the fix rather than the bug, and it passes whether or not the original symptom is gone. The recorded failing run is the evidence, not the order of the commits.

Should an agent file the bugs it finds while working?

Yes, as separate items rather than silent fixes inside the current one. A fix folded into unrelated work is invisible to review and to the backlog. Laimonade's create_bug records which item a bug was discovered from, so the provenance survives without the two pieces of work being merged into one.