CI failure tracking: one card per failure, not per push

Stefan-Iulian Tesoi · · 6 min read

A worn steel control panel with a lit red lamp beside a green and an amber one and an analogue meter, each lamp a single labelled fault that is either showing or clear

CI failure tracking works when each distinct failure gets exactly one card, titled from the repository, branch and check name so a recurring failure finds its own card again. The card carries the run link, the commit and the failing lines, and the next green run of the same check moves it to review. Filing it is a lookup and an insert, not a judgement.

Most teams that automate this get it wrong in a way that looks fine for a week. The board fills with cards saying CI is failing and nothing else, and nobody can tell which of them describe the same fault.

What goes wrong with CI failure tracking?

They multiply, and they say nothing. Two cards sat on one board as P1 bugs, titled "CI failing on main" and "Fix CI failures on main branch". Their descriptions were boilerplate: "As a developer, I want the CI pipeline to pass reliably on main". Between them they named no repository, no workflow, no job, no commit and no failing assertion.

They were almost certainly the same failure filed twice, and nothing on either card made that decidable.

That was not an agent ignoring the facts. It was an agent never given any. The instruction it received carried three fields, the repository, the branch and the check name, and a title template of CI failing on <branch>. Every failure on main produced the same title by construction, so a lint break and a broken integration test became one card.

What should a CI failure card carry?

Enough to start fixing without opening GitHub, and nothing anyone has to scroll past. The test is whether the person who picks it up can go straight to the failing line.

The card is written as facts, not as a task. Anything phrased as an instruction is something a model reading the card may later report having done. The same rules apply as to a bug report written for a coding agent: the observed behaviour, where it happened, and the evidence, with the diagnosis left to whoever has the code open.

How do you stop duplicate tickets?

Make the title a function of the failure, and look for an open card with that exact title before filing.

The title is CI failing: <check> on <repo>@<branch>. A recurring failure computes the same string and finds its own card. Two checks failing on one branch, or one check failing on two branches, stay separate. A card that is done or archived does not count, so a failure that returns after its card was closed is a new occurrence and gets a new card.

The way this breaks is by letting anything reword the title. On 30 September a red check fired again while the previous day's card was still open. Laimonade handed the filing to its agent as a free-text intent, the agent improved the title while enriching it, and the exact-title duplicate guard compared the improved title against the board and found nothing. A second card was filed. The two drafts of the agent's reply were then wrong in opposite directions, one claiming a card under a title it had not used and the other claiming nothing had changed, and the reply check blocked both.

A check that stays red fires on every push. Whatever handles it has to be idempotent, or the board pays for every commit.

Why does the green run matter as much as the red one?

Because the green run is the only event that knows the failure has gone, and a card nothing can clear stays open until someone stumbles on it. The green run of the same check on the same branch recomputes the same title, finds the open card, and moves it to In Review. A person closes it from there.

That half had never worked. Green events carried an empty title, so the instruction asked the agent to resolve "an open backlog item titled """, which nothing could satisfy, and it ran anyway on every green check. On 14 September that came to 341 agent runs in one morning for one workspace, 253 of them in a single hour, and 6.65 million input tokens against 13,000 output tokens in total. A CI pipeline emits a check run per push per branch, so the bill tracked how much anybody was building rather than how much there was to do.

The fix was to ask the database whether there was a card to close before spending a model call to ask the same question.

EventOpen card for this title?What happens
Red checkNoOne P2 bug card is filed
Red checkYesNothing new; the existing card stands
Green checkYesThe card moves to In Review for a person to close
Green checkNoNothing. By far the commonest event

Should a model file the card at all?

No. Everything the card needs is computed before a model is involved, and a model adds a call, a reworded title and a chance to claim something it did not do. A failing CI check is a fact, and a webhook that turns a fact into a prompt has turned it into a guess.

Laimonade now files these cards directly: one indexed lookup, one insert, no agent run. The model earns its place later, when somebody picks the card up and needs the failure read and a fix proposed, or when a coding agent gets stuck on it and needs to say so.

This is one piece of the coding agent workflow a team can actually run: signals reach the backlog as cards a person can act on, and a card the green run clears still waits in review for a person to close. It is also the backlog-side twin of quality gates that fail quietly. A gate that never goes red hides failures; a tracker that files every red twice buries them. What the GitHub connection does with CI status is in integrations.

Frequently asked questions

What priority should an automatic CI failure card get?

Lower than urgent, until a person raises it. A webhook cannot tell a broken release branch from a lint warning on an experiment, and a priority that pulls work onto the current sprint is a decision about the team's week. Laimonade files these as P2 bugs in the backlog, and whoever triages decides whether one jumps the queue.

What happens with a flaky test?

It shows up as churn, which is the honest signal. A flaky check goes red and files a card, then goes green and moves it to review. If it fails again after the card was closed, that is a new occurrence and a new card. Three cards for one check in a week is a flake, counted, and the fix is the test rather than the tracking.

Should the job log go on the card?

No. The log is unbounded and mostly noise, and the failures in it are already on the card as annotations, each with a file and a line. A link to the run carries the rest for anyone who needs it. A card that has to be scrolled past a log to be read is a card nobody reads.