Pull request acceptance criteria: a verdict, not a gate

Stefan-Iulian Tesoi · · 6 min read

A railway crossing's hooded warning lamps and red-and-white crossbuck against a clear sky, the striped barrier arm standing raised, a signal that warns without closing the road

Pull request acceptance criteria are worth checking on the pull request itself, because that is where the merge decision is made, and the result belongs in a comment rather than a required check. A false fail stops honest work and teaches people to route around the check; a false pass is a recorded mistake a reviewer can still catch.

Most teams check criteria at two moments, if at all: when the work is handed over, and when somebody rereads the ticket after the merge. The first happens before a reviewer has looked. The second happens after it stopped mattering.

Why check pull request acceptance criteria at merge time?

Because the merge is the last point at which a verdict can change an outcome cheaply. Before it, a missed criterion is a comment and a fix. After it, the same miss is a bug report, a revert or a customer noticing.

Laimonade's verifier ran only at hand-over for its first month. It read the item's criteria, the linked commits and the tests, and wrote its verdict onto the backlog card. The card is read by whoever triages the board. The reviewer reads the diff, the description and the CI status, in a different tool, and never saw it.

That gap widens with coding agents. When an agent wrote the change, the pull request is often the first place a person looks at the work at all, so it is where the criteria have to be visible.

Should the check block the merge?

No, and the reason is an asymmetry in what each kind of error costs. A verdict produced by a model will sometimes be wrong, and a blocking check turns every wrong verdict into someone else's stalled afternoon.

GitHub makes blocking easy. Branch protection can mark any check as a required status check, and a pull request cannot merge until it passes. That is the right tool for a test suite, which is deterministic. It is the wrong tool for a judgement.

ErrorAs a required checkAs an advisory check
False failHonest work is stopped until someone overrides the check, and then disables itThe reviewer reads it, disagrees, and merges
False passMerged behind a green tick that reads as proofMerged with a comment a reviewer could have questioned, and it is on the record
Cannot verifyIndistinguishable from a failSays what evidence was missing

The first row decides it. A check that blocks good work twice gets switched off, and every correct verdict it would have given afterwards goes with it.

A comment that reads as a gate while being advisory is how an advisory gate gets switched off.

So tone is part of the design. The verdict reports what it found and does not tell the reviewer what to do. The merge decision stays with the person who has read the diff.

Laimonade's verifier was measured on twelve labelled cases in September: no false passes, no false fails, and two abstentions where the truth was a fail. It under-calls failure rather than inventing it. Twelve cases is a small sample and the model behind it has changed since, which is one more reason for it to advise.

What does a verdict have to read?

Everything a machine observed, plus the agent's own claims under a label that says they are claims. A verdict built from the agent's summary only grades the summary.

Each criterion then gets one of three marks, met, not met or cannot verify, with the evidence cited by file and line. Unmet criteria come first, so a reviewer who reads one line learns the thing that would change their decision. Tests that establish nothing get their own section.

This does not replace reading what an agent actually did. It tells the reviewer where to read first.

How does a verdict find the right backlog item?

Only through a link somebody made on purpose: an item id in the branch name or the pull request description, or a commit already matched to the item. Anything cleverer fails in the worst direction.

A guess that misses does not produce a missing comment. It produces a confident verdict about the wrong item on somebody's pull request, which is worse than silence. Issue references are the tempting shortcut, and they collide: #12 exists in every repository a team owns.

Most pull requests have no backlog item behind them, and those get no comment at all. Neither does an item with no acceptance criteria. A note saying "nothing to check" on every push is true and is noise.

A verdict is about one item. A pull request that carries two pieces of work is judged against one of them, and the other goes unexamined.

Why should the comment edit itself?

Because every new comment is an email. GitHub notifies everyone following the pull request: the author, the reviewers, anyone who commented and anyone watching the repository. A check that runs on every push and posts each time sends that whole list an email per push.

Laimonade's verdict carries a hidden marker per backlog item and edits its own comment instead. An edit sends no new notification, so a pull request produces one email when the first verdict lands. Every later push updates the same comment, and the current verdict is always in the same place.

The cost of getting this wrong lands on someone else. The thread being filled is a customer's review conversation, and a tool that makes it harder to read gets uninstalled.

This is the merge-time half of a definition of done when an agent wrote the code: evidence collected by something other than the author, put in front of the person who decides. The advisory check is useful only while it can never fail anyone's pull request, which is the same property quality gates that fail quietly lose in the other direction. What the GitHub connection reads and writes is listed in integrations.

Frequently asked questions

Does an advisory check just get ignored?

Some of the time, and that is cheaper than the alternative. A verdict that is right gets read, because its first line names the unmet criterion. A verdict that is wrong gets disagreed with and merged past. A blocking check that is wrong gets overridden, then disabled, and every right verdict after that is lost with it.

What does cannot verify mean on a verdict?

That the evidence did not settle the criterion either way. No test touched it, no check covered it, or the change sat outside what the verifier could read. It is not a failure and should not be read as one. It marks the place a reviewer has to check by hand.

Can one pull request carry two backlog items?

It can, but the verdict judges it against one. Two items in one pull request mix two sets of criteria in one review, and whichever item the link resolves to is checked while the other is not. One item per pull request keeps the verdict, the review and the history about the same piece of work.