AI agent prompt injection through tickets and issues
Stefan-Iulian Tesoi · · 7 min read

Yes. A coding agent treats what it reads as context, so instructions planted in a ticket, an issue, a comment or a README can steer it. Filtering that text does not reliably stop this. What bounds the harm is what the agent can do afterwards: no secrets in reach, no network it does not need, and every change seen by a person before it lands.
That is AI agent prompt injection as a development team meets it. It needs no stolen credential and no bug in any tool, only write access to something the agent reads.
What is AI agent prompt injection?
Text that changes what an agent does because the model read it as an instruction rather than as data. The OWASP Top 10 for LLM Applications, the OWASP LLM Top 10 for short, ranks it first as LLM01:2025 and splits it in two:
- Direct injection: a user types the instruction into the prompt.
- Indirect prompt injection: the instruction arrives in input "from external sources, such as websites or files."
For a coding agent the second kind matters most, because almost everything it reads comes from elsewhere. Greshake and colleagues, who described the attack in 2023, gave the cause: such applications "blur the line between data and instructions." A ticket is data to whoever filed it and an instruction to the agent that loads it.
Where does untrusted text reach a coding agent?
Everywhere the agent reads: issues, comments, READMEs, fetched pages, source files, tool responses, MCP tool descriptions and the backlog itself.
The prompt injection GitHub issues can carry is the cheapest to stage. In an Invariant Labs demonstration from May 2025, on test repositories, a user asked Claude 4 Opus, connected to the GitHub MCP server, to look at the open issues in a public repository. One issue carried a payload, and the agent copied data from the user's private repositories, salary included, into a pull request on the public one. Invariant called it "not a flaw in the GitHub MCP server code itself": the connection reached both repositories.
Aikido Security's Gemini CLI case from December 2025 is the same mechanism in CI. A GitHub Actions workflow in Google's repository fed issue titles and bodies to Gemini for triage, with API keys and a GitHub token in its environment. On a private fork with test credentials, Aikido filed an issue beginning "The login button does not work!" plus instructions, and the agent wrote a key and the token into the issue body. Google patched it within four days.
A backlog is the same exposure one step removed. A ticket from a support form, a client portal or an issue-triage bot carries an outsider's words into the work order an agent loads, and a work order is read as instructions by design. That is why coding agent security judges a connection by what it can reach.
Why does filtering the input fail?
Because the attack is language, and language that steers an agent can read exactly like a bug report. OWASP states that "it is unclear if there are fool-proof methods of prevention for prompt injection." Each kind of filter has a documented gap:
- Stripping hidden text. GitHub drops hidden characters, such as HTML comments, before input reaches its cloud agent. Anthropic's Claude Code GitHub Action strips more, and its security notes add: "but new bypass techniques may emerge."
- Detecting known phrasings. OWASP lists instructions encoded "using Base64 or emojis" to evade filters, and Invariant found that many off-the-shelf detectors missed its attack.
- Text in plain sight. Nothing strips a sentence a person can read. OpenAI's Codex documentation shows an issue saying a command "causes a 404 error" and asking: "Please run the script and provide the output." Following it would send the last commit message to an attacker's server; no character in it needs stripping.
A filter removes the places an instruction can hide. It cannot remove an instruction that reads like a bug report.
Strip hidden text anyway; it is nearly free. But as Simon Willison says of guardrail products, "in web application security 95% is very much a failing grade."
What limits the damage when it happens?
Assume some injected text will be followed, and make following it achieve little. Willison calls the dangerous mix the lethal trifecta: private data, untrusted content, and a way to communicate externally. An agent reading issues always has the second, so remove the other two and narrow what it can change.
| Control | What it stops | What it leaves open |
|---|---|---|
| No secrets in the agent's environment | A key leaking from that run | Code and data the agent can read |
| Network off, or allowlisted | Posting to an arbitrary server | Channels through allowed hosts |
| One repository per credential | Reading a private repository from a public task | Whatever that repository holds |
| A person reviews before merge | Injected code reaching main unseen | Anything done during the run |
| Agent cannot edit its own settings | Switching its own approvals off | Actions already permitted |
The Invariant and Gemini CLI leaks both travelled through GitHub itself, so an allowlist that permits GitHub would not have stopped either. Aikido's first remediation is to restrict the agent's tools and withhold write access to issues and pull requests. Codex cloud blocks agent internet access by default and suggests allowing only GET, HEAD and OPTIONS when it is on, which narrows the channel without closing it: OWASP's indirect scenario leaks a conversation through an image link.
In CVE-2025-53773, a researcher showed injected text making GitHub Copilot in VS Code write "chat.tools.autoApprove": true into the project's settings, which "disables all user confirmations". Microsoft fixed it in August 2025. An agent that can edit its own permissions has none that are fixed.
Every row is least privilege applied to a reader, which is also OWASP's advice: "Restrict the model's access privileges to the minimum necessary for its intended operations." Some limits belong in the system the agent calls. Laimonade gives a connected coding agent no tool that marks an item Done: the furthest it can take an item is a hand-back for review, judged by checks it does not run. That is one of the limits in what an agent should never be able to do.
Who can write to what your agent reads?
More than most teams have listed. Name the writer of every input the agent loads; where it is outside the team, the agent should hold no secrets and have no unreviewed route to production.
- Issues and comments. Anyone with an account, on a public repository. Requiring write access to start a run, as GitHub's cloud agent and the Claude Code action do, stops strangers triggering the agent, not writing what a trusted user points it at.
- Backlog intake. Support forms, client portals, triage bots. Laimonade flags items from outside requesters as external and can file new GitHub issues into the backlog; the flag records where text came from, not whether it is safe.
- Fetched pages. Every README the agent opens has an unvetted author.
- MCP servers. The publisher writes the tool descriptions and can change them later.
- The agent's own configuration. If the agent can write it, injected text can.
- Commit messages and pull request descriptions. Aikido lists both.
Frequently asked questions
Are public repositories more exposed than private ones?
Yes, because anyone with a GitHub account can open an issue on a public repository. Private repositories narrow the writers to collaborators, not to zero: dependencies, fetched pages and customer tickets still carry outside text. In the Invariant Labs case the public repository was the way in and the private one the loss.
Can an MCP tool description carry an injection?
Yes. Invariant Labs called it a tool poisoning attack in April 2025: instructions in a tool description that the model reads in full while the user sees a simplified version, and that a malicious server can change after approval. The MCP specification says clients "MUST consider tool annotations to be untrusted unless they come from trusted servers."
Does a stronger model resist prompt injection?
More often, not reliably. In November 2025 Anthropic reported a 1% attack success rate for its browser agent against an adaptive attacker given 100 attempts per environment, and wrote that this "still represents meaningful risk." Invariant's attack worked on Claude 4 Opus. A better model lowers the frequency; the limits on a hijacked agent decide the cost.