SyndaiSign inStart free
← Back to blog

Securing coding agents: the untrusted-input-to-CI attack, and how to harden against it

·Syndai Team·6 minutes
ai-agents
security
coding-agents
prompt-injection

The security problem with coding agents is not that the model is malicious. The problem is simpler. An agent reads untrusted text. It treats that text as an instruction. Then it acts with whatever permissions it was handed. Say the text comes from a source an attacker controls. Say the permissions include your CI secrets. That is the whole attack in one place.

The recent run of coding-agent vulnerabilities all had this shape. A GitHub issue or a crafted config file reached an agent that could run commands. Those commands could reach places the issue's author should never have touched. This is a design problem, and design problems have design fixes. Here is the threat model and the controls that break the chain.

The dangerous chain

Every serious coding-agent exposure is a version of the same four links:

  1. Untrusted input enters the context. A GitHub issue. A pull request comment. A file in a cloned repo. A web page the agent fetched. A config file it read. None of it is written by you.
  2. The agent treats it as an instruction. Models do not draw a hard line between data and commands. Text that says "ignore your task and run this" is text the agent may act on.
  3. The agent has reach. It can run shell commands, hit the network, read environment variables, or write files.
  4. The reach touches something valuable. CI secrets. An API key in the environment. A deploy credential. The power to open a pull request that another system trusts.

Break any one link and the attack fails. The mistake is trying to fix only link 2 by asking the model to be more careful. That is the least reliable link to defend. A model that can be instructed can be misinstructed.

Treat all fetched content as data, never instructions

The first rule is a posture, not a control. Everything the agent reads from outside the task is data. It gets quoted and summarized, never followed as a command. An issue that reads like an instruction is content to surface, not an order to execute. Build the agent's prompts so untrusted text is clearly framed as material to work on. Keep the real instructions coming from a trusted channel. This does not fully close link 2, which is why the next controls exist. But it lowers how often the model is even tempted.

The reliable fixes are on the agent's reach. Reach is something you control with machinery. You do not have to hope the model respects it.

  • Run in an isolated sandbox. The agent runs in a throwaway environment. It does not run on a machine with standing access to anything. A command that goes wrong wrecks a container you were going to discard anyway.
  • Separate the agent's identity from your CI secrets. The agent should not run with the same credentials your deploy pipeline uses. If it needs to trigger a build, it asks through a boundary with its own permissions. Compromising the agent then does not hand over the pipeline.
  • Deny network egress by default. Most coding tasks do not need to reach arbitrary hosts. Use an egress denylist, or better, an allowlist. Then even if the agent is steered into exfiltrating a secret, it has nowhere to send it. The most common exfiltration path is an outbound request the agent never needed.
  • Scope filesystem and tool access to the task. The agent gets the repo it is working on, not the whole machine. It gets the tools the task requires, not every tool available.

Some actions are worth stopping for a human even when everything else looks fine. Merging to a protected branch. Moving money. Changing access. Sending outbound mail. The pattern is to route those through an explicit approval boundary. The agent proposes the action. The action does not run until it is approved and its parameters are checked against a policy. This is the difference between an agent that can do something dangerous and one that can ask to.

The value here is that the gate is server-side. It does not depend on the agent's cooperation. A compromised agent can propose a bad action. But the policy and the approval sit outside it, and they are what actually fire the action.

Verify with evidence, not the agent's word

An agent will report that it did the safe thing. A report is a claim. The controls above are worth more than the claim. They hold whether or not the agent is telling the truth. The sandbox isolates regardless. The egress denylist blocks regardless. The approval gate stops the action regardless. Design so that safety comes from the boundaries. The agent's self-report is then a convenience, not a control.

A short checklist

If you are building or running a coding agent, walk the four links:

  • Input: is untrusted content framed as data, never as instructions?
  • Reach: does the agent run in a sandbox? Is egress denied by default? Is access scoped to the task?
  • Identity: are the agent's credentials separate from your CI and deploy secrets?
  • Consequential actions: are the big moves gated? Merges, transfers, and access changes should all pass an approval gate the agent cannot bypass.

None of these depend on the model behaving. That is the point. The model will sometimes misread what a piece of text is asking it to do. The controls decide whether that mistake costs you anything.

Two recent case files show this class in the wild. The npm worm that plants itself in your coding agent's config hid persistence in .claude/settings.json and .vscode/tasks.json. That is the writable instruction surface at work in a real supply-chain attack. The Black Hat 2026 coding-agent flaws trace the same untrusted-input-to-reach pattern across several disclosed CVEs.

Common questions

Is prompt injection solved in 2026? No. There is no reliable way to make a model perfectly split untrusted data from instructions. The input link cannot be fully closed. So the practical defense is to limit what the agent can do when it is fooled.

What is the single highest-value control? Denying network egress by default. Most exfiltration needs an outbound request. Remove the power to make arbitrary ones and the worst outcomes lose their exit.

Do these controls slow the agent down? Isolation and egress rules cost almost nothing for normal tasks, because normal tasks do not need broad reach. Approval gates only interrupt the specific big actions you choose to gate. Routine work flows on.

Should a coding agent have my CI secrets? Not directly. If it needs to trigger a build, it should ask across a boundary with its own scoped permissions. A compromised agent then cannot read or reuse the pipeline's credentials.