The Black Hat 2026 coding-agent flaws, and the pattern behind them
At Black Hat USA on August 5, 2026, Novee Security disclosed vulnerabilities in three of the most widely used coding agents. The three are Anthropic's Claude Code, Google's Gemini CLI, and OpenAI's Codex. Two received CVEs. All three share one shape underneath. Untrusted input reached an agent that could do more than the task required. That shape is the thing worth learning. It is not tied to any one vendor.
Below is what each flaw was, drawn from the official advisories. Then the pattern, and the hardening that follows from it.
CVE-2026-54316: Claude Code exfiltration through a pre-approved domain
The Claude Code flaw is tracked as CVE-2026-54316. Anthropic's advisory titles it "Out-of-Band Data Exfiltration via Pre-Approved HuggingFace Domain in WebFetch."
Here is the mechanism. huggingface.co was pre-approved as a bare hostname for the WebFetch tool. So any path on that domain was auto-approved, with no permission prompt. That includes model repositories an attacker controls. HuggingFace counts requests as downloads on its servers. That counting became a covert out-of-band channel. An agent could be steered into encoding data and leaking it through that channel. One request at a time, never touching a blocked destination.
It affects Claude Code from version 0.2.54 up to (not including) 2.1.163. The fix shipped in 2.1.163. The severity scores diverge by system. Anthropic rates it 6.0 (Moderate) under CVSS v4.0. The NVD rates it 9.1 (Critical) under CVSS v3.1. If you cite a number, say which system it came from.
The lesson is the pre-approval. An allowlisted domain is still a domain an attacker can host content on. A trusted host can still serve untrusted content.
CVE-2026-12537: Gemini CLI command injection before the sandbox
The Gemini CLI flaw is CVE-2026-12537. It is an OS command injection in the container launcher. A crafted .gemini/.env file reaches it. It fires in the launcher, so it runs at the host level before the sandbox is in place. That is pre-sandbox host code execution.
It is rated 10.0 (Critical) under CVSS v4.0, the maximum score. The fix shipped in Gemini CLI 0.39.1 and in the run-gemini-cli GitHub Action 0.1.22.
The lesson is where the boundary sits. A sandbox only protects what runs inside it. Setup code runs outside that protection. Reading a config file or launching a container happens before the walls go up. And a config file in a cloned repository is untrusted input.
The Codex finding: a writable instruction file that outlives the task
OpenAI's Codex was part of the same disclosure and did not receive a CVE. Here is the issue. Two passes can run in one job and share a checkout. The first pass can write an AGENTS.md file. The second pass then loads that file as instructions. That hijacks the next agent run. OpenAI's stance is that the sandbox did exactly what its docs say it does.
CVE or not, the shape is familiar. The agent's instructions came from a file that earlier untrusted work could write. Persistence plus a writable instruction surface lets one run steer the next.
The pattern
Read the three together and the same chain appears each time:
- Untrusted input enters. A model repo on an allowlisted domain. A
.envin a cloned repo. AnAGENTS.mdwritten by a prior pass. - The agent acts on it with standing reach. A pre-approved fetch. A launcher that runs shell. Instructions loaded without question.
- The reach touches something it should not. A covert exfil channel. Host-level execution. The next run's behavior.
None of these is a model being dumb. They are design choices about what the agent could do when handed input nobody checked. That is why "make the model more careful" does not fix it. The fix is on the reach.
What hardens against this class
The same controls apply to all three. Each cuts the chain at a link you control by machine, rather than a link you hope the model respects:
- Treat an allowlisted host as an untrusted content source. Approving a domain does not approve every path and payload on it. Scope approvals to what the task needs. Prefer denying network egress by default, so a covert channel has nowhere to send.
- Put the trust boundary before the setup code, not after. If a config file or launcher can run commands, it sits inside the attack surface. Validate or sandbox the thing that builds the sandbox.
- Treat instructions from a writable, task-reachable file as data. An
AGENTS.mdan earlier pass could write is plain input. It should never act as policy. - Separate the agent's identity from your secrets and CI. Least privilege means a compromised run lands in a throwaway environment. It never reaches your pipeline.
We wrote the general version of this threat model in a companion post on securing coding agents. Black Hat 2026 makes the same argument with real CVE numbers attached. The danger in a coding agent is rarely the model. It is untrusted input meeting standing reach. For a live example of the same writable-config shape in the wild, see the npm worm that plants itself in your coding agent's config.
Common questions
What is CVE-2026-54316? A Claude Code vulnerability in the WebFetch tool. The HuggingFace domain was pre-approved. An agent could be steered into leaking data through HuggingFace's server-side download counting, used as a covert channel. Fixed in Claude Code 2.1.163.
What is CVE-2026-12537? An OS command injection in the Gemini CLI container launcher. A crafted .gemini/.env file triggers it. It gives host-level code execution before the sandbox starts. Rated 10.0 under CVSS v4.0. Fixed in Gemini CLI 0.39.1.
Did Codex have a CVE? No. In the Codex finding, one pass writes an AGENTS.md that a later pass loads. It was disclosed in the same research but received no CVE. OpenAI stated the sandbox behaved as documented.
What is the common thread? In all three, untrusted input reached an agent with more reach than the task required. The durable defense is limiting that reach, rather than asking the model to be more cautious.