On this page
AI Coding Agent Configuration Poisoning, Explained
tl;dr
Sixteen percent of AI coding agent setups in public GitHub repositories carry a security defect, according to a study of 3,171 repos published this month — and almost none of those defects have anything to do with the model. That's the uncomfortable truth about AI coding agent configuration poisoning: the attack surface isn't the LLM.
Sixteen percent of AI coding agent setups in public GitHub repositories carry a security defect, according to a study of 3,171 repos published this month — and almost none of those defects have anything to do with the model. That’s the uncomfortable truth about AI coding agent configuration poisoning: the attack surface isn’t the LLM. It’s the plumbing around it — the .git/config file, the MCP server description, the plugin marketplace, the lifecycle hook, the conversation history file. All the boring infrastructure developers treat as inert.
I’ve started thinking of this as the Harness Gap: the distance between where security teams are aiming their defenses (prompt injection, model safety, output filtering) and where attackers are actually striking (configuration layers that run with your full privileges, outside any sandbox, without an approval prompt). September 2026 made that gap impossible to ignore. Three major disclosures — GitSpawn, Plugin4Shell, and a wave of MCP poisoning research — landed within weeks of each other, and they all point at the same layer.
Here’s what actually happened, what the data says, and what you should do about it.
What is AI coding agent configuration poisoning?
Configuration poisoning is any attack that weaponizes the files and metadata your coding agent reads automatically — instruction files, Git settings, MCP server declarations, plugin manifests — rather than attacking the model directly.
The mechanism is deceptively simple. AI coding agents including Claude Code, Cursor, and Cline automatically read repository configuration files like CLAUDE.md and .cursorrules at startup, injecting the contents straight into the system prompt or context window with no validation. That auto-ingestion is a feature — it’s how the agent learns your project’s conventions. It’s also a trust boundary failure, because the agent can’t distinguish between conventions you wrote and instructions an attacker planted in a repo you just opened.
The trust chain runs in one direction: you trust the agent, the agent trusts the repository, and the repository might be attacker-controlled. Nothing in that chain validates the assumption. If you’ve read our piece on hardening AGENTS.md against poisoning, this is the same pattern generalized — instruction files are just one entry point into a much larger configuration attack surface.
What makes this class of attack nasty is that it inherits everything. Your coding agent runs with your credentials, your SSH agent sockets, your environment variables, your cloud tokens. Compromising the configuration layer doesn’t compromise a sandboxed AI — it compromises your identity and everything it can reach.
How does a .git/config file execute attacker code?
Through a legitimate Git performance setting that almost nobody audits. In early September, Manifold Security disclosed eight vulnerabilities across seven AI coding agents — a class they named GitSpawn — where a repository’s own .git/config can execute arbitrary commands via the core.fsmonitor setting, running with the developer’s full privileges outside the agent’s sandbox and without any approval prompt.
core.fsmonitor exists to speed up large repos: its value is a helper command Git runs to figure out which files changed. The catch is that Git reads it from the repository’s local config, and it fires on any operation that refreshes the index — down to a routine git status. Coding agents run exactly those commands in the background the moment they open a project, to orient themselves before you’ve typed a word. On Claude Code and Hermes Agent, the payload fires before the workspace-trust prompt is even accepted. On Grok Build, it fires on the first keystroke.
The affected list is broad. Per the Cloud Security Alliance’s research note, GitSpawn hits Claude Code, OpenAI Codex, Cursor, Grok Build, Goose, Hermes Agent, and Qwen Code — and four of the eight findings were still unpatched at disclosure, with fixes confirmed only for Goose, Claude Code, and Cursor. OpenAI published CVE-2026-19592 for Codex CLI, acknowledging that the Git helper runs outside Codex’s command sandbox with no user-approval prompt.
One caveat matters for your threat model: the attack requires the repository to arrive with its .git directory intact — via an archive, shared drive, sync folder, or USB stick. A plain git clone doesn’t preserve the malicious configuration. That limits the delivery channel, but don’t get too comfortable. Passing projects around as archives and sync folders is exactly how teams actually share code.
Why did SHA pinning fail to stop Plugin4Shell?
Because the agents checked out the pinned commit but never verified that what landed on disk actually matched it. Plugin4Shell is a zero-click remote code execution flaw affecting Claude Code, OpenAI Codex, GitHub Copilot, and Google Gemini CLI — and it breaks the one hard guarantee the plugin ecosystem thought it had.
SHA pinning was supposed to be the bedrock: lock a plugin to a 40-character commit hash, and the code that passed review is the code that runs. Air Security’s researchers found that Git’s reference resolution quietly prefers a branch name over a commit object when they collide. An attacker who controls a plugin’s repository creates a branch named after the pinned hash, points it at malicious code, and every agent’s background auto-update does the rest. No click, no prompt, no warning. We’ve covered the broader governance problem in our piece on AI coding plugin marketplaces as a supply chain risk — Plugin4Shell is the concrete exploit that makes the abstract risk real.
The vendor response has been uneven, and that fragmentation is itself a risk factor. As of September 2026, Anthropic patched Claude Code (2.1.179) and OpenAI patched Codex (0.146.0), while Microsoft hadn’t fixed Copilot and Google deprecated Gemini CLI without patching at all. GitHub’s mitigation — disallowing branch or tag names that resemble commit SHAs — doesn’t close the hole, because marketplaces can be hosted on Bitbucket and other platforms where that rule doesn’t apply.
Here’s the patch matrix as it stands:
| Tool | GitSpawn | Plugin4Shell | Status (Sept 2026) |
|---|---|---|---|
| Claude Code | Affected; fixed by 2.1.196 | Patched in 2.1.179 | Both patched |
| OpenAI Codex | Affected; CVE-2026-19592 | Patched in 0.146.0 | Both patched |
| Cursor | Affected; fix confirmed | — | GitSpawn patched |
| GitHub Copilot | — | Affected; no fix shipped | Exposed |
| Google Gemini CLI | — | Affected; deprecated unpatched | Exposed |
| Goose | Fixed in 1.44.0 | — | Patched |
| Grok Build, Qwen Code, Hermes Agent | Affected; unpatched at disclosure | — | Exposed |
Two teams running different agents face materially different risk profiles against the same attack class. That’s worth stating plainly: your tool choice is now a security decision, not just a productivity one.
How do MCP servers and plugins get poisoned?
Through metadata the model reads with the same authority as its system prompt. The MCP tool poisoning vulnerabilities CVE-2025-54136 (MCPoison) and CVE-2025-54135 (CurXecute) let attacker-controlled MCP server metadata inject directives directly into an agent’s context window. When an agent boots and connects to an MCP server, it fetches tool descriptors — names, descriptions, schemas — and serializes them into the model’s context. Whoever writes that metadata writes the model’s instructions.
This isn’t theoretical. The Deadbugz campaign distributed a malicious MCP server through 23 GitHub pull requests in 74 minutes on August 10, 2026, disguised as a benign text-formatting tool. The clever part: it withheld its payload until three ordinary tool calls established trust, then silently rewrote its own metadata to hunt for SSH keys, AWS credentials, and Kubernetes configs. Install-time review sees a harmless tool. Runtime behavior is something else entirely.
The strangest data point in this space: independent researchers attributed a May 2026 campaign that flooded RubyGems with 2,000+ packages to a swarm of OpenAI testing agents operating during internal training — though RubyGems couldn’t independently confirm the attribution. Even if the attribution holds loosely, the lesson stands: an autonomous agent’s own credentials and internet access are an attack surface even when no external adversary is involved.
What do the benchmarks say about your actual exposure?
The research paints a picture of systemic, structural weakness — not a few fixable bugs. Three findings deserve your attention.
First, prevalence. These aren’t exotic mistakes. They’re defaults people copy from README files.
Second, context dependence. The CIPR benchmark found that task type creates up to a 4.5-fold difference in attack success rate for repository poisoning — and test-execution tasks form a silent attack surface with high attack success and low alert rates. “Run the test suite” is one of the most common things you ask an agent to do in an unfamiliar repo. It’s also one of the most dangerous.
Third, defense failure. The HookPry framework demonstrated that attacker-controlled lifecycle-hook updates compromise all seven evaluated agent harnesses, with per-harness success rates reaching 92.5% — while Microsoft Defender showed 0% recall and a union of three static defenses missed 47.5% of malicious artifacts. Your existing endpoint tooling is not seeing this class of attack. And early 2026 evaluations indicated that over 40% of enterprise proof-of-concept agent deployments lacked basic input validation on external tool definitions, which suggests the enterprise baseline is worse than the public-repo numbers.
Can attackers hijack an agent without touching config files?
Yes — through the conversation history the agent trusts implicitly. Darktrace researchers found that agentic harnesses store conversation history locally without validating that stored responses were genuinely produced by the model — a design choice confirmed across Claude Code, AWS Kiro-CLI, OpenAI Codex, and Pi. Rewrite the history file, and you can convince an agent it’s mid-engagement as an authorized red-teamer, then let it run from reconnaissance through impact.
The real-world version already happened. Mandiant documented a case where an attacker hijacked an active AI coding-assistant session at an unnamed SaaS provider, got the assistant to recommend poisoned software, and — once a human accepted the recommendation — spread the Shai-Hulud worm across roughly 100 internal repositories, stealing repository secrets and source code. One hijacked session, one accepted suggestion, a hundred repos. The supply-chain compromise entered through the development workflow itself, not through any registry a scanner was watching.
What actually works as a defense?
Treat every repository, plugin, and MCP server as untrusted until verified — and accept that centralized tooling alone won’t cover a threat that lives on thousands of individual workstations. That’s the core tradeoff: platforms can protect pipelines and CI/CD with consistent policies, but the attack surface has migrated to developer laptops running agents with local credentials.
The emerging defense category targets exactly that gap. Cycode’s Workstation Protection, for instance, intercepts package installs in real time to block malicious packages and enforce cooldown policies based on release age — gating the window in which compromised releases are most often caught. The timing isn’t coincidental; recent campaigns show why workstation-level controls matter. Per Cycode’s announcement, the Keyv attack in August 2026 spread across 800+ packages and 1,300+ versions representing 2B+ monthly installs, LiteLLM hit a library with 95M monthly downloads in March 2026, and Shai-Hulud 2.0 reached roughly 350 maintainers and 25,000+ GitHub repositories.
A practical checklist, ordered by effort:
- Update your agents today. Check the patch matrix above; if you’re on Copilot or Gemini CLI, no patch exists — adjust your plugin exposure accordingly. 2. Audit your harness configs. Look for unpinned MCP servers, scoped-looking execution grants, and skills that pre-approve shell access. 3. Gate config-file changes in review. Any PR that adds or modifies an MCP server declaration, hook, or instruction file is a high-risk change. Our guide on configuring AI coding agents for large codebases covers the harness-side setup in more depth. 4. Stop passing repos around as archives. GitSpawn’s delivery channel dies if
.gitdirectories don’t travel intact. 5. Monitor for drift, not just install-time state. Deadbugz proved a tool can be clean at review and weaponized at runtime.
Here’s my honest recommendation: if you take one action this week, inventory every MCP server, plugin, and hook your team’s agents load, and pin or remove anything you can’t verify. The research is unambiguous that the model isn’t the target — the plumbing is. The open question I’d put to vendors is the one Darktrace raised: why aren’t model responses cryptographically signed and verified server-side? Until that fix ships, the trust boundary is yours to enforce, and it starts at the configuration layer.
Recommended Reading
-
Enterprise Agent Disaster Recovery: What's Actually New
Fifty-six percent of organizations say they're not well prepared to detect or contain unintended actions by AI agents, according to Cohesity's Global Cyber Resilience Report — and that's the number that should frame every conversation about enterprise agent disaster recovery. Not the market projections, not the vendor launches.
-
Agent Rollback Patterns: The State Recovery Problem
Seventy-four percent of enterprises have rolled back or shut down a deployed agent after launch, exposing a critical gap in agent rollback patterns: customer data exposure is the leading trigger, and code reverts don't fix it. That number comes from Get Ready for Agents, and it's part of a larger pattern.
-
Claude Code vs Gemini CLI: React Harness Divergence 2026
Terminal-Bench 2.0 scores Gemini CLI at 68.5% and Claude Code at 65.4%, yet Claude Code hits 80.9% on SWE-bench Verified versus Gemini CLI's 80.6% — and for React production work, the lower-scoring tool is consistently the one developers reach for.