On this page
A Practical Guide to Securing AI Agent Plugins at Work
tl;dr
Only 13.5% of enterprise AI agent platforms document all four critical governance capabilities, leaving plugin security gaps unaddressed. These plugins can execute code, access credentials, and inherit user permissions, creating widespread supply-chain risk for organizations.
Only 13.5% of 565 agentic AI platforms document all four enterprise governance capabilities in full, exposing a weak operational baseline for securing AI agent plugins in enterprise environments. A plugin isn’t a passive integration anymore. It can execute code, retain context, inspect telemetry, and inherit the permissions of the employee running the agent.
The governance data points in the same direction. Only 33% of platforms document deployment and data residency capabilities in full, compared with 49% for security and identity governance. Of 109 platforms that missed the four-capability bar by one item, 67 failed only on residency. The issue is less about whether a vendor has a familiar compliance badge and more about where its runtime, telemetry, and enforcement decisions actually live.
That’s what I’d call supply-chain inversion: enterprises assume plugins, marketplaces, error events, forms, and conversation history are trusted inputs, while attackers increasingly target exactly those distribution points. Your control model needs to answer four operational questions for every plugin:
- What code actually executes? A marketplace label or requested commit isn’t enough.
- What can that code reach? Files, repositories, credentials, networks, and production systems matter.
- Where does its telemetry go? A convenient control plane can become a new data boundary.
- Can you stop and inspect the run? Continuous enforcement beats a one-time review.
The first control isn’t a fancy agent-security platform. It’s an accurate inventory of plugins, source repositories, installed versions, permissions, and update paths. Artifact verification matters here because a successful installation and a successful deployment are very different claims.
What does Plugin4Shell prove about SHA pinning?
Plugin4Shell turns “we pinned the plugin” into a documentation exercise. The zero-click, high-severity remote-code-execution flaw affected Claude Code, Codex, GitHub Copilot, and Gemini CLI by exploiting a gap between requesting a pinned Git commit and verifying the code checked out under the host.
The attack became operationally serious because agents aren’t passive package viewers. AIR Security also found that background plugin updates are enabled by default in Claude Code and Codex, allowing malicious code to arrive on existing installations without user interaction. A security review performed during installation therefore says little about the code delivered six months later.
The vendor response remains uneven. Anthropic patched the flaw in Claude Code 2.1.179 and OpenAI patched it in Codex 0.146.0; Google deprecated Gemini CLI without patching, while Microsoft did not fix Copilot. That leaves buyers with a messy patch-management obligation: identify the agent, establish whether the vulnerable path is reachable, update where possible, and disable or isolate the path where it isn’t.
The broader lesson is simple: a pin is meaningful only if the installer proves that the installed bytes match the approved commit. Security teams should also record both the marketplace and the underlying source repository, since approving a catalog doesn’t necessarily mean approving the code source behind every entry. Our analysis of AI coding agent plugin marketplaces as a supply-chain risk goes deeper into why repository trust and marketplace trust are separate—and both fragile.
Which trusted inputs can become attack paths?
Public telemetry and ordinary business data can enter an agent context with no obvious sign that they’re executable instructions. The agent treats them as text, but its tools may give that text consequential power.
Sentry Seer’s “PhantomFix” is the clearest telemetry example. CVE-2026-90999 lets an unauthenticated attacker post fabricated events to a public Sentry DSN, after which Seer can embed attacker-controlled instructions in the prompt sent to a coding agent and induce installation of an attacker-chosen package. The attacker doesn’t need repository access or a Sentry account. They enter through the observability pipeline.
SalesBleed demonstrates the same pattern in CRM software. Zenity Labs disclosed three zero-click flaws in Salesforce Agentforce that hid instructions in public Web-to-Lead forms and exfiltrated CRM data through DNS subdomains embedded in image tags. The agent eventually performed the harmful action, but the malicious instruction originated in a field that humans had historically treated as data rather than authority.
These examples change the plugin threat model. You can’t sanitize only executable code and assume the context window is safe. Error messages, issue descriptions, web forms, support tickets, retrieved documents, and stored chat history all need provenance labeling and behavioral controls. A useful rule is blunt: if untrusted data can influence an agent, it must never be allowed to silently acquire tool permissions, package-install authority, or outbound network access.
The practical control is data-source isolation. Tag each context input, restrict the tools exposed while processing it, and require approval when the agent wants to convert external instructions into code or consequential actions. That’s more robust than asking a model to distinguish “data” from “instructions” after both have been blended into one context window.
Why do runtime and sandbox failures dominate plugin risk?
Runtime infrastructure is another trusted ingestion point, and recent vulnerabilities show why a plugin review can’t stop at package installation. CVE-2026-58138 is a critical unauthenticated remote-code-execution flaw in affected Orkes Conductor versions, with roughly 7,000 attack attempts recorded between September 2 and September 9, 2026. The workflow engine that coordinates agents and human approvals can itself become the execution path attackers use.
Sandboxes have the same problem when untrusted strings cross into shell commands. Two Amazon Bedrock AgentCore Python SDK vulnerabilities, CVE-2026-12530 and CVE-2026-16796, allowed command injection through crafted package names and could expose execution-role credentials from the metadata service. Package installation looked like an ordinary agent action. The implementation turned an untrusted name into a shell command.
Local context introduces a quieter version of the same failure. Agentic harnesses for Claude Code, AWS Kiro-CLI, OpenAI Codex, and Pi store conversation history locally without validating its origin, enabling conversation-history poisoning. A modified historical response can tell the next agent that an otherwise suspicious action is part of an authorized workflow.
You’ll get more protection from layered execution boundaries than from a longer prompt. Constrain package installation, isolate credentials, deny metadata access, separate untrusted retrieval from privileged tools, and monitor behavior against the agent’s normal pattern. “The model was instructed not to do that” is not a runtime control.
What should a runtime control plane actually enforce?
A useful control plane has to sit in the execution path. The core Model Context Protocol specification doesn’t define authentication, rate limiting, tool allowlisting, or audit logging, leaving those controls to the surrounding infrastructure. For plugin-heavy agents, that surrounding layer isn’t optional decoration; it’s the security boundary.
An MCP gateway is a proxy between agents and the Model Context Protocol servers they call. Unlike a static scanner, which inspects server code or manifests at one point, a gateway applies authentication, role-based access control, rate limits, and audit logging continuously. That distinction matters because plugin behavior changes after installation, permissions change after a role update, and data can be poisoned between reviews.
The minimum policy set should include:
- Agent and human identity bound to every request
- Tool-level and action-level authorization
- Filesystem, repository, and network boundaries
- Package-install and credential-use restrictions
- Human approval for privileged or irreversible actions
- Structured logs that preserve inputs, decisions, and outcomes
- Emergency revocation without waiting for a marketplace review
Connection-level authorization also isn’t enough. Enterprise MCP authorization in practice explains why an approved connection can still contain dangerous individual actions. A gateway must decide what the agent can do next, not merely whether the channel is open.
Which security architecture should an enterprise choose?
The main architectural split is data location. NeuralTrust’s TrustLens runs in the customer’s private environment and derives posture from observed behavior, while Noma Security’s Kong gateway plugin routes traffic, enforcement decisions, and audit logs to Noma’s cloud by default. One favors an in-perimeter native control plane; the other favors faster integration through an existing gateway ecosystem.
| Security approach | Published pricing | Primary security role | Best fit |
|---|---|---|---|
| Microsoft Agent 365 | $15 per user per month | Identity, governance, and security controls in Entra ID | Organizations centered on Microsoft identity |
| TrueFoundry MCP Gateway Pro | $499 per month | Runtime gateway with up to 25 server registrations and configurable rate limiting | MCP-heavy teams needing centralized policy |
| IBM ContextForge | — | Open-source gateway option | Teams prioritizing an open-source control path |
| NeuralTrust TrustLens | — | Private-environment posture and observed-behavior monitoring | Data-residency-sensitive deployments |
| Noma Security | — | Kong-based enforcement with vendor-cloud traffic and audit handling | Kong-centered teams accepting cloud telemetry |
Microsoft’s Agent 365 launched on May 1, 2026, making identity governance a separate procurement decision rather than something bundled invisibly into the agent runtime. It’s a sensible choice for Microsoft-centered fleets, but it won’t by itself verify plugin bytes, sanitize MCP traffic, or contain code execution inside every sandbox.
Funding and integration maturity deserve attention too, though they aren’t security evidence. Lasso Security has raised $28 million in total funding, but capital doesn’t establish data boundaries or enforcement quality. My preference is still an open or portable enforcement layer wherever feasible, paired with explicit policies that you own and can test.
What does agent control cost at scale?
Control cost comes in three forms: per-seat identity management, platform subscriptions, and runtime usage. The cheapest layer depends on the boundary you’re protecting. A plugin operating on a developer laptop creates endpoint and supply-chain risk; a managed agent creates sandbox, credential, residency, and spend risks.
Managed services trade infrastructure work for concentration risk. Claude Managed Agents costs $0.08 per session-hour on top of standard token rates, with an always-on session costing roughly $58 per month in runtime charges before token usage. That may be attractive for a continuously active workflow, yet you’re also accepting the vendor’s sandbox, execution, and state-management boundary.
Data requirements can be a harder gate than price. OpenAI Agents API is US-only without Zero Data Retention at launch, while Anthropic Managed Agents offers per-session dollar caps and isn’t restricted to US-only data residency. For regulated workloads, contract terms and processing boundaries should determine the shortlist before feature checklists do.
Native enterprise products can be simpler because identity and operations are already familiar. The Agentforce Sales Max edition costs $550 per user per month and includes 2.75 million Flex Credits. That bundles convenience with substantial credit exposure, so teams should understand how quickly automated workflows can consume the allocation rather than treating the included volume as unlimited capacity.
What should security teams do before expanding plugins?
Start with an execution-path inventory, not a list of approved marketplace names. For each agent-plugin pair, record the source repository, installed artifact, update mechanism, tools exposed, credentials available, network destinations, telemetry destination, and human approval points. If you can’t trace an action through that chain, you don’t yet understand its blast radius.
Then separate controls by where they execute. Use endpoint policy for laptop-installed agents, sandbox boundaries for code execution, an identity-aware gateway for tool calls, and per-service data controls for managed runtimes. A single dashboard can correlate those events, but correlation isn’t enforcement if the agent can bypass the path where the decision should happen.
Pilot with reversible permissions. Make package installation manual, restrict outbound traffic, deny metadata access, and use short-lived credentials. Add autonomy only when observed behavior supports it. This approach catches design flaws before an agent can quietly turn a public form, telemetry event, or compromised plugin into privileged execution.
The first architecture decision is concrete: which agents can install code, retain conversation history, or route telemetry outside your perimeter? Build the security boundary around those agents first, because they sit closest to the failure modes attackers are already exploiting.
Recommended Reading
-
MCP Authentication Explained: 2026 Security Best Practices
41% of public MCP servers have no authentication, creating critical enterprise security risks. The 2026 MCP specification and new enterprise-managed authorization extensions fix structural gaps in per-user OAuth consent. This guide outlines 2026 security best practices for MCP deployments.
-
Agent Delegation Patterns That Survive Production Reality
Bounded delegation tokens with enforced scope narrowing are critical to prevent inherited standing privilege, as only 13% of organizations currently have adequate AI agent governance. Reusable credentials passed between agents expand access at every handoff, while standards like Open Agent Passport D-004 mandate signed, traceable chains that shrink authority with each hop.
-
Tenant-Isolated Agent Tools: The 2026 Field Guide
Tenant isolation is the hidden cause of 40% of canceled agentic AI projects, per Gartner. Choose silo, pool, or bridge isolation models based on your regulatory exposure and platform capacity to avoid costly cross-tenant leaks.