On this page
Tenant-Isolated Agent Tools: The 2026 Field Guide
tl;dr
Tenant isolation is the hidden cause of 40% of canceled agentic AI projects, per Gartner. Choose silo, pool, or bridge isolation models based on your regulatory exposure and platform capacity to avoid costly cross-tenant leaks.
Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, and a surprising share of those failures trace back to a problem buyers didn’t budget for: tenant isolation. If you’re evaluating tenant-isolated agent tools right now, you’re really evaluating whether one customer’s agent can ever see, touch, or infer another customer’s data — and whether the vendor’s answer to that question is architecture or marketing.
The stakes aren’t hypothetical. Gartner also expects the average Fortune 500 enterprise to run more than 150,000 agents by 2028, while only 13% of organizations believe they have adequate agent governance in place. That gap is where cross-tenant leaks live.
Here’s what the production evidence actually shows, what the tooling landscape looks like, and how to choose without overpaying for the wrong isolation model.
Why does tenant isolation break down with agents?
Traditional multi-tenant SaaS validates identity once at the edge, then trusts the request. Agents destroy that assumption because they generate their execution path at runtime — deciding which tools to call, which sub-agents to spawn, and what context to carry along.
The clearest evidence comes from a September 2026 arXiv paper formalizing what the authors call the stochastic deputy problem. In a 373-trial ablation across eight model configurations, a correctly validated tenant parameter in the tool schema served every out-of-scope attempt — 26 of 26. The model itself became the vulnerability: when the tenant parameter was removed from the MCP tool schema, 12 of 56 trials escaped the interface entirely by forging writable scope. The lesson is uncomfortable. If tenant identity is something the agent can see, it’s something the agent can rewrite.
Google’s engineering guidance for multi-tenant agentic systems reaches the same conclusion from the architecture side: tenant identity must survive every hop through the planner, sub-agents, tool calls, and model invocations, and it must be immutable once set. An agent that can rewrite its own tenant context is an agent with no tenancy at all.
And model-level safety training won’t save you. After the Hugging Face incident in September 2026, Anthropic itself pledged stronger containment — network egress controls, filesystem restrictions, process-level sandboxing — effectively conceding that model behavior is insufficient without infrastructure-level isolation. I’ve started calling this the session-boundary isolation pattern: the isolation primitive has to live at the session and credential layer, below the model, not in the prompt. We covered a related failure mode in our review of cross-tenant leaks in self-hosted AI tools, and the same logic applies here with more force.
Which isolation pattern should you choose?
Multi-tenant agent architectures have converged on three patterns, per AWS’s multi-tenant agent guidance: silo (dedicated resources per tenant, strongest noisy-neighbor protection), pool (shared resources with session isolation via unique session IDs), and bridge (a hybrid where some components run siloed and others pooled).
Silo is the audit-friendly answer. Axonius, which reconciles data from over 1,400 systems into one source of truth, runs each customer workload in a dedicated Amazon VPC — a silo deployment — and extended that model to its agent workloads rather than re-architecting tenancy. You pay more in infrastructure, but your compliance story is one sentence long.
Pool mode is where the interesting engineering happens. Amazon Bedrock AgentCore’s approach is session-isolated microVM-based compute — lightweight microVMs launched per session, without the cost or latency of full VMs. Benchling’s production numbers are the strongest public proof point: their AgentCore-based architecture handles more than 600 code execution sessions per day across 250+ tenants per week with zero security incidents. Their stack combines account-level isolation, Route 53 Resolver DNS Firewall, and VPC endpoint policies to block exfiltration — including the DNS-resolution vector that standard egress rules miss.
One honest caveat on pool mode: structural isolation can wreck database performance if you implement it naively. The arXiv study measured a 57x latency ratio from set-valued tenant scope under function-wrapped membership predicates, recovering index access only through a JSON_TABLE lateral join. Strict isolation and query performance are a real tradeoff, not a strawman — but Benchling’s zero-incident record at scale shows it’s a solvable one.
What do the managed platforms actually give you?
The managed layer has matured fast in 2026. Microsoft Foundry’s hosted agent isolation exposes two independent controls: user isolation (the caller’s Entra identity or a delegated identity from a trusted middle tier) and session isolation (a VM-isolated sandbox keyed by agent_session_id). Note the deliberate separation — a session is sandbox compute and persisted files, not a conversation.
OpenAI’s Agents API, built on the Codex harness, lets you choose where agents run: OpenAI-managed sandboxes, your own infrastructure, or sandbox partners. One early customer reported evaluation scores improving from 0.71 to 0.85 with a 4x latency reduction from subagent support. DigitalOcean Managed Agents went the consumption-billing route with per-second active CPU billing — you pay only for CPU actually consumed, with session creation to first response under two seconds and paused-work resumption around 300 milliseconds. For developer workstations specifically, Docker Sandboxes wraps coding agents like Claude Code, Gemini CLI, Copilot CLI, and Codex in isolated VMs with separately configurable network and filesystem permissions, plus a governance tier for MCP policy across the fleet.
Here’s how the main options compare:
| Platform | Isolation primitive | Pricing | Best for |
|---|---|---|---|
| Amazon Bedrock AgentCore | Session-isolated microVMs, VPC mode | — | SaaS providers on AWS needing pool-mode scale |
| Microsoft Foundry hosted agents | VM-isolated sandbox per agent_session_id + Entra user isolation | — | Entra-centric enterprises, middle-tier apps |
| Claude Managed Agents | Anthropic-run harness and sandbox | $0.08 per session-hour plus token rates | API-first teams wanting zero ops |
| DigitalOcean Managed Agents | Isolated sandbox per session | Per-second active CPU billing | Bursty workloads sensitive to idle cost |
| Docker Sandboxes | Per-agent isolated VMs on workstations | — | Governing local coding agents at fleet scale |
| OpenAI Agents API | OpenAI-managed sandbox, own infra, or partners | — | Teams already on the Codex harness |
A dash means the research didn’t surface a verified price — treat that as a procurement question, not an oversight. If you’re weighing these runtimes more broadly, our breakdown of agent execution engines and memory-layer lock-in covers the switching-cost angle.
Where do identity and governance layers fit?
Isolation primitives handle the “where code runs” question. They don’t answer “who is this agent acting for, and can I shut it off.” That’s a separate layer, and 2026 produced real products for it.
Okta for AI Agents added a Kill Switch that revokes active tokens at the Agent Gateway, extended shadow AI discovery to endpoints, and added visual configuration of agent connections. Solo.io’s Enterprise for agentgateway ships five capabilities from one deployment: STS-backed token exchange and OAuth flows, a read-write management UI, MCP tool governance, traffic observability, and Kubernetes-native control plane integration. WSO2 went the open route — Agent Manager is fully open-source, governing agents across any framework with per-agent, per-environment identity controls and a sandboxed runtime at GA.
On the data-access side, two implementations are worth studying. Datris brokers credentials through HashiCorp Vault so agents never handle secrets directly, runs agent-written scripts in a sidecar with no network route to the database, and enforces per-action policy at the platform rather than in prompts. Historis binds every request to the identity inside a verified token, resolving visible records through a backend-only Postgres SECURITY DEFINER function revoked from authenticated and anon roles — with a deploy-time invariant that fails the build if the boundary weakens. That last detail matters more than it looks: it turns isolation from a code-review hope into a CI-enforced property. For the runtime authorization piece specifically, our guide to building a custom EMA action-level authorization gateway walks the build-versus-buy decision.
What does tenant-isolated agent infrastructure actually cost?
Vendor sticker prices are close to meaningless here, and the variance is wild. AaaS pricing in 2026 has converged on three structures — per-token orchestration fees, per-action billing, and per-agent-hour rates — with search interest in the category climbing from 90 queries a month in June 2025 to 260–390 by Q1 2026 at a $32.88 cost-per-click, which tells you how hard vendors are fighting for these buyers.
The data points worth anchoring on:
- Claude Enterprise runs $20 per seat per month billed annually plus usage at standard API rates, with 20-seat minimums for self-serve and 50 for sales-assisted plans.
- Claude Managed Agents charges $0.08 per session-hour on top of token rates — no monthly fee, no per-agent license. A genuinely 24/7 agent lands around $58 a month before a single token.
- The most instructive number is self-built: German insurance broker MRH Trowe rolled secure agents out to roughly 400 employees in the first month at about $14 per seat in infrastructure-and-token cost, with a path to cut roughly 40% more through right-sizing and scheduled scaling. They explicitly rejected standard SaaS AI tools because those trade data isolation for setup speed — German financial regulations left no room for that trade.
Efficiency also compounds. On SWE-bench Pro with the same Opus 5 model, Teradata’s Tera consumed 73% fewer tokens than Claude Code while completing work 42% faster at 58% lower total cost — a reminder that harness design, not just model choice, drives your isolation-adjusted bill. My contrarian read: per-seat and per-token quotes are the smallest numbers vendors can legally publish, and buyers optimizing for them are exactly the ones feeding that 40% cancellation rate.
How should you decide?
Match the isolation model to your regulatory exposure and platform capacity, in that order:
- Regulated data, unforgiving auditors? Silo or managed-with-VPC-mode. MRH Trowe and Benchling are your templates — per-session compute isolation, identity passed server-side, DNS-level egress control.
- SaaS provider at tenant scale? Pool mode on session-isolated microVMs, but only if you can enforce context propagation below the agent. The stochastic deputy results say application-layer tenant parameters will eventually fail you.
- Thin platform team? Managed runtimes (AgentCore, Foundry, DigitalOcean) plus a governance layer (Okta, Solo.io, or open-source WSO2) beats a half-built custom stack every time.
- Whatever you choose, bind tenant scope to verified credentials, remove it from anything the model can see, and make the boundary fail CI rather than rely on review.
The open question I’d put to any vendor in this space: show me what happens when the agent itself tries to change its tenant context. If the answer involves the word “prompt,” keep looking.
Recommended Reading
-
Enterprise AI Risk Assessment: Unignorable Compliance Gap
68% of organizations that suffered 2026 data breaches had no AI governance policy, and unapproved shadow AI incidents cost $5.39 million per event. EU AI Act transparency rules now mandate immediate AI inventories, and compliance tooling costs are far lower than breach penalties.
-
Agent Delegation Patterns That Survive Production Reality
Bounded delegation tokens with enforced scope narrowing are critical to prevent inherited standing privilege, as only 13% of organizations currently have adequate AI agent governance. Reusable credentials passed between agents expand access at every handoff, while standards like Open Agent Passport D-004 mandate signed, traceable chains that shrink authority with each hop.
-
AI Search Entity Resolution: Costs, Tools, and Tradeoffs
A multi-stage cascaded architecture beats LLM-only entity resolution for AI search, resolving 90% of entities for under $20 per 1M daily turns. Unlike naive LLM-only approaches that incur high API costs, this tiered method reserves expensive model judgment for only the most ambiguous cases.