On this page
GitHub's Own Docs Admit Its MCP Registry Can Be Bypassed
tl;dr
Baseline MCP governance is free via client-side allowlists, not costly gateways. GitHub's own documentation confirms its MCP registry can be bypassed by editing configuration files, so registries provide no runtime control. Gateways only justify their cost for servers holding shared service credentials.
GitHub’s documentation states that its registry-only policy “can be bypassed by editing configuration files,” per a governance comparison published last week — which is a polite way of saying the control layer most enterprises were told would govern their AI agents doesn’t actually stop anything at runtime. MCP server routing architecture is in the middle of two connected shifts right now: the protocol went stateless on July 28, 2026, and the governance story is quietly moving from centralized registries to client-side enforcement. Vendors are still selling you the old story. The evidence says otherwise.
I’ve started thinking about this as a control shift — the point where the thing that actually blocks a bad request moves from a catalog somewhere to the client making the call. Here’s what the data shows, and where the new spec opens gaps nobody’s tooling closes yet.
How Should MCP Server Routing Architecture Work After the Stateless Shift?
The short answer: plain HTTP routing, at the header layer, with no session affinity anywhere in your request path. The MCP 2026-07-28 specification removed the initialize/initialized handshake and the Mcp-Session-Id header, which means any server instance behind your load balancer can now answer any request. The sticky sessions, the shared session stores, the Redis cluster you stood up just to keep protocol state alive — none of it is required anymore.
Every request is now self-contained. Protocol version, client identity, and capabilities travel in a _meta field on each call rather than being negotiated once at connection time, as the stateless spec’s security analysis lays out. A client’s first message can be the actual tool call. That’s a genuine simplification, and if you’ve been running session-affine routing, it’s worth auditing what you can delete.
The piece most teams miss is the header layer. The new spec requires Mcp-Method and Mcp-Name headers on Streamable HTTP requests, which means a gateway can route and meter MCP traffic without parsing JSON bodies. Before this revision, your load balancer had to crack open every request body to figure out what kind of operation it was carrying. Now the operation type sits in the headers, which is exactly where WAF rules, rate limits, and routing tables want it. If you’re rebuilding your routing layer this year, build it header-first.
Where Do Routing Decisions Actually Happen Now?
Three places, and the order matters: the client allowlist decides whether a call should exist, the gateway decides what the call may touch, and the load balancer just moves packets. Conflating these is how teams end up paying for infrastructure they don’t need.
The protocol change also reshaped how servers talk back. Multi Round-Trip Requests (MRTR) replace server-initiated streaming: a server that needs input mid-call returns an input_required result, and the client retries with its answers attached, per the spec’s security writeup. We’ve covered the token cost mechanics of MRTR separately, because the retry pattern turns out to be the primary cost constraint in the new protocol — worth reading before you commit to a routing design that assumes one request equals one billable call.
One deadline should be on your calendar. The legacy HTTP+SSE transport has a 12-month deprecation window from the 2026-07-28 release, with removal planned for July 2027. If your routing stack still terminates SSE connections, you’re maintaining infrastructure with a known sunset date. Teams mid-migration on other priorities can reasonably wait a release cycle — but “reasonably wait” and “forgot to plan” are different postures, and only one of them survives an audit.
Do You Need an MCP Gateway, or Is That Vendor Messaging?
Here’s the contrarian position, and I think the evidence backs it: for most deployments, no. MCP itself provides no built-in authentication, rate limiting, audit logging, or access control — that’s a widely held view across the gateway landscape, per a 2026 gateway comparison. But the conclusion vendors want you to draw from that fact — “therefore buy a gateway” — doesn’t follow. The same gap can be closed at the client.
Enterprise allowlists enforced through Claude Code managed settings (allowedMcpServers plus allowManagedMcpServersOnly) and GitHub Copilot’s enterprise managed settings block unapproved MCP servers at runtime, at no additional cost beyond licenses you’re already paying. Compare that to the registry story: GitHub’s registry-only policy matches on server name or ID and can be bypassed by editing configuration files, and Azure API Center describes itself as design-time governance, not a runtime gateway. A registry is a catalog. Catalogs are useful. They are not controls.
Where gateways earn their keep is the narrow case of remote servers holding shared service credentials — a Jira bot token, a read-replica password. There, you need centralized credential vaulting and per-tool ACLs, and the gateway market has split accordingly: open-source projects like Bifrost, ToolHive, and agentgateway on performance, enterprise platforms like TrueFoundry, Composio, and Lunar MCPX on governance and managed operations.
Here’s the comparison, measured against a 300-developer workload with 25 approved servers:
| Option | What it controls | Blocks unapproved servers at runtime? | Price |
|---|---|---|---|
| Claude Code managed settings | Which servers the client may load, by URL or exact command | Yes, when the allowlist uses serverUrl/serverCommand | No separate price; a config file |
| GitHub Copilot enterprise managed settings | Which servers Copilot app, CLI, and VS Code may run | Yes, and fails closed on a malformed file | No separate price listed |
| Kong Konnect MCP Registry + AI Gateway | Registry plus per-consumer tool ACLs at the gateway | Yes, for traffic routed through it | From $25/serverless control plane/month; Enterprise custom |
| Azure API Center (MCP registry) | Inventory and discovery; feeds Copilot’s registry list | No — it is design-time | Free plan; Standard included with linked APIM Standard/Premium |
Notice the pattern in that last column. The two options that actually block at runtime for ordinary client traffic cost nothing beyond existing licenses. The one that costs money is justified only by the credential-vaulting case. If a vendor’s pitch starts with “you need a gateway for baseline governance,” that’s marketing, not architecture.
What Does Stateless MCP Break for Identity and Authorization?
The stateless spec is a net positive for security in one dimension and a genuine regression risk in another. On the positive side: no session state means no session to hijack, and header-based routing gives your WAF and gateway a clean policy enforcement point. On the risk side, the client identity carried in _meta is a self-reported claim — if you don’t bind it to an authenticated principal, your audit trail records whatever the client decided to say about itself.
That’s not a theoretical concern. With no session, there’s no session-scoped authorization state, so every request must be authorized on its own merits. Stricter in principle. Easier to get wrong in practice. A malicious actor who can spoof the identity field in _meta can potentially bypass access controls and poison your logs with actions attributed to someone else — and most current gateway and client tooling wasn’t built to bind these claims to real authenticated principals. The spec created the gap; the ecosystem hasn’t shipped the filler yet.
This is also where the stateless shift intersects with the governance question from the last section. If your runtime control is a client-side allowlist, the client’s identity claims are somewhat less load-bearing — the allowlist already decided which servers are approved. But the moment you put a gateway in front of shared-credential servers, identity binding becomes the difference between an audit log and a work of fiction. We’ve written about what changed in the stateless spec and the migration risks in more depth; the identity-binding problem is the one I’d prioritize, because it’s the failure mode that’s invisible until you need the audit trail in a dispute.
What Does MCP Governance Actually Cost?
Less than you’d think for the baseline, and impossible to budget precisely for the gateway tier — which is itself a data point. Kong’s Konnect MCP Registry + AI Gateway starts at $25 per serverless control plane per month, with Enterprise at custom pricing. Azure API Center has a free plan, with Standard included if you already run linked APIM Standard or Premium. The client-side allowlist tier costs a config file.
Here’s the uncomfortable finding for anyone building a business case: a 50-developer gateway deployment cost cannot be projected from published pricing. No per-seat or per-developer rates exist for the primary governance tools — gateway pricing is either contact-sales or custom-per-user with no disclosed rates. I can’t show you the math because the vendors won’t show the inputs. When an entire category hides its pricing behind a sales call, treat that as information: the buyers who sign are the ones who didn’t comparison-shop against a free allowlist.
The honest cost model, then, looks like this. Baseline governance — blocking unapproved servers — is effectively free with licenses you hold. Gateway cost applies only to the subset of remote servers with shared credentials, and for that subset you should demand per-seat or per-request pricing in writing before you commit. If a vendor won’t disclose rates for a 50-developer pilot, that tells you what renewal will look like.
How Should You Build Your Routing Stack?
Start with the free layer, add the expensive layer only where credentials force it, and validate the whole thing against a published standard. Concretely:
- Enforce at the client first. Push allowlists through Claude Code and Copilot managed settings, matching on server URL or exact version-pinned command. This is your default runtime control.
- Treat registries as catalogs. Use them for inventory and discovery. Never count them as blocking controls — the vendors’ own documentation says they aren’t.
- Gateway only the credential-holding servers. If a remote server holds a shared service credential, it goes behind a gateway with per-tool ACLs and audit logging. Everything else doesn’t.
- Route on headers, not bodies. Build your gateway and WAF rules around
Mcp-MethodandMcp-Nameso policy enforcement doesn’t require JSON parsing. - Bind identity before you trust audit logs. Whatever sits in your request path must tie
_metaidentity claims to an authenticated principal, or your compliance story is decorative. - Audit against the CIS benchmark. The CIS MCP Server Benchmark v1.0.0 provides 55 prescriptive recommendations across 10 security domains, covering both local and remote deployments including gateway and proxy layers — a vendor-neutral checklist for exactly the stack you’re assembling.
If you’re running one or two internal tool servers, stop reading and ship: a single stateless process behind a reverse proxy is the correct architecture at that scale, and a gateway would be overhead masquerading as diligence. If you’re running dozens of servers across hundreds of developers, the scaling decisions get harder, and the allowlist-plus-selective-gateway split is the pattern I’d defend with the evidence above.
The open question I’d leave you with: when your gateway vendor ships “stateless MCP support,” ask them specifically how they bind _meta identity to an authenticated principal. If they can’t answer in one sentence, their audit logs aren’t ready for the protocol you’re now running.
Recommended Reading
-
Agent Delegation Patterns That Survive Production Reality
Bounded delegation tokens with enforced scope narrowing are critical to prevent inherited standing privilege, as only 13% of organizations currently have adequate AI agent governance. Reusable credentials passed between agents expand access at every handoff, while standards like Open Agent Passport D-004 mandate signed, traceable chains that shrink authority with each hop.
-
MCP Cacheable Tool Results: A Production Caching Playbook
Caching MCP tools/list results cuts agent input-token costs by 30–60% for fleets with repetitive workloads. Incorrectly marking permission-filtered catalogs as public creates cross-user authorization leaks, so default to private scope unless you can prove responses are identical across all callers.
-
AI Agent Dependency Security: A Practical Control Guide
Effective AI agent dependency security requires an end-to-end control path from source code through sandbox execution, with enforceable financial and permission limits, not just standalone inventory tools. Autonomous agents expand attack surfaces beyond traditional CVE scanners, with documented incidents including 2,090 malicious RubyGems published in hours and unconstrained recursive loops incurring 50,000 USD in cloud costs in under an hour.