• 8 min read

MCP Stateless Architecture Explained: What Changed and Why

tl;dr

Stateless MCP simplifies infrastructure by eliminating session stores and sticky routing, but shifts state management to application code. Migrations risk hidden reliability issues like lost stream resumability and duplicated side effects for non-idempotent tools. Building a production stateless MCP server costs $100K to $1M upfront plus $5K to $25K monthly maintenance, with low-scale infrastructure at $100 to $500 per month.

Featured image for "MCP Stateless Architecture Explained: What Changed and Why"

GitHub deleted its Redis session layer this summer, and the deletion is the most honest summary of MCP stateless architecture you’ll find. When the GitHub MCP Server adopted the 2026-07-28 specification, it removed Redis sessions and deep packet inspection — database writes on initialize disappeared, and database reads disappeared from every subsequent call. That’s the promise of the stateless revision in one changelog entry: less infrastructure, faster calls, nothing lost for users.

The catch is that the state didn’t vanish. It moved. And where it moved — into your application code, your gateways, and your retry logic — determines whether this migration saves you money or quietly introduces a reliability regression you won’t see until production.

What did the 2026-07-28 specification actually change?

On July 28, 2026, the MCP project published specification version 2026-07-28, making the protocol core stateless by removing the initialize/initialized handshake and the Mcp-Session-Id header. Under the old model, a client’s first request negotiated a session, the server minted an ID, and every later request had to echo that ID back to the same instance that issued it. Under the new model, a client’s first message can be the actual tool call.

Every request is now self-describing — it carries its own protocol version and client context, so any server instance behind a standard load balancer can handle it. The revision eliminates the protocol-level requirement for sticky sessions and shared session stores, which is the infrastructure tax that made remote MCP servers painful to scale in the first place.

Three smaller additions matter as much as the subtraction:

  • server/discover: an optional method that returns supported protocol versions, capabilities, and identity in one response, for clients that want to check before calling.
  • Mcp-Method and Mcp-Name headers: method and tool names now travel in HTTP headers, so gateways can route, authorize, and throttle without parsing JSON bodies.
  • Cacheable list results: ttlMs and cacheScope hints let clients stop refetching tool catalogs on every reconnect.

If you want the full migration picture, we’ve covered the July 2026 spec changes and their cost implications separately. The short version: the protocol now behaves like an ordinary HTTP API. Whether your deployment gets simpler depends on what you do with the state the protocol stopped managing.

Does stateless MCP mean your application is stateless?

No — and this is the misunderstanding that will burn the most teams over the next year. Statelessness describes the protocol, not the application. When a workflow spans multiple calls, application state must be managed explicitly via opaque handles passed as tool arguments and stored in external datastores.

AWS’s framing is the clearest: think of it as a coat check. The old protocol was a valet who remembered your face, so you had to keep dealing with that same valet. Now you get a numbered ticket, and any attendant can serve you because the ticket carries the reference. A tool returns an identifier — a basket_id, a job_id — the model includes it on the calls that follow, and the actual state lives in your datastore.

This is a pattern I’ve started calling Explicit State: the stateless revision transfers state management responsibility from infrastructure (sessions, sticky routing, shared stores) to application developers (explicit handles, idempotency keys, durable tasks). Infrastructure gets simpler. Applications get harder to build correctly. ArchCrux’s breakdown separates four distinct kinds of state that survive the revision — interaction state from multi-round-trip flows, durable asynchronous Task state, business effect state in external systems, and the routing metadata the protocol now exposes. Each has a different owner, lifetime, and recovery rule. Treating them as one generic “agent session” reintroduces exactly the coupling the spec removed.

There’s a genuine upside here, and it’s easy to miss: the identifier now sits in the model’s context rather than hidden in a header. The model can reason about it and thread it across tools. That’s a better design than implicit session state — but only if you build for it deliberately.

What silently breaks when you migrate?

Stream resumability, and nothing will warn you. The Last-Event-ID header and SSE event IDs are gone from Streamable HTTP, which means an interrupted response stream now loses the in-flight request outright, and clients must re-issue it. Under the old transport, a dropped stream could resume and redeliver undelivered messages. Now the client retries with a new request ID.

Here’s why that’s dangerous: your tests pass. SDKs maintain backward compatibility. The migration looks mechanical. What you get instead is a quiet reliability regression that only surfaces under real network conditions, as occasional lost tool calls. For an idempotent read, harmless. For a tool that charges a card or provisions infrastructure, a lost request followed by a client retry is a duplicated side effect. The spec doesn’t solve this — side-effecting tools need an idempotency key supplied as a parameter, and that’s application work nobody will prompt you to do.

The second trap is elicitation. If you built on the previous spec’s elicitation design, the rewrite is an architectural inversion requiring significant rework, not a mechanical rename — notifications/elicitation/complete and the elicitationId field were removed outright, with no deprecation window. The replacement, Multi Round-Trip Requests (MRTR), turns server-initiated callbacks into a retry loop: the server returns an input_required result, and the client comes back with the answers attached. That changes your control flow, and it changes your cost model too — we’ve broken down how MRTR makes token cost the primary constraint in a separate analysis.

Where should you deploy a stateless MCP server?

The protocol now natively fits serverless and edge runtimes, yet most industry guidance still defaults to containers. That mismatch is the most interesting deployment story of 2026.

AWS explicitly identifies Lambda as a suitable deployment option because the protocol no longer requires persistent session connections. Cloudflare goes further, stating that MCP servers can now run in just a Worker, with no stateful infrastructure needed. And yet the default production path remains containers on Amazon ECS or EKS, carrying autoscaling tuning and per-session IAM complexity into a workload that is now mostly request/response.

Deployment optionWhy it fits stateless MCPWhat you still manageDocumented cost data
AWS LambdaRequest/response model matches the protocol; no persistent connections requiredFunction duration caps, cold starts, gateway auth—
Cloudflare WorkersGlobally distributed, no stateful infrastructure neededDurable Objects or KV for application state—
Containers on ECS/EKSFamiliar tooling, existing pipelinesAutoscaling, IAM scoping, the full Kubernetes operational surface—

My take: the stateless protocol eliminates the technical justification for container orchestration as the default. The actual cost drivers — authentication, audit logging, governance — don’t change with your runtime, and they’re better handled by managed gateways than by platform teams rebuilding them on Kubernetes. If you’re designing fresh, start serverless and reach for containers only when you hit a concrete constraint, like long-running streaming that outgrows function duration limits. For deeper pattern-level guidance, see our breakdown of stateless MCP server architecture patterns.

What does MCP stateless architecture cost to build and run?

The protocol is free. Everything around it isn’t, and the numbers are bigger than most teams expect.

According to World Programming’s 2026 cost analysis, a level-1 read-only MCP server costs $100K to $300K partner-built, a level-2 actions server runs $300K to $700K, and a level-3 agent-resident build starts at $1M — plus ongoing maintenance of $5K to $25K per month. The server code itself is the cheap part; auth, audit, and the safety story are where the budget goes.

On the implementation side, per designedbyai.io’s architect guide, a first production server takes 3–6 engineer-weeks, with 1–2 weeks per additional server, and infrastructure at low-to-medium scale typically runs $100 to $500 per month. The same source recommends keeping tools per agent context at 20–25 for reliable selection.

Performance overhead is modest but real: protocol overhead runs 8 to 40 milliseconds per tool call, and the same guide suggests MCP adoption is justified when you’re integrating three or more external services, with a recommended 30-day parallel testing period for rollouts. Below that threshold, a thin API wrapper is probably simpler and cheaper.

Who is already running this in production?

The adoption signals are stronger than the typical “new spec” cycle. GitHub shipped support ahead of the official release and, as noted up top, used the migration to delete session infrastructure entirely. Morgan Stanley has deployed over 110 APIs in production using MCP paired with FINOS CALM for architecture governance, with deployment gates and automated checks across more than 100 production services.

The most consequential signal may be institutional: CNCF TOC Initiative #1746 is evaluating MCP as the default wire specification for distributed agentic systems on Kubernetes, with workstreams covering protocol interoperability, agentic gateway conformance, and runtime abstraction. If that evaluation lands, MCP stops being a vendor-originated protocol and becomes cloud-native plumbing — which raises the stakes on getting your stateless architecture right now rather than after the reference patterns harden.

Should you migrate, and to what?

Yes, but not mechanically. The migration is safe for read-only, idempotent tools and genuinely risky for anything with side effects. My recommended sequence:

  1. Audit your tools for side effects first. Anything that charges, sends, or provisions needs an idempotency key before you cut over — this is the change that produces no errors and no warnings.
  2. Rewrite elicitation flows as MRTR retry loops. Budget real engineering time here; it’s an inversion, not a rename.
  3. Move application state to explicit handles backed by your existing datastore, and delete the session infrastructure only after legacy client traffic drains.
  4. Default to serverless plus a managed gateway unless you have a measured reason for containers.

The open question I’d put to your team: which of your current tools would duplicate a side effect if a stream dropped mid-call and the client retried? If you can’t answer that from your codebase today, that’s where your migration plan starts.