On this page
MCP Multi-Round-Trip Requests Explained
tl;dr
MCP's new multi round-trip request pattern replaces deprecated in-flight server requests, fixing the stateless protocol core but making token cost the primary constraint. Per-call costs vary up to 10x across models, and multi-server schema bloat can push monthly costs to $717 at 200 daily requests.
Close to half a billion SDK downloads a month, per the 2026-07-28 specification announcement — that’s the scale of the install base that just absorbed the largest MCP revision since launch. The headline change is the stateless protocol core, but the piece that will touch your application code most directly is Multi Round-Trip Requests, the mechanism that lets a server ask a client for input mid-task without holding a stream open. If you build MCP servers or clients, this is the feature that decides whether your interactive tools survive the migration window.
Here’s the short version: the old way of doing server-to-client requests is deprecated, the new way is a request/response loop with an explicit state token, and the tradeoffs are about token economics as much as protocol hygiene. What I call the MCP Cost Shift pattern is playing out in real deployments right now — the spec fixed the infrastructure bottleneck and immediately made recurring token cost the constraint everyone talks about.
What Are Multi Round-Trip Requests in MCP?
They’re the replacement for in-flight server-to-client requests. Under the old protocol, a server processing your tool call could pause and fire a request back at the client — sampling/createMessage, roots/list, or elicitation/create — over a bidirectional stream that had to stay open the whole time. The 2026-07-28 spec replaces that mechanism under SEP-2322, because a server that initiates requests mid-flight is a server that needs session state, and session state is exactly what the new protocol removed.
The new flow works like this: instead of issuing an in-flight request, a server answers with a result whose resultType is "input_required", carrying an inputRequests map and an opaque requestState token. The client fulfills those embedded requests and re-issues the original request with inputResponses and the echoed requestState. The server picks up where it left off, using the state token to reconstruct what it was doing.
Think of it like a support ticket that gets bounced back to you for missing information. You don’t stay on hold with an open phone line — you get a case number, you gather the documents, and you resubmit with the case number attached. Any support agent can pick up your resubmission, because the ticket carries its own context.
There’s a guardrail: round trips are capped at input_required_max_rounds, which defaults to 10, matching the TypeScript and Python SDKs. A server can’t trap a client in an infinite clarification loop — after ten bounces, the client gives up.
Why Did the Old In-Flight Model Have to Go?
Because it was structurally incompatible with the rest of the new spec. The initialize/initialized handshake is removed under SEP-2575, and the Mcp-Session-Id header along with the protocol-level session are gone under SEP-2567. Protocol version, client info, and client capabilities that used to be exchanged once at connection time now travel in _meta on every request, making each request fully self-describing. A server that can answer any request on any instance can’t also demand a live channel back to a specific client — those two designs don’t compose.
The supporting cast matters too. A new server/discover method lets clients fetch server capabilities, supported protocol versions, and identity when they need them up front. Method and tool names now travel in Mcp-Method and Mcp-Name HTTP headers, so gateways can route and authorize on headers without parsing JSON-RPC bodies. We covered the mechanics of the stateless shift in detail in How MCP Actually Works, including where validation responsibility moves to application code.
The part that should focus your planning: HTTP+SSE transport, Roots, Sampling, and Logging are deprecated with a minimum twelve-month migration window. Sampling is one of the three request kinds the round-trip pattern replaces, so if your server leans on in-flight sampling today, you’re on that clock. Teams that deferred the migration math last quarter should read our architecture breakdown of the spec changes — the migration cost for remote deployments is not trivial, and the deprecation window is a deadline, not a suggestion.
How Do Clients Drive input_required Results?
Two ways, and the difference is where I’d focus your code review. The Ruby client drives these results automatically: once a handler is registered through on_elicitation, on_sampling, or on_roots, the call_tool, get_prompt, and read_resource methods fulfill the embedded requests and re-issue the original request with inputResponses plus the echoed requestState. You register a handler, declare the matching capabilities on connect — a server embeds only the request kinds the client declared — and the loop disappears behind a single method call.
Without a matching handler, the client raises MCP::Client::InputRequiredError instead of returning the result as final, and the error exposes input_requests, request_state, and the raw result so you can drive the exchange manually. That manual path is the escape hatch you want for agent runtimes where a human or a model decides how to answer an elicitation, rather than a canned form handler.
Backward compatibility is handled sensibly. Servers on legacy protocol versions never send resultType, so existing client behavior is unchanged, and the Ruby SDK ships MCP::ResultType::COMPLETE and MCP::ResultType::INPUT_REQUIRED for forward compatibility. On the server side of the ecosystem, the TypeScript, Python, Go, and C# SDKs are updated to match the spec, with the C# SDK v2.0 defaulting HttpServerTransportOptions.Stateless to true while keeping v1 code compiling and running.
What Does the Round-Trip Model Cost You in Tokens?
Here’s where the analysis gets interesting, and where the ecosystem’s loudest claims deserve scrutiny. Tool definitions are re-billed as input tokens on every API call whether or not tools are used — a standing cost, with typical multi-server setups consuming 41,000–72,000 tokens of tool schemas before the first prompt. One developer’s measurement of nine connected servers landed at about 41,000 tokens of standing context, and the same measurement found merging two duplicate custom servers cut schema tokens roughly in half, from ~41k to ~21k, with no functional loss.
The per-call math is model-dependent in a way most teams miss. The real GitHub MCP server’s 26 tools bill $0.0302 per call on Claude Opus 4.8 versus $0.0029 on Gemini 3.6 Flash — a 10x spread driven by the model, not the protocol. That’s why the widely cited benchmarks claiming MCP burns 4–32x more tokens than equivalent CLI calls feel outdated for this spec: a 2026 measurement of an identical social media posting job found MCP used 22% fewer input tokens than REST, with total costs within 4%, and the remaining variance traced to one flaky REST run rather than protocol overhead.
The spec also gives you a caching lever the old protocol lacked. List responses for tools, prompts, and resources now carry cache hints (ttlMs) and a deterministic order, so clients can cache tool catalogs and keep upstream prompt caches stable across reconnects. That matters because, per the CalculatorAI analysis, a five-server setup with 58 tools runs $0.1178 per request — $717 a month at 200 requests a day, against $47 a month with no MCP servers at all. The tools cost fifteen times what the actual work costs.
Here’s the measured comparison from that social media posting job, run five times per path:
| Path | Input tokens (median) | Cost per run (median) | Turns | Wall time | Best for |
|---|---|---|---|---|---|
| Hosted MCP server | 143,031 | $0.1020 | 8 every run | 19s | Predictable, agent-driven jobs |
| REST via Bash + curl | 182,782 | $0.0981 | 4 to 9 | 26s | Scripts you control end to end |
| Plain script, no model | 0 | $0.0000 | 0 | 1.7s | Fully deterministic workflows |
Read the spread before the medians. The MCP path’s input tokens varied by 321 across five runs; the REST path varied by 289,497, because one run went sideways. The round-trip pattern’s determinism — a fixed number of bounces, each carrying explicit state — is a cost-stability feature, not just a protocol nicety.
Where Does the Round-Trip Pattern Get Fragile?
Three places, and they’re all worth engineering around. First, consolidation has an accuracy cliff. Merging servers to deduplicate schemas cuts token bloat, but Anthropic’s own documentation and WorkOS’s analysis show tool selection accuracy drops once available tools exceed 30–50 per server. Uber’s engineering team found third-party vendors bundle 34–49 tools per server — one workspace suite ships 49 tools requiring roughly 22K tokens of schema — so consolidation that pushes you over the threshold degrades agent correctness even as it trims the bill. You’re trading a visible cost for an invisible error rate.
Second, prompt caching is powerful but brittle. Explicit caching cuts carried schema token costs to roughly a tenth of base input price, but the same developer measurement that produced the 41k baseline caught a single flaky server restart breaking the prompt cache prefix, causing input tokens to jump by roughly the size of the full context window for one request — costing more than an hour of stable work. The ttlMs hints help stabilize the prefix across reconnects, but a server that reorders its tool list on restart still invalidates everything downstream.
Third, small toolsets can’t benefit at all. Sub-floor toolsets — too small to meet provider cache minimums — cannot cache, so a two-tool server pays full input price on every call regardless of how carefully you configure hints. If your server is small, the round-trip mechanism is cheap; the schema overhead is what isn’t.
How Do Gateways and Platforms Handle the New Spec?
Mostly by making it opt-in, which is the right call. AgentCore Gateway supports the 2026-07-28 spec via UpdateGateway with a configuration field that advertises supported protocol versions — and adding 2026-07-28 alongside 2025-11-25 does not change behavior for clients requesting the older version. Nothing breaks the day you flip it on; only clients that explicitly request the new version get the new behavior.
That opt-in design matters more than it looks, because the governance layer is converging on the protocol fast. ServiceNow, Rubrik, and Microsoft all shipped policy enforcement through MCP in a single week — gateway-level control of which servers an agent can discover and which methods it can invoke. The Mcp-Method and Mcp-Name headers make that enforcement cheap, since a gateway can authorize without deserializing the body. If you’re thinking about where enforcement should live in your stack, our security risks breakdown covers why the protocol layer is the right surface — and why shipping without it was the original sin of early MCP deployments.
One gap to watch: the stateless shift also moved logging and monitoring responsibility onto implementers, and most native server logs fail enterprise auditability requirements. A round-trip exchange that bounces between client and server across stateless requests is harder to reconstruct after the fact than a session-scoped log — plan your observability before you plan your migration.
Should You Migrate to the Round-Trip Pattern Now?
My recommendation: yes, but in a specific order. Register your handlers first — on_elicitation, on_sampling, on_roots or their equivalents in your SDK — so the automatic driving path works before any server you depend on starts emitting input_required results. Then audit your tool counts against the 30–50 accuracy threshold, because migration is the natural moment to consolidate, and consolidation is where teams accidentally trade correctness for savings. Finally, instrument your cache: watch cache_read_input_tokens per request, because a restart-driven prefix break is the single most expensive failure mode in the new model, and it’s invisible until the invoice arrives.
The open question I’d leave you with: the spec solved session infrastructure and gave you cache hints, but schema bloat is still your problem — there’s no protocol-level deduplication or lazy tool loading yet. Until the maintainers treat schema compression as a first-class concern, every team is hand-rolling the same optimizations. If your tool catalog is the expensive part of every request, the round-trip pattern is the cheap part. Budget accordingly.
Recommended Reading
-
MCP vs GraphQL
MCP enables runtime tool discovery for AI agents but carries a steep token tax that can cost 19-40x more than direct API calls for common workflows. Hybrid architectures pairing MCP for agentic tasks with GraphQL or REST for deterministic integrations balance functionality and cost for most teams.
-
Build Multi-Tenant SaaS with AI: 2026 Guide
Totalum is the only AI app builder with a public API and MCP server for embedding. Multi-agent architectures outperform single-agent setups by 2.4x on complex tasks, proving orchestration beats solo chatbots.
-
AI Gateway Architecture Explained: The 2026 Control Plane
AI agents trigger dozens of model calls per request, making gateways mandatory control planes. Sub-millisecond overhead, caching, and unified governance for LLM, MCP, and A2A traffic define production AI infrastructure. Open-source AI-native gateways win for performance and compliance.