• 8 min read

MCP Cacheable Tool Results: A Production Caching Playbook

tl;dr

Caching MCP tools/list results cuts agent input-token costs by 30–60% for fleets with repetitive workloads. Incorrectly marking permission-filtered catalogs as public creates cross-user authorization leaks, so default to private scope unless you can prove responses are identical across all callers.

Featured image for "MCP Cacheable Tool Results: A Production Caching Playbook"

That SCOUT deployment report shows why MCP cacheable tool results matter: discovery data can consume more context and money than the task it’s supposed to support.

The catch is that “cache it” isn’t one decision. You have to decide what can be shared, how long it stays valid, whether tool ordering remains stable, and whether caching saves input tokens or accidentally breaks authorization.

What are MCP cacheable tool results?

MCP cacheable tool results are discovery responses that carry freshness and scope metadata, allowing clients to reuse them instead of repeatedly requesting the same catalog. The MCP 2026-07-28 specification announcement introduced ttlMs and cacheScope on cacheable list results.

This landed alongside a stateless redesign. MCP removed the initialization handshake and Mcp-Session-Id header, while Multi Round-Trip Requests, or MRTR, replaced constantly open bidirectional streams with short, independent calls. The same release added deterministic ordering for list responses, giving clients a better chance of preserving upstream prompt caches.

It’s crucial to distinguish discovery metadata from execution data. The specification doesn’t make tools/call responses cacheable; tool calls themselves remain uncached. A cached tool list tells the client what exists. It doesn’t mean the result of invoking a tool can safely be replayed.

That boundary is sensible. A catalog may remain unchanged for minutes, while a tool call can create a payment, modify a record, or return user-specific data. The first is metadata. The second is a side effect.

How do ttlMs and cacheScope work?

ttlMs declares how long a list result remains fresh in milliseconds, much like HTTP’s Cache-Control: max-age. According to SEP-2549’s caching explanation, cacheScope has two defined values: public for results shared across users and private for results scoped to a user or authorization context.

Use the fields as a contract about variability, not merely a performance setting:

  • If every authorized user receives the identical catalog, public is defensible.
  • If roles, tenants, feature flags, or identity change the tool list, use private.
  • If you can’t prove the response is identical across callers, stay private.
  • If freshness depends on an external control plane, set a short TTL rather than pretending the catalog is static.

The dangerous mistake is treating a permission-filtered tools/list response as shared. Marking that response with a shared scope—including the public equivalent—can let one user’s cached tool set serve another, creating an authorization leak.

Private caching wastes less than an incident. You may reduce cache reuse, but you keep the authorization boundary attached to the cached object. That’s the right trade for tool catalogs that reflect role-based access.

Where does the cost savings come from?

The savings come from avoiding repeated discovery and stable tool-schema processing, not from caching tool execution results. For a reported fleet of 50 agents with 25 tools averaging 800 tokens per definition and 200 turns per day, repeated schema boilerplate costs range from $1,500 per month to $9,000 per month.

The same source reports provider prompt-cache discounts of 50–90% for cached input tokens, with cache lifetimes from 5 minutes to 1 hour. It also reports 30–60% reductions in overall input-token spending for agent workloads using gateway caching.

Here’s the scenario projection from those inputs:

These are scenario estimates, not a universal benchmark. Your cache-hit rate, schema size, provider pricing, and prefix stability determine the result.

The main caching choices work differently, so compare them before choosing a layer:

Caching optionPricing / cost evidenceMain featuresBest fit
MCP list-result cache—Reuses tools/list until ttlMs expires; supports public or private scopeStable remote tool catalogs
Provider prompt cacheReported 50–90% discountReuses stable prompt prefixes for a limited lifetimeStable tool ordering and repeated prefixes
Gateway token cacheModeled fleet ranges from $1,500 per month to $9,000 per month before optimizationNormalizes requests and reuses tokenized representationsLarge fleets with repetitive agent workloads
SCOUT searchable catalog—Selects relevant tool definitions rather than loading the full catalogCatalogs too large for direct injection

A protocol cache and a provider cache solve different problems. MCP metadata tells the client when it can reuse tools/list; provider caching discounts repeated prompt processing. Neither makes an uncacheable tool call safe to replay.

Can progressive discovery and prompt caching coexist?

Yes, but only if you treat tool-array stability as an architectural constraint. Progressive discovery can slash context consumption, while mutation of the tools array can invalidate the provider cache prefix.

The scale is substantial. MCP client documentation illustrates roughly 150,000 tokens loaded upfront versus about 2,000 tokens with progressive discovery. That’s an illustrative order-of-magnitude comparison, not a reproducible benchmark, but it shows why catalog governance belongs in production design.

Here’s the contradiction: adding or removing tool definitions mid-conversation invalidates prompt caches, and the resulting miss can cost more tokens than the definitions you removed. That warning comes directly from the same MCP client guidance.

A practical design keeps the frequently used prefix stable and moves dynamically selected definitions behind a breakpoint whenever the provider supports that approach. It also avoids re-sorting the same tools differently for equivalent requests. Finally, the design keeps authorization-filtered catalogs out of shared prompts and caches, and measures accuracy after discovery changes—not just token consumption.

The MCP client guidance discussed here suggests treating a catalog larger than 1–5% of the context window as a signal to switch to progressive discovery. That’s a useful trigger, not a law. A small catalog doesn’t need a retrieval layer; an enormous stable catalog may still justify caching without dynamic discovery.

For server selection, our guide to the best MCP servers for Claude Code emphasizes a small, non-colliding tool set. For integration architecture, MCP versus GraphQL hybrid designs covers the related choice between agentic discovery and deterministic direct calls.

What does stateless MCP change for cache design?

Stateless MCP removes protocol sessions, not application state. Requests can reach any server instance, but anything an application still needs—conversation history, approvals, task status, or idempotency controls—must live outside the protocol session.

The redesign makes infrastructure simpler. The official specification announcement says any request can land on any instance behind ordinary load balancing. The release also adds Mcp-Method and Mcp-Name headers, allowing gateways to route operations without parsing request bodies, as explained in Tigera’s gateway analysis.

That doesn’t make caching trivial. With a session ID gone, the cache key must capture everything that makes a list response meaningfully different. Protocol version, server identity, request parameters, and authorization context matter because changing any one of them can produce a different catalog.

MRTR also shifts long-running coordination into independent calls. That improves load-balancer compatibility, while application developers handle pause, retry, and resume behavior themselves. Our production failover architecture guide covers the broader consequence: moving a session out of the protocol doesn’t move business state out of the system.

Keep two concerns separate. Use stateless protocol behavior to remove session affinity. Use application-level storage, authorization, and idempotency mechanisms to preserve correctness. Conflating them creates a system that scales horizontally but replays work.

How do tool renames interact with cached catalogs?

A cached tool name becomes a versioned API contract. Renaming a tool and deploying the change isn’t atomic for clients that still hold the previous tools/list response.

The specification recommends names containing 1 to 128 case-sensitive ASCII letters, digits, underscores, hyphens, or dots, according to the MCP tool-schema guidance. Those permissive names are easy to change—and easy to break when prompts, allowlists, saved workflows, or cached catalogs still reference the old value.

More importantly, the cache-lifetime analysis treats a server’s advertised ttlMs as a deprecation commitment for its tool names. A long cache lifetime therefore extends the period during which a renamed tool can remain live in client state.

MCP’s separate formal deprecation policy establishes a twelve-month minimum window between formal deprecation and removal, with narrow exceptions. That improves ecosystem planning, but it doesn’t automatically retire a name already cached by a client.

Treat names as durable interfaces. Keep an old name answering through a compatibility path when clients may still hold it, publish the replacement with clear precedence, and shorten the catalog TTL when your deployment process needs faster convergence. A clean rename on the server can still be dirty in every client that cached the old list.

What decision framework should teams use?

Start with the narrowest cacheable surface: tools/list. Don’t begin with semantic response caching, broad gateway policies, or dynamically assembled tool sets. Classify every list result by whether it changes with time, tenant, identity, feature flags, or rollout state.

Then apply this sequence:

  1. Measure the catalog. Record tool count, serialized definition size, repeated discovery calls, and the percentage of context occupied before the user query.
  2. Classify variability. Mark identical cross-user results as public; use private whenever authorization or tenant context can alter the response.
  3. Choose a TTL from deployment behavior. Keep it short enough that normal tool changes propagate without manual cache eviction.
  4. Preserve prompt stability. Keep ordering deterministic and avoid adding or removing definitions unnecessarily during a conversation.
  5. Test negative paths. Attempt cross-user cache hits, stale-name calls, TTL expiration, catalog rollback, and rollout-time schema changes.
  6. Evaluate accuracy beside savings. A cheaper prompt that selects the wrong tool isn’t an optimization; it’s an incident waiting for a larger sample size.

If the catalog consumes too much context, add searchable discovery rather than continuing to advertise everything. If the catalog is small and stable, direct caching may be enough. If repeated prompt processing dominates, add provider or gateway caching—but preserve the prefix conditions that make those mechanisms work.

My default recommendation is to enable cache hints for tools/list only, default to private scope whenever authorization can change the catalog, preserve stable tool ordering, and expand to public caching only after automated tests prove that every eligible response is identical.