• 9 min read

AI Agent Memory Tiers Explained: What Actually Matters

tl;dr

Postgres + pgvector is a strictly better default than commercial agent memory stores for most 2026 enterprise use cases. At 10,000 monthly active users, the baseline costs $163 to $332 monthly, 2-6x less than managed options like Zep or Letta, with no independent confirmation of better retrieval from paid tiers.

Featured image for "AI Agent Memory Tiers Explained: What Actually Matters"

At a defined workload of 10,000 monthly active users, a baseline Postgres + pgvector memory tier runs ~$163 to $332 a month, while Zep’s managed store runs $375 to $750 and Letta lands around $1,020 — and no independent party has confirmed the commercial options retrieve anything better. That gap is the whole story of AI agent memory tiers in 2026. The tiers themselves are real engineering, and you need them. What you probably don’t need is an expensive managed product to implement them.

Here’s the thing most vendor landing pages skip: memory tiers are an architecture, not a product. Once you understand the layers and what each one costs, you’ll find that most of the “memory platform” market is managed hosting of a pattern any competent team can build on infrastructure they already run. If you want the full conceptual grounding first, our guide to how agents remember context across tasks covers the mechanics; this post is about the tiers, the tradeoffs, and where the money actually goes.

What Are the Memory Tiers in an AI Agent?

The short answer: agents need at least four distinct memory layers, and conflating them is the root cause of most “my agent forgot” complaints. The CoALA taxonomy — the framework most of the field borrowed from cognitive science — names four memory types: working, episodic, semantic, and procedural. That split traces back to psychology’s original distinction between episodic memory (what happened, when), semantic memory (general facts), and procedural memory (how to do things), which researchers ported to LLM agents almost wholesale.

Storage-wise, one widely cited breakdown of agent memory systems describes four types: in-context memory (short-term, up to 1M tokens on the largest windows), external vector or key-value stores for long-term recall, parametric memory baked into model weights, and episodic logs of past interactions. Translate that into production infrastructure and you get something more concrete — production agents need four distinct layers: short-term context windows, mid-term scratchpads that survive a single run, long-term vector stores for facts and decisions, and structured state for things the agent must not forget.

Why does the distinction matter so much? Because each layer has a different write rate, a different retrieval shape, and a different cost profile. Working memory is free but ephemeral — it dies with the session. A scratchpad is cheap but unstructured. A vector store is persistent but returns semantically similar text, including stale text. Structured state is precise but rigid. An agent that treats these as one undifferentiated bucket will retrieve the wrong thing at the wrong time, and the failure is silent: there’s no error message when the model starts ignoring a constraint buried 80K tokens back.

Why Can’t a Bigger Context Window Replace Memory Tiers?

Because the window degrades well before it fills, and the degradation is invisible. Frontier models with 200K-token windows show measurable reasoning degradation above roughly 130K tokens — well below their advertised limits — and the threshold drops further on complex multi-step tasks. Worse, coding agents accumulate 80,000 to 150,000 tokens within roughly thirty-five minutes of ordinary work, so a long session can cross the degradation line before lunch.

Cost compounds the problem. Every token in the window is paid for on every turn, so a 200K-token context sent repeatedly costs two hundred times what a 1K-token context costs in input alone. Stuffing full history isn’t just architecturally lazy — it’s a cash burn spiral at any real volume. And the failure mode isn’t a clean cliff: it’s context drift, where earlier decisions and constraints gradually lose influence as new content piles up. The agent doesn’t crash. It just gets quietly worse.

This is why the 2025–2026 production consensus converged on tiered memory architectures inspired by operating-system memory hierarchies, with explicit token budgets and hybrid compression-plus-retrieval strategies. The analogy is genuinely useful: an OS doesn’t keep everything in RAM because RAM is expensive and finite. It keeps hot data close and pages cold data out to disk, with a policy deciding what moves. Agent memory is the same discipline applied to tokens — the policy that decides what gets into the window each turn, not just the stuff outside it.

How Do the Tiers Map to Real Infrastructure?

Three complementary strategies, and you’ll typically want all of them working together. Per research on token budgets in persistent agents, the production landscape has settled on tiered memory architectures that keep hot state cheap and cold state out-of-band, prompt caching that amortizes stable system content at roughly a 90% discount, and selective compaction that preserves reasoning continuity without dragging full history along.

In practice, the mapping looks like this:

  • Working memory → the context window itself, plus prompt caching for the stable prefix (system prompt, tool definitions).
  • Episodic memory → session transcripts or structured event logs, queryable by time and outcome.
  • Semantic memory → a vector store (Postgres + pgvector, or a dedicated vector database) holding extracted facts and decisions.
  • Procedural memory → skills, playbooks, or tool definitions the agent reuses across tasks.

One wrinkle worth knowing: the free baseline is better than most teams realize. Anthropic now ships a client-side Memory tool (identifier memory_20250818) where Claude reads and writes files in a /memories directory backed by your own storage. That’s a directory of markdown files — which means any paid memory platform has to beat, in effect, a well-organized folder.

The honest caveat is that no single architecture dominates. A systematic evaluation of 12 memory systems across five benchmark workloads found that effectiveness depends heavily on how well the memory structure aligns with the workload’s bottleneck. Relational questions (“who approved the last renewal”) favor graph-shaped storage. Temporal questions favor validity windows on facts. Plain recall favors vectors. Match the tier to the query pattern, not to the vendor’s demo.

What Do Commercial Memory Stores Actually Sell?

Mostly the same thing, it turns out. All four major managed memory layer launches in mid-2026 — Weaviate Engram, Cloudflare Agent Memory, Microsoft Foundry Memory, and Mem0 — converged on the identical architecture: an asynchronous extract-and-store write path plus a retrieval-first recall path, deliberately refusing to stuff the context window. When four companies ship the same design inside seven weeks, that’s not innovation. That’s a commoditized pattern getting wrapped in managed hosting.

The genuine differences are philosophical, not functional. Zep stores memory as a temporal knowledge graph that tracks when facts were true, so a superseded fact is invalidated rather than overwritten. Mem0 extracts facts from conversations and hands them back at prompt time. Letta is an agent runtime where the agent itself edits its own memory using tools. Those are three different answers to the same question — but the core extract-and-retrieve machinery underneath is the same everywhere.

Then there’s the benchmark problem. Vendor-reported scores do not survive independent harnesses: Mem0’s claimed 94.4% on LongMemEval reproduces at 57.5%, and its 92.5 LoCoMo score reproduces at 66.9%. We’ve covered the benchmark inflation problem in agent memory systems in depth, but the short version is this: the retrieval-quality claims justifying premium pricing are the one thing nobody outside the vendors has confirmed.

How Much Do Memory Tiers Cost at Production Scale?

Here’s the comparison at a concrete workload — 10,000 monthly active users, 20 assistant turns each per month, roughly 50 retained memories per user, with all prices checked on the vendors’ own pricing pages in September 2026:

Postgres + pgvectorZep CloudMem0 PlatformLetta Cloud
Memory modelYour rows plus an embeddingBitemporal knowledge graph of entities and edgesLLM-extracted facts, optionally graph-linkedStateful agent memory blocks and files
Cost at this workload~$163–$332/mo all-in$375–$750/mo$249/mo Pro tier, 4x over its retrieval quota~$1,020/mo before LLM tokens
Self-host pathIt’s your databaseGraphiti only; bring your own graph DBApache 2.0, minus platform optimizationsApache 2.0, self-hostable server
Right to erasureDELETE ... WHERE user_id = $1One API call deletes user, threads, graphDelete APIs per memory and per userPer-agent; cascade undocumented

The pattern across that table is consistent: commercial stores run roughly 2–6x the baseline for standard workloads. And the pricing models create their own traps. Usage-based metering — per operation, per stored memory, per ingested byte — means your bill grows with agent adoption rather than with headcount, which looks cheap in a pilot and gets expensive in production.

Google offers the cautionary tale. Memory Bank bills storage at $0.30 per GiB-month on total data stored including revisions, and parent memories carry no default TTL — so four months of accumulated agent state landed on the first metered invoice, with the charge depending on configuration decisions made months before billing switched on. When your October invoice reflects an April config choice rather than September traffic, that’s not a pricing model. That’s a liability. The same lock-in dynamics show up in agent execution engines, where memory-layer compatibility drives switching costs — worth reading before you commit state to any managed runtime.

When Should You Buy a Memory Store Instead of Building Tiers Yourself?

When you can name the specific capability you couldn’t reasonably write yourself — and not before. No independent party has confirmed that a commercial memory store retrieves better than baseline Postgres + pgvector, so the retrieval-quality argument for buying is currently unsupported. What remains is the capability argument, and it’s real but narrow. Zep’s bitemporal fact invalidation, Letta’s self-editing memory blocks, Mem0’s entity deduplication — each requires genuine custom engineering to replicate. If temporal fact invalidation is core to your product, that’s a legitimate reason to buy. If you just need your agent to remember user preferences across sessions, it isn’t.

The question that sorts the field fastest isn’t a benchmark — it’s the boundary question: how much of your application should the memory system own? A memory layer you bolt on (Mem0) owns little. An agent runtime (Letta) owns nearly everything. A temporal graph (Zep) owns your fact model. Answer that first, and most of the comparison table collapses to one or two rows.

Here’s the decision framework I’d actually use:

  1. Start on Postgres + pgvector. Build the four tiers — context window, scratchpad, vector store, structured state — on infrastructure you already operate. Costs are predictable, data control is total, and erasure is a SQL statement.
  2. Add prompt caching and compaction before adding vendors. These two strategies cut token costs dramatically and cost nothing but engineering time.
  3. Buy only on a named, unmet capability. Write down the specific feature — temporal invalidation, self-editing memory, graph-aware retrieval — that you cannot reasonably build. If you can’t fill in that sentence, you’re buying hosting, not technology.
  4. Price the lock-in before the pilot. Check the self-host path, the erasure API, and what happens to your bill when usage grows 10x.

My honest read: for the vast majority of enterprise agent use cases in 2026, Postgres + pgvector is a strictly better default than any dedicated commercial memory store — what I’ve come to think of as the memory-store-last pattern. Equivalent retrieval at a fraction of the cost, full data sovereignty, pricing you can forecast. The open question worth watching: if independent benchmarks ever do confirm a retrieval advantage for the commercial stores, does the 2–6x premium become justified for everyone, or only for the workloads whose query patterns actually stress the difference? Until someone runs that study, the burden of proof sits squarely on the vendors.