The Memory Tax

In the span of about a year, the three largest AI assistants all shipped the same feature: persistent memory that survives the conversation. Claude Memory reached the free tier on March 2, 2026. Gemini's Personal Intelligence expanded January 14, 2026. OpenAI's rebuilt Dreaming system rolled out to Plus and Pro in June 2026 and reached Enterprise that February. Covered as a product story, it's about an assistant that finally remembers you. Underneath it is a cost story nobody priced in: a single session's context resets when the conversation ends, but memory, by design, does not, it accumulates for as long as the account exists. Mem0's own benchmarking puts numbers on what that costs done naively: a 24-entry memory store adds 594 tokens to every single call, a 500-entry store adds roughly 8,000, and production agents commonly hit 80,000 to 120,000-token contexts within two to three weeks of continuous operation, bills growing 5 to 10x past what teams budgeted. A June 2026 paper, TokenPilot (arXiv:2606.17016), shows a cache-aware retrieval-and-eviction design recovers 61-87% of that cost in the exact always-on deployment mode every major assistant now runs by default. Here's the math, the paper, and a four-step framework for building memory that doesn't get more expensive every week it stays alive.

Published 2026-08-16 by Dor Amir on the Nadir blog.

Filed under Agents.

In the span of about a year, the three largest AI assistants all shipped the same feature: persistent memory that survives the conversation. Claude Memory rolled out to Team and Enterprise plans in September 2025, reached Pro and Max in October 2025, and landed on the free tier on March 2, 2026. Gemini's version launched as "Personal context" on Gemini 2.5 Pro in August 2025, then got rebranded and expanded as Personal Intelligence on January 14, 2026. OpenAI's rebuilt system, internally called Dreaming, replaced the old saved-memories list and Reference Chat History with a single background process in June 2026 and reached Enterprise the following February, reporting 82.8% factual recall and 71.3% preference adherence in its own benchmarks. By the middle of 2026, an assistant that forgets you between sessions is the exception, not the default.

That rollout got covered as a product story: assistants that know your preferences, your projects, your history. What got almost no coverage is what it does to the bill. A single conversation's context resets when the session ends. Memory, by design, does not. It accumulates for as long as the account exists, and the naive way to implement it, inject everything the system remembers into every prompt, means the cost of a call stops tracking what the user asked and starts tracking how long the agent has been running.

What memory actually costs, entry by entry

Mem0's own benchmarking, published across two 2026 posts on token optimization for agent memory, put numbers on the growth curve for naive full-context injection: a 7-entry memory store adds about 146 prompt tokens to every call; 24 entries adds 594; 500 entries adds roughly 8,000. None of that is optional overhead the model can skip, it's the literal cost of reminding the model who you are and what it already knows, on every single request, whether or not that context is relevant to the one being asked.

Memory bills for existing, not for asking. Left: naive full-context injection grows from 146 to roughly 8,000 tokens per call as a memory store goes from 7 to 500 entries. Right: the same 24-entry store costs 594 tokens injected whole, 293 tokens retrieved top-10, or 166 tokens retrieved top-5, a 51-72% cut with no change to what the agent remembers.
Memory bills for existing, not for asking. Left: naive full-context injection grows from 146 to roughly 8,000 tokens per call as a memory store goes from 7 to 500 entries. Right: the same 24-entry store costs 594 tokens injected whole, 293 tokens retrieved top-10, or 166 tokens retrieved top-5, a 51-72% cut with no change to what the agent remembers.

Run that same math at production scale and the numbers stop looking like a benchmark table. Mem0 reports continuously running agents commonly reach 80,000 to 120,000-token contexts within two to three weeks of live operation, purely from memory accumulation, with API bills growing 5 to 10x past what teams initially budgeted for. A well-built retrieval architecture holds around 7,000 tokens per retrieval at that same scale against 25,000 to 100,000 or more for a full-context approach, with recall staying above 91%. The gap between those two numbers is not a tuning parameter. It's the difference between an agent that gets more expensive to run every week it stays alive and one that doesn't.

Memory strategy24-entry storeWhat it does
Full dump (naive)594 tokens/callEvery entry, every call, regardless of relevance
Retrieval, top-10293 tokens/call (-51%)Semantic search, keep the 10 most relevant
Retrieval, top-5166 tokens/call (-72%)Same search, tighter cutoff

The retrieval versions don't make the agent remember less. They make the agent stop paying, on every call, for entries that call didn't need.

TokenPilot: memory management as a routing decision, not a UX feature

A June 15, 2026 paper, TokenPilot: Cache-Efficient Context Management for LLM Agents (arXiv:2606.17016), part of the fifteen-author LightMem series, treats this as a systems problem rather than a product one. Its argument: as agents run in long-horizon sessions, context accumulation drives up inference cost on its own, independent of task difficulty, and the fix has to manage two things at once instead of one.

TokenPilot's design has two parts working together. Ingestion-Aware Compaction keeps the prompt prefix stable as new content arrives, so eviction doesn't quietly shred the exact byte-identical prefix a provider's prompt cache depends on to give you a discount. Lifecycle-Aware Eviction tracks how much each context segment is actually still being used and offloads the ones that aren't, rather than evicting on age alone. Tested on PinchBench and Claw-Eval, the two components together cut cost 61% and 56% in isolated-session mode, and 61% and 87% in continuous mode, the deployment shape every production memory feature in the previous section actually runs in, while holding task performance competitive with the uncompacted baseline.

The 87% figure in continuous mode is the one worth sitting with. That's the mode where memory keeps running across sessions indefinitely, exactly what Claude Memory, Gemini Personal Intelligence, and ChatGPT's Dreaming are all built to do by default now. The larger the memory store gets and the longer an agent stays alive, the more there is for a well-designed eviction policy to find and remove, which is the opposite of how most naive memory implementations behave.

A four-step framework for building this yourself

None of the techniques above require waiting for a published library. They compose into an ordinary pipeline:

  1. Budget the memory slice before you write the prompt. Don't let memory content grow to fill whatever's left of the context window. Mem0's token-budgeting example caps memory at 80 tokens, selects the 3 highest-relevance entries that fit, and drops the rest as overflow, landing at 149 prompt tokens against a 600-token full-dump baseline, a 75% cut, with the cap enforced regardless of how large the underlying store gets.
  1. Retrieve top-k, don't inject the store. The table above already shows the mechanism: semantic retrieval against the 5 or 10 most relevant entries instead of the whole history. This is the single highest-leverage change, and it's the one every retrieval-augmented memory system already does by default; the mistake is skipping it and dumping the raw store because it's simpler to ship.
  1. Evict on a decay curve, not FIFO. Deleting the oldest entries first assumes age correlates with irrelevance, which isn't reliably true, a fact stated three months ago is often more relevant than a preference set a year ago. Mem0's Ebbinghaus-curve eviction example kept the 9 of 24 entries still above a relevance-decay threshold and cut prompt tokens from 600 to 249, a 59% reduction, by scoring what's actually still useful instead of what's oldest.
  1. Keep the retrieved prefix stable across calls. This is TokenPilot's actual contribution on top of steps 1 through 3: an eviction policy that changes the prompt prefix on every call breaks a provider's prompt cache and pays the write-tax on every request, which can erase most of what the eviction just saved. Structure retrieval so the stable parts of the prompt, system instructions, tool schemas, the retrieval query format, stay byte-identical across calls, and let only the retrieved memory payload vary.
def build_memory_context(query, memory_store, token_budget=150, top_k=5):
    # Retrieve first, never inject the raw store.
    candidates = memory_store.search(query, limit=top_k * 2)

    # Rank by relevance and recency decay, not insertion order.
    ranked = sorted(candidates, key=lambda m: m.relevance_score(query) * m.decay_weight(), reverse=True)

    selected, used_tokens = [], 0
    for memory in ranked[:top_k]:
        cost = memory.token_count()
        if used_tokens + cost > token_budget:
            break
        selected.append(memory)
        used_tokens += cost

    # Fixed template keeps the prefix stable for prompt-cache reuse.
    return MEMORY_PREFIX_TEMPLATE.format(entries=selected)

Where Nadir fits

Nadir doesn't decide what an agent remembers or how it evicts, that's the memory layer's job, and the four steps above are how you build one that doesn't compound. What Nadir does is score each request as it actually arrives, memory payload included, against a quality floor and current provider pricing, before it's sent. A memory store that's grown to 8,000 tokens over three weeks of live use doesn't get to silently push every call onto the same expensive model just because the prompt got longer than it was when the routing decision was last reviewed. The classifier re-evaluates each request. Complete non-streaming responses report actual cost in nadir_metadata.cost.total_cost_usd; when a benchmark is configured and priced, nadir_metadata.benchmark_comparison.savings_usd reports the comparison. If retrieval-based memory isn't built yet, routing is the layer that stops the naive version from being billed at frontier rates on each call while it gets built.

Related reading


Sources: Zhang et al., ["TokenPilot: Cache-Efficient Context Management for LLM Agents"](https://arxiv.org/abs/2606.17016), arXiv:2606.17016, June 15, 2026. Mem0, ["The 2026 Token Optimization Playbook: Cut AI Agent Memory Costs 3-4X"](https://mem0.ai/blog/the-2026-token-optimization-playbook-cut-ai-agent-memory-costs-3%E2%80%934x), 2026. Mem0, ["6 Techniques to Cut AI Agent Memory Cost Beyond Basic Retrieval"](https://mem0.ai/blog/6-techniques-to-cut-ai-agent-memory-cost-beyond-basic-retrieval), 2026. Anthropic, Claude Memory rollout timeline (Team/Enterprise September 2025, Pro/Max October 2025, free tier March 2, 2026). Google, Gemini Personal Intelligence, rebranded and expanded January 14, 2026 from "Personal context" (August 2025). OpenAI, Dreaming memory system, rolled out to Plus/Pro June 2026, expanded to Enterprise February 2026.

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.