In the span of about a year, the three largest AI assistants all shipped the same feature: persistent memory that survives the conversation. Claude Memory rolled out to Team and Enterprise plans in September 2025, reached Pro and Max in October 2025, and landed on the free tier on March 2, 2026. Gemini's version launched as "Personal context" on Gemini 2.5 Pro in August 2025, then got rebranded and expanded as Personal Intelligence on January 14, 2026. OpenAI's rebuilt system, internally called Dreaming, replaced the old saved-memories list and Reference Chat History with a single background process in June 2026 and reached Enterprise the following February, reporting 82.8% factual recall and 71.3% preference adherence in its own benchmarks. By the middle of 2026, an assistant that forgets you between sessions is the exception, not the default.
That rollout got covered as a product story: assistants that know your preferences, your projects, your history. What got almost no coverage is what it does to the bill. A single conversation's context resets when the session ends. Memory, by design, does not. It accumulates for as long as the account exists, and the naive way to implement it, inject everything the system remembers into every prompt, means the cost of a call stops tracking what the user asked and starts tracking how long the agent has been running.
What memory actually costs, entry by entry
Mem0's own benchmarking, published across two 2026 posts on token optimization for agent memory, put numbers on the growth curve for naive full-context injection: a 7-entry memory store adds about 146 prompt tokens to every call; 24 entries adds 594; 500 entries adds roughly 8,000. None of that is optional overhead the model can skip, it's the literal cost of reminding the model who you are and what it already knows, on every single request, whether or not that context is relevant to the one being asked.
Run that same math at production scale and the numbers stop looking like a benchmark table. Mem0 reports continuously running agents commonly reach 80,000 to 120,000-token contexts within two to three weeks of live operation, purely from memory accumulation, with API bills growing 5 to 10x past what teams initially budgeted for. A well-built retrieval architecture holds around 7,000 tokens per retrieval at that same scale against 25,000 to 100,000 or more for a full-context approach, with recall staying above 91%. The gap between those two numbers is not a tuning parameter. It's the difference between an agent that gets more expensive to run every week it stays alive and one that doesn't.
| Memory strategy | 24-entry store | What it does |
|---|---|---|
| Full dump (naive) | 594 tokens/call | Every entry, every call, regardless of relevance |
| Retrieval, top-10 | 293 tokens/call (-51%) | Semantic search, keep the 10 most relevant |
| Retrieval, top-5 | 166 tokens/call (-72%) | Same search, tighter cutoff |
The retrieval versions don't make the agent remember less. They make the agent stop paying, on every call, for entries that call didn't need.
TokenPilot: memory management as a routing decision, not a UX feature
A June 15, 2026 paper, TokenPilot: Cache-Efficient Context Management for LLM Agents (arXiv:2606.17016), part of the fifteen-author LightMem series, treats this as a systems problem rather than a product one. Its argument: as agents run in long-horizon sessions, context accumulation drives up inference cost on its own, independent of task difficulty, and the fix has to manage two things at once instead of one.
TokenPilot's design has two parts working together. Ingestion-Aware Compaction keeps the prompt prefix stable as new content arrives, so eviction doesn't quietly shred the exact byte-identical prefix a provider's prompt cache depends on to give you a discount. Lifecycle-Aware Eviction tracks how much each context segment is actually still being used and offloads the ones that aren't, rather than evicting on age alone. Tested on PinchBench and Claw-Eval, the two components together cut cost 61% and 56% in isolated-session mode, and 61% and 87% in continuous mode, the deployment shape every production memory feature in the previous section actually runs in, while holding task performance competitive with the uncompacted baseline.
The 87% figure in continuous mode is the one worth sitting with. That's the mode where memory keeps running across sessions indefinitely, exactly what Claude Memory, Gemini Personal Intelligence, and ChatGPT's Dreaming are all built to do by default now. The larger the memory store gets and the longer an agent stays alive, the more there is for a well-designed eviction policy to find and remove, which is the opposite of how most naive memory implementations behave.
A four-step framework for building this yourself
None of the techniques above require waiting for a published library. They compose into an ordinary pipeline:
- Budget the memory slice before you write the prompt. Don't let memory content grow to fill whatever's left of the context window. Mem0's token-budgeting example caps memory at 80 tokens, selects the 3 highest-relevance entries that fit, and drops the rest as overflow, landing at 149 prompt tokens against a 600-token full-dump baseline, a 75% cut, with the cap enforced regardless of how large the underlying store gets.
- Retrieve top-k, don't inject the store. The table above already shows the mechanism: semantic retrieval against the 5 or 10 most relevant entries instead of the whole history. This is the single highest-leverage change, and it's the one every retrieval-augmented memory system already does by default; the mistake is skipping it and dumping the raw store because it's simpler to ship.
- Evict on a decay curve, not FIFO. Deleting the oldest entries first assumes age correlates with irrelevance, which isn't reliably true, a fact stated three months ago is often more relevant than a preference set a year ago. Mem0's Ebbinghaus-curve eviction example kept the 9 of 24 entries still above a relevance-decay threshold and cut prompt tokens from 600 to 249, a 59% reduction, by scoring what's actually still useful instead of what's oldest.
- Keep the retrieved prefix stable across calls. This is TokenPilot's actual contribution on top of steps 1 through 3: an eviction policy that changes the prompt prefix on every call breaks a provider's prompt cache and pays the write-tax on every request, which can erase most of what the eviction just saved. Structure retrieval so the stable parts of the prompt, system instructions, tool schemas, the retrieval query format, stay byte-identical across calls, and let only the retrieved memory payload vary.
def build_memory_context(query, memory_store, token_budget=150, top_k=5):
# Retrieve first, never inject the raw store.
candidates = memory_store.search(query, limit=top_k * 2)
# Rank by relevance and recency decay, not insertion order.
ranked = sorted(candidates, key=lambda m: m.relevance_score(query) * m.decay_weight(), reverse=True)
selected, used_tokens = [], 0
for memory in ranked[:top_k]:
cost = memory.token_count()
if used_tokens + cost > token_budget:
break
selected.append(memory)
used_tokens += cost
# Fixed template keeps the prefix stable for prompt-cache reuse.
return MEMORY_PREFIX_TEMPLATE.format(entries=selected)
Where Nadir fits
Nadir doesn't decide what an agent remembers or how it evicts, that's the memory layer's job, and the four steps above are how you build one that doesn't compound. What Nadir does is score each request as it actually arrives, memory payload included, against a quality floor and current provider pricing, before it's sent. A memory store that's grown to 8,000 tokens over three weeks of live use doesn't get to silently push every call onto the same expensive model just because the prompt got longer than it was when the routing decision was last reviewed. The classifier re-evaluates each request. Complete non-streaming responses report actual cost in nadir_metadata.cost.total_cost_usd; when a benchmark is configured and priced, nadir_metadata.benchmark_comparison.savings_usd reports the comparison. If retrieval-based memory isn't built yet, routing is the layer that stops the naive version from being billed at frontier rates on each call while it gets built.
Related reading
- The Context Garbage Collector
- Multi-turn conversations bill your first message 20 times. Most teams have never calculated what that costs.
- An FSE 2026 paper cut agent input tokens up to 60% by deleting context the agent already read.
- Prompt caching's write tax
Sources: Zhang et al., ["TokenPilot: Cache-Efficient Context Management for LLM Agents"](https://arxiv.org/abs/2606.17016), arXiv:2606.17016, June 15, 2026. Mem0, ["The 2026 Token Optimization Playbook: Cut AI Agent Memory Costs 3-4X"](https://mem0.ai/blog/the-2026-token-optimization-playbook-cut-ai-agent-memory-costs-3%E2%80%934x), 2026. Mem0, ["6 Techniques to Cut AI Agent Memory Cost Beyond Basic Retrieval"](https://mem0.ai/blog/6-techniques-to-cut-ai-agent-memory-cost-beyond-basic-retrieval), 2026. Anthropic, Claude Memory rollout timeline (Team/Enterprise September 2025, Pro/Max October 2025, free tier March 2, 2026). Google, Gemini Personal Intelligence, rebranded and expanded January 14, 2026 from "Personal context" (August 2025). OpenAI, Dreaming memory system, rolled out to Plus/Pro June 2026, expanded to Enterprise February 2026.