Persistent Memory for AI Agents: What It Takes to Make It Stick

John Rood··6 min read

An agent helps you settle an approach on Tuesday. On Wednesday, in a fresh session, it asks you why that approach was chosen.

The work from Tuesday is in the repo. The reasoning is not. That is the default behavior of every agent that starts cold, and it is the reason persistent memory has moved from nice-to-have to the first thing people evaluate when they run agents for real work.

If you want the broader picture first, what AI memory is covers the category across assistants, products, and agents. The rest of this post is the working view: what persistent memory actually requires, the places it breaks, and how to add it to an agent that does not have it yet.

The context window is working memory, and it ends

Everything an agent knows during a session lives in its context window: the system prompt, the conversation, tool results, and whatever it read from disk. When the session ends, that stack is gone. When the window fills, compaction rewrites it, keeping a summary and discarding the rest.

Compaction is the worst possible archivist. It runs under a hard token budget, at the moment of maximum pressure, and it keeps outcomes while dropping reasons. "We chose the queue-based design" survives. "Because the direct write path was hitting lock contention under load" does not, and that second sentence is the one you need three weeks later.

The repo records what changed. Nothing in a stock setup records why, or what was rejected, or which constraint was non-negotiable. Persistent memory is the name for closing that gap.

Four jobs, and none of them are optional

Persistent memory is a small system with four jobs. Skip any one and the others stop mattering.

  • Capture. Something writes down what happened, automatically, at the end of each turn: your request plus the final answer, with tool calls and intermediate noise filtered out. If capture depends on the model deciding to remember, it will fail silently the one time it matters.
  • Store. The record lives outside the context window, scoped to a person or a project, in a place that survives restarts, reinstallations, and model swaps.
  • Recall. When a new request arrives, relevant records come back ranked, semantic meaning first and recency close behind. On a typical vault that retrieval lands in 300 to 500 ms. The number matters because recall sits in front of every request you make.
  • Inject. What comes back is placed into the prompt before the model runs, inside a budget that respects the context limit. The agent should not have to ask for its own history.

Four functions, no magic. The difficulty is doing all four automatically and consistently, which is where most homegrown setups quietly give up.

Where implementations break

The failures are consistent enough to list.

Memory that depends on the model. A surprising number of memory features are really a tool the model may call. That works until the model is deep in a task and does not call it. Tool-call memory makes remembering something the user has to request, which defeats the point.

Memory that lives in one tool. If the only copy sits inside one product, the moment you open a second one you are back to re-explaining. A decision made in Claude Code should be available to Codex and your chat assistant without an export step. One vault, many surfaces.

Memory that is only a summary. Bullet-point session summaries feel tidy and lose the load-bearing details. The useful memory is full fidelity: what was asked, what was decided, what got rejected and why. Short is not the same as good.

Memory with no lifecycle. Facts go stale. Two memories contradict each other. If nobody can edit, delete, or expire a record, the vault becomes a liability, and the agent starts citing last month's truth with this month's confidence.

How to add it without building storage infrastructure

There are three practical insertion points, and they all end in the same vault.

For coding agents, hooks. Claude Code, Codex, OpenClaw, and DeepSeek Harness all expose lifecycle events, and the MemoryRouter packages use them. Recall runs before each prompt, capture runs after each completed turn, and both are deterministic rather than model-decided. Installation is one command per tool, and the hooks fail open: if the memory service is unreachable, the agent proceeds as if nothing happened.

For chat surfaces, MCP. The remote server at mcp.memoryrouter.ai connects ChatGPT, Claude, and Cowork over OAuth, so the same memory is available in the places you talk as well as the places you code.

For your own agents and products, the API. The endpoint is OpenAI-compatible, so the integration is usually a base_url change plus a Memory Key. /v1/memory/prepare returns ranked context before inference. /v1/memory/ingest stores the exchange after it. A Memory Key maps to exactly one vault, and no other key can read it.

Portability is the part worth pausing on: because every surface writes to one vault, AI memory stops being a per-tool feature and becomes a layer you own. The same mechanism that gives one agent continuity gives a fleet of them a shared base of context. The core concepts guide covers keys, vaults, and scoping in detail.

The two-session test

Do not trust a memory setup you have not tested across a session boundary. The test takes four minutes.

  1. In a connected session, state a disposable fact: "Store this: the demo exporter uses a 40-second flush window because CI throttles bursts."
  2. Close everything. Open a fresh session. Paste nothing. Ask: "What flush window does the demo exporter use, and why?"
  3. Run the same question on a profile with no memory connected. It should come back empty. The difference between the two answers is what you just added.

If the fact returns with its reason attached, all four jobs worked. If only a vague summary comes back, you found the summary-only failure. If nothing comes back, capture is the first place to look.

What persistent memory does not do

  • It does not replace project rules. Build commands and conventions belong in CLAUDE.md or AGENTS.md, where they load every time, unchanged. Memory is for what happened, not for what is always true.
  • It does not fix a bad prompt. It removes re-explaining, not ambiguity.
  • It does not decide what matters. Capture writes the turn. Ranking decides what returns. Both are mechanical.
  • It does not run itself. Some process has to call capture, which is why hooks and plugins exist.

The measure of a working setup is simple: you stop noticing that sessions start. The agent opens with your project loaded, and the first minute of a session goes to work.

Create your MemoryRouter account, connect one agent, and run the two-session test before real decisions go in.