What Is an AI Agent Memory Layer, and How Do You Add One?

John Rood··6 min read

Ask five people what an "AI agent memory layer" is and you will get five answers: a vector database, a summary table, a note file, a product feature, or a whole platform. The term is doing a lot of work, so it is worth pulling apart before you pick one, because the wrong mental model leads to the wrong build.

If you want the category view first, what AI memory is covers how this layer fits across assistants, products, and agents. This post is the infrastructure view: what a memory layer is, what it is not, and the three ways to add one to a system that already runs agents.

The definition that holds up

A memory layer is the component between your agent and the model that turns what happened into stored context, and turns stored context back into the prompt.

Mechanically, it interposes in two places. Before inference, it retrieves relevant records and injects them. After inference, it captures the exchange. Around both, it manages the boring parts that decide whether the whole thing is usable: who owns a memory, how long it lives, how it gets edited, and what happens when two of them disagree.

That is the whole shape. Everything else is implementation quality.

What it is not

Not a vector database. A vector database is a storage primitive. It gives you similarity search and leaves you the responsibilities that actually make memory work: deciding what to write, when to write it, how to rank results that conflict, and how to guarantee that one user's memories never surface in another user's request. Those decisions are the layer. The index is where the bytes land.

Not retrieval over your documents. RAG answers "what do the sources say" from a corpus you author and update deliberately. Memory answers "what did we decide, prefer, and try" from work that happens as a side effect of using the agent. Both are useful. They are different jobs with different write patterns, and they deserve different systems. (If you want that comparison in depth, it is its own post: vector memory vs RAG for agents.)

Not the built-in memory of a model. Vendor memory is real and often good, and it dies at the vendor boundary. It cannot follow a decision from Claude Code into Codex, and it cannot be exported, audited, or moved when you change providers.

Where it sits in the stack

your agent or app
      |
      v
memory layer            capture  ->  store  ->  recall  ->  inject
      |                                    |
      v                                    v
model provider                        vault (per user, per project)

Two properties make this placement meaningful. First, the layer is provider-agnostic: swapping the model underneath does not touch the memory, and the same vault keeps serving requests through a different provider. Second, the credential is the isolation boundary. A Memory Key authenticates a request and identifies exactly one vault. You are not trusting a caller to pass the right user id.

The requirements checklist

When you evaluate a memory layer, or build one, the requirements are consistent:

  • Isolation enforced by key. One key, one vault, no cross-reads. This is the requirement that is expensive to retrofit.
  • A real lifecycle. Edit a memory in place, delete specific ids, clear a vault, export in bulk. Deletion must propagate to retrieval, or your compliance story is fiction.
  • Ranking with recency. Semantic similarity alone will serve a decision you reversed three weeks ago. Meaning first, recency close behind, conflicts resolved rather than duplicated.
  • A latency budget. Recall runs in front of every request, so it adds to every reply. For a typical vault, expect 300 to 500 ms. Ask what holds at ten times the size.
  • Fails open. If the memory service is slow or down, your agent must proceed as if memory were not configured. Memory should never be a gate on the agent running at all.
  • Observability. You should be able to open a dashboard and read exactly what got stored, then delete it if it is wrong. Memory you cannot inspect is memory you cannot trust.

Three ways to add one

The insertion point depends on what you already run.

MCP server. mcp.memoryrouter.ai is a remote server with OAuth, and any MCP-compatible client can attach it: ChatGPT, Claude, Cowork, and others. You get explicit memory tools in the toolbelt and the model reaches for them when relevant. This is the fastest path for chat surfaces and for clients whose only extension point is MCP.

Lifecycle hooks. Coding agents expose session events, and the MemoryRouter plugins use them: recall before each prompt, capture after each completed turn, flush before compaction. Claude Code, Codex, OpenClaw, and DeepSeek Harness all have packages, installed with one command per tool. The important property is that these paths are deterministic. Remembering does not depend on the model choosing to call a tool.

Direct API. For your own agents and multi-user products, the endpoint is OpenAI-compatible: a base_url change plus a Memory Key. /v1/memory/prepare returns context before inference, /v1/memory/ingest stores the exchange after it, and the rest of the surface covers search, list, edit, upload, and delete. This is the path where you own the integration and the memory is invisible to the model. The architecture guide walks through how the layer sits between your app and the providers.

All three write to the same vault, which is the point. One agent's decisions are another surface's context, and the AI memory layer is portable across providers instead of being rented from one.

Build versus buy, stated honestly

Build your own if your ranking requirements are genuinely unusual, or if data residency rules make a hosted service a non-starter. Those are real reasons.

Otherwise, price the whole job before you commit to building it. The parts people underestimate: extraction and cleaning of raw turns, deduplication, recency weighting, conflict resolution, deletion propagation, per-tenant isolation, and keeping recall fast as vaults grow. Each is a small problem. Together, they are the reason memory vendors exist, and they are why teams that estimate two weeks are still maintaining this eighteen months later.

Buying has its own checklist, which is the evaluation above. If a provider cannot answer the deletion question and the latency question with specifics, keep looking.

Verify, then stop thinking about it

Whatever path you choose, run the same acceptance test. State a disposable fact in one session with its reasoning attached. Close everything. Open a fresh session with nothing pasted, and ask for the fact and the reason. Then delete the record and confirm retrieval stops returning it.

A memory layer that passes that test twice, on real work, is doing its job. The threshold is when you stop noticing sessions starting: your agent opens with the project loaded, handoffs between agents stop requiring a briefing, and the reasoning behind decisions survives the week it was made.

Create your MemoryRouter account, attach the surface you use most, and run the acceptance test before real decisions go in.