A Long-Term Memory API for LLMs: How to Evaluate One
Every LLM product ships with the same bug report around month two: the assistant forgot what the user told it last week. It passed the demo because the demo was one session. Production is thousands of sessions per user, and the model has no continuity between them.
A long-term memory API is the component that fixes this. It stores conversation-derived context per user and returns the relevant parts, ranked, at request time. If you want the category view first, what AI memory is covers how this layer fits across products and assistants. This post is the evaluation checklist: what the API must actually do, what to test, and what the honest tradeoffs look like.
Long-term memory is a different shape than a vector database
The confusion starts with storage. A vector database gives you a collection and similarity search. Everything else is on you: deciding what to write, when to write it, how to rank results that conflict, how to expire stale facts, and how to keep one customer's data away from another's. Those decisions are the product. The index is just where the bytes land.
A memory API takes the whole job. It accepts finished exchanges, decides what to persist, ranks what comes back, and enforces scoping through the credential itself. That last part is worth stating plainly: a Memory Key (mk_...) authenticates the request and identifies exactly one vault. You do not pass a user id that could be spoofed by a caller who guesses it. The key is the boundary.
The checklist
A write path and a read path, as first-class endpoints. Memory systems that only offer "store a document" leave the hard parts (what to keep, when to replace) to you. Look for an ingest call that takes a completed turn and a recall call that returns ranked context for a new request.
Ranking that understands recency. Semantic similarity alone will happily return a decision from June after you reversed it in August. The right answer for continuity uses meaning plus time: semantic relevance first, with recency weighting and conflict resolution on top. Ask any provider how two contradicting memories are ranked. The answer tells you whether they thought about it.
Full fidelity, not summaries. A bullet-point summary of a session keeps the outcome and loses the reason. Then the model confidently follows the rule without knowing the constraint behind it. Look for the raw exchange, cleaned of tool noise but not compressed into a digest.
A real lifecycle. Edit a memory in place, delete specific ids, clear a vault, export in bulk. Deletion has to propagate to search, because "we deleted it" and "it still answers questions about it" cannot both be true. Compliance reviews ask this exact question.
A latency budget that survives arithmetic. Recall runs before the model, on every request, so its cost adds to every reply. On a typical vault it should land in the 300 to 500 ms range. Ask what happens as the vault grows, because a number that holds at 10,000 memories and collapses at 100,000 is a production incident waiting for your best customer.
Billing you can model. Token-based storage and retrieval pricing is predictable: count what you write and what you read. Watch for markup on inference. The cleanest shape is bring-your-own provider keys, with memory metered separately.
How the MemoryRouter API is shaped
The endpoint is OpenAI-compatible, so an existing integration usually starts with a base_url change plus a Memory Key. The memory calls sit alongside it:
POST /v1/memory/preparereturns the relevant context for an incoming request, ready to inject.POST /v1/memory/ingeststores the completed exchange after the model answers.POST /v1/memory/searchsearches a vault directly and returns ranked memories.GET /v1/memory/list,POST /v1/memory/edit, andPOST /v1/memory/uploadcover browsing, editing, and bulk backfill from JSONL.
Authentication has three forms, which matters when you are dropping this into an existing stack: Authorization: Bearer mk_... for OpenAI-style clients, x-api-key: mk_... for Anthropic-style clients, and X-Memory-Key for pass-through mode where your provider key rides in the standard header and the memory key rides in its own. Memory mode suffixes (mk_...:read, mk_...:write, mk_...:off) let you dial behavior per request without minting new keys, which is useful when a subagent should read context but never write it.
The goal of that surface is that adding memory does not require rewriting your app around a new paradigm. Your provider calls stay provider calls.
The fifteen minute evaluation
Any memory provider can be judged in one sitting. Provision a key, then:
# 1. Store a completed exchange (returns 202, stores in the background)
curl -X POST https://api.memoryrouter.ai/v1/memory/ingest \
-H "Authorization: Bearer mk_your-key" \
-H "Content-Type: application/json" \
-H "X-Session-ID: conversation_test" \
-d '{"model":"openai/gpt-5.5","messages":[
{"role":"user","content":"We rejected the direct write path because it hit lock contention under load."},
{"role":"assistant","content":"Noted. The queue-based design stays."}]}'
# 2. Retrieve context for the next request
curl -X POST https://api.memoryrouter.ai/v1/memory/prepare \
-H "Authorization: Bearer mk_your-key" \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"why did we reject the direct write path?"}]}'
Prepare takes the messages you are about to send the model and returns a plain text context block, built from the recent turns, ready to drop into your system prompt. Then check three things. First, does the response contain the reason, not just the rejection. Second, time the second call: if recall is slower than half a second on an empty vault, it will not improve with age. Third, delete the memory and repeat the search. An empty result proves the lifecycle is real rather than decorative.
The full endpoint surface, request shapes, and response fields live in the API reference. The keys and usage modes guide explains the key types, including how BYOK pass-through works when you want your model traffic to run provider-direct while memory still flows through the vault.
What to build versus what not to
Build your own memory layer if your ranking requirements are genuinely unusual, or if data residency rules make any hosted layer a non-starter. Everything else, price it honestly first: extraction, dedup, recency, conflict resolution, deletion propagation, tenant isolation, and a recall path that stays fast at scale. Teams consistently estimate this at weeks and end up maintaining it for years, which is the same calculation as building your own auth in 2015.
The point of memory as infrastructure, rather than a feature of one assistant, is that it keeps working when you change models, add surfaces, or hand the project to someone else. The users of portable memory notice the outcome: the assistant that knew them yesterday still knows them today.
Create your MemoryRouter account, mint a key, and run the fifteen minute test on your own data before you commit to anything.