How to Give Your AI Chatbot Memory Across Sessions
Session one, your chatbot is charming. Session two, it greets the same user like a stranger, and the user decides it is a toy. No amount of prompt engineering fixes that, because the problem is architectural: a new conversation starts from a blank context window, and nothing in it knows that a person said anything before.
If you want the category view first, what AI memory is covers how persistent memory works across assistants, products, and agents. This post is the build guide for people shipping a chatbot: the identity model, the two API calls, the lifecycle work that compliance will ask about, and the tests that prove it works.
The shape of the fix
One mapping carries the whole design:
your user.id -> memoryKey (mk_...) -> private vault -> AI requests with memory
Your app owns users. The memory layer owns user-scoped memory. A Memory Key both authenticates a request and identifies exactly one vault, which means isolation is enforced by the credential rather than by a parameter your code has to pass correctly. That key-per-user model is also what lets memory follow a person beyond your product, which is the cross-AI memory story from your users' side.
Step 1: Provision a key per user
Create the key when the user signs up or sends their first message, and store it mapped to your internal user id. Two ways exist: programmatically, on your backend, with an account key you never expose to clients, or by hand in the dashboard for early testing.
async function getOrCreateMemoryKey(userId: string) {
const existing = await db.memoryKeys.findUnique({ where: { userId } })
if (existing) return existing.memoryKey
const res = await fetch('https://api.memoryrouter.ai/v1/keys', {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.MEMORYROUTER_ACCOUNT_KEY}`,
'Content-Type': 'application/json'
},
body: JSON.stringify({ name: `user:${userId}` })
})
const { key } = await res.json()
await db.memoryKeys.create({ data: { userId, memoryKey: key } })
return key
}
The new key belongs to your account and inherits any provider keys you have stored, so one account key can mint a key for every end user. The quickstart covers the full flow, including the dashboard path.
Step 2: Recall before you generate
When a message arrives, call POST /v1/memory/prepare with the messages you are about to send the model. You do not write a query string: retrieval is built from the recent non-system turns. The response is a plain text block you drop into your system prompt.
curl -X POST https://api.memoryrouter.ai/v1/memory/prepare \
-H "Authorization: Bearer mk_user-123" \
-H "Content-Type: application/json" \
-H "X-Session-ID: conversation_abc" \
-d '{"messages":[{"role":"user","content":"anything on the pricing page change today?"}]}'
Two details keep this cheap. First, density (low, default, high, xhigh) controls how much context comes back, so chat surfaces can run lighter than a coding agent. Second, if nothing relevant is found, context comes back null and you skip injection entirely. Recall for a typical vault averages 300 to 500 ms, which is a real cost in front of every reply, so treat the retrieval budget as part of your latency plan rather than an afterthought.
Step 3: Capture after the turn
Once the model answers and you have sent the reply to the user, store the exchange:
curl -X POST https://api.memoryrouter.ai/v1/memory/ingest \
-H "Authorization: Bearer mk_user-123" \
-H "Content-Type: application/json" \
-H "X-Session-ID: conversation_abc" \
-d '{"model":"openai/gpt-5.5","messages":[
{"role":"user","content":"anything on the pricing page change today?"},
{"role":"assistant","content":"The Starter plan still includes 200M memory tokens each month."}]}'
Ingest returns 202 immediately and stores in the background, so this never sits between your user and their reply. Store the user turn and the final answer, not your internal tool chatter. The value of full-fidelity turns over tidy summaries is that the reason behind a decision survives; "we picked the queue design because the direct path hit lock contention" is worth ten cleaned-up bullet points.
Step 4: Run the lifecycle before you need it
Memory is a data store, so it inherits every obligation your other stores have.
- Delete a single memory when a user points at something wrong.
POST /v1/memory/editreplaces text in place;POST /v1/memory/delete-idsremoves specific records. - Wipe a vault when an account closes.
DELETE /v1/memoryclears memory for the key, and it should be wired into the same backend path that deletes the user everywhere else. - Export on request.
GET /v1/memory/listpaginates through everything, and JSONL upload works in the other direction for migrations. - Rotate keys without losing memory, so a leaked key is an inconvenience rather than an incident.
The user lifecycle guide documents each of these with the exact calls, including creation, rotation, deletion, export, and migration paths.
Step 5: Say the privacy story out loud
Users ask where their chat history went the moment you announce memory. Have the answer on your own site before they ask.
In the MemoryRouter case, memories are encrypted at rest and in transit, isolated to the account behind one key, and never trained on. The user can read everything stored in the dashboard and delete any of it, including the whole vault. If your chatbot handles anything regulated, let that be the beginning of your review rather than the summary, but do not ship memory without a sentence users can find.
The tests that count
Three tests, ten minutes total, and they catch almost everything:
- Two-session test. State a disposable fact in one session with its reason. Close everything. Ask cold in a new session, pasting nothing. You should get the fact and the reason. A control session without memory connected should come back empty.
- Wrong-user test. Send user A's question with user B's key. You should get B's context or nothing, never A's. This is the test that proves your isolation is real.
- Deletion test. Delete a memory, repeat the question that used it, confirm it is gone from retrieval. If it still shows up, your lifecycle is decorative.
What memory does not fix
- It does not fix a chatbot that never writes. Capture is what fills the vault, so ingest has to run reliably even when the model answer is short.
- It does not replace structured state. Orders, balances, and entitlements belong in your own database. Memory is for context, not for the system of record.
- It does not make retrieval psychic. It surfaces what resembles the current request; the model still does the reasoning.
- It does not run itself. The two calls above are your integration, and the plugins exist because manually wiring capture is exactly the thing people forget.
The bar users hold you to is simple: the second conversation should feel like the second conversation, not the first one again. Two API calls, one key per user, and a deletion path are what separate a chatbot people keep using from a demo people try once.
Create your MemoryRouter account, provision your first test key, and run the two-session test on your own product before you ship it.