Your Retriever Knows the Docs. Your Agent Forgets the Customer.
The first working version of a LangChain app usually stands on a folder of documents, and it is good enough to demo. Then the same user comes back a week later and has to introduce themselves again. The retriever never failed. Nothing in the stack was holding anything about the person using it.
Retrieval answers questions. Memory is about a user.
A retriever turns a query into source chunks. Every visitor gets the same chunks, and none of them are about the person asking.
Conversational memory is a different input: the preferences and decisions a user carries between sessions, stored under their own key and recalled when relevant.
Confusing the two is easy, because from inside the app both look like the same operation. A query goes out, text comes back, the text enters the prompt. The difference only shows up in the next session, when the same user asks for something you never put in a document.
LangChain documents its own long-term memory patterns, and a Postgres table keyed by user ID is a legitimate way to build this yourself. A managed path also exists: langchain-memoryrouter connects a LangChain app to MemoryRouter vaults, with the setup and the exact tool signatures in the LangChain docs. The package is optional. LangChain is not missing a feature, and this is one way to supply the layer.
Prove recall from a second process
The fastest way to learn whether memory is actually wired is two scripts in separate processes, sharing one Memory Key. The full walkthrough lives in the docs linked above. The short version:
# retain_demo.py
import os
from langchain_core.messages import AIMessage, HumanMessage
from langchain_memoryrouter import create_memory_tools
retain, _ = create_memory_tools(os.environ["MEMORYROUTER_API_KEY"])
print(retain.invoke({
"messages": [
HumanMessage(content="Synthetic test: Project Kestrel's launch color is copper."),
AIMessage(content="Noted: Project Kestrel uses copper for its launch color."),
],
"session_id": "langchain-proof-a",
}))
# recall_demo.py
import os
from langchain_memoryrouter import create_memory_tools
_, recall = create_memory_tools(os.environ["MEMORYROUTER_API_KEY"])
print(recall.invoke({
"query": "What launch color did we choose for Project Kestrel?",
"limit": 5,
"session_id": "langchain-proof-b",
}))
Run the first script, give ingestion a moment, then run the second. The first should report Retained 2 conversation message(s). That confirms the ingest request succeeded, not that indexing has finished. The second should return memory containing Project Kestrel and copper, and it does not pass the original conversation to the search. If the match has not landed yet, wait briefly and rerun only the recall script. Re-ingesting the same test over and over will not speed it up.
Separate processes and separate session IDs, so nothing from the first run can tag along. A test that keeps the previous messages in the list is testing the list.
Tools run when the model calls them
create_memory_tools returns a retain tool and a recall tool. Binding them to a model supplies definitions only. Your executor still has to run the calls the model returns, and a model can skip a call on the turn where it mattered most.
So pick the behavior the product needs. Model-selected memory suits open-ended assistants. Recall on every turn is the deterministic option, through LangGraph nodes or a direct recall.invoke before the model, and it is the one to choose when a skipped lookup is a bug.
One key per user
Once real users arrive, the mapping moves to the server: authenticated user, their Memory Key, tools built from that key. A session ID labels a conversation; it does not separate two users who share a key. And a Memory Key or user-to-key mapping never comes from the client.
Then run the negative control. Point the recall script at a second user's key and confirm it returns nothing about Kestrel. Isolation you have watched work beats isolation you assume works.
What this is not
- MemoryRouter is a hosted service. Running a local LangChain model does not keep the retained conversation local.
- The retain path keeps human and assistant text and drops system messages, tool results, and tool-call metadata. It is not a secret-redaction system, so sensitive text still needs your own controls.
- Retrieval is semantic and can return no match. An empty recall is a possible result, not always a defect.
- The package does not run your agent loop, and memory does not appear because a library is installed. The wiring above is the part that stores it.
Try it on a test key
Install langchain-memoryrouter==0.1.2 against a test key and make the second script return what the first one stored. The integrations directory has the wider setup, including the other clients that can read the same vault.
When a durable home for per-user context is the piece your app is missing: create your MemoryRouter account and run both scripts against it.