A Memory Proxy for LLM Apps: One URL Change, Per-User Vaults

John Rood··5 min read

Your app stores every message. The model still starts each request with none of them.

That gap is where most shipped AI features quietly break. A user explains their setup once, and weeks later the product asks them to explain it again. The database holds everything from that first session, and nothing decides which parts of it belong in front of the next call.

There are two places to make that decision. Inside your application, where the store, the ranking, and the injection are yours to build. Or in the request path, where a memory layer sits between your app and the provider and handles all of it on every call. This post is about the second shape, the memory proxy, and how to tell whether your app should use it.

What changes in your code

One URL and one key.

from openai import OpenAI

client = OpenAI(
    api_key="mk_xxxxxxxxxxxxxxxx",  # the Memory Key for this end user
    base_url="https://api.memoryrouter.ai/v1",
)

Everything else about the call survives once the SDK points at MemoryRouter. The messages array, the model IDs, streaming, retries, the code that parses the response. The response returns in the standard format, so nothing that reads it notices the move. Provider credentials are configured once in the dashboard. The Memory Key is the only credential your app carries, and it belongs on your server, like the key it replaced.

Key suffixes tune behavior when you need it. :read retrieves without storing, :write stores without retrieving, :off skips memory and passes the call through untouched. A plain key does both jobs.

What happens on one request

Five things happen between your app sending the call and reading the answer.

  • The Memory Key resolves to one vault, so the request reads and writes exactly one user's history. There is no user id in the payload for a caller to spoof.
  • Relevant memories are retrieved for the incoming request.
  • Retrieved context is injected into the prompt, and the call is forwarded to the provider configured for that vault.
  • The completed exchange is stored for later.
  • The provider's response comes back in the standard shape.

From your app's side, the call looks like any other completion call. From the user's side, the next session starts with the last one's context. Keep sending the live conversation in the messages array; memory covers everything that came before it.

What recall adds to the request

Retrieval runs before the provider call, so its cost sits inside every request your users wait on. Budget for it the way you budget for any dependency in the path, and measure it on your own vault before you commit to a timeout. On a typical vault, average recall lands in the 300 to 500 ms range.

Routing and memory are different decisions

If your stack already routes model calls through a gateway, LiteLLM, OpenRouter, or one you wrote yourself, the routing question is settled. Keep it separate from the memory question:

  • Routing decides where the call runs: which provider serves it, with what fallback and spend rules.
  • Memory decides what the call carries: what this user already decided and explained.

A gateway that routes cleanly still delivers every call with empty memory, and no amount of routing fixes a user who repeats themselves every week. If you evaluate a memory layer as a gateway feature, hold it to the same test as anything else in the path: two requests, a fresh messages array, and a second call that recalls what the first one stored.

Prove it with two requests

Run this on fake data before real history is involved. It takes a few minutes.

  1. Send one request: "For the demo, the Quill account wants shipping summaries grouped by project." Let the response come back and the background ingestion finish.
  2. Send a second request with a fresh messages array and only the new question: "How does the Quill account want shipping summaries grouped?" Do not replay the first request's history, and do not paste the answer into the question. If you replay the original prompt, you tested chat history, not memory.
  3. Confirm the recall. If it misses, inspect the vault and retry before suspecting anything else; ingestion completes in the background, and a memory still in the pipeline is not searchable yet.
  4. Send the same question with a different Memory Key. Quill's preference must not come back. That is the isolation receipt, and it costs one extra call.

Four steps on fake data, and the two properties that matter here are proven: recall across requests and isolation between keys.

When proxy mode does not fit

Proxy mode puts MemoryRouter in the model path, and inference requests pass through it on the way to your provider. If a policy requires the provider call to leave from your own infrastructure, use the retrieval and storage pair (POST /v1/memory/prepare and POST /v1/memory/ingest) that keeps the model call inside your stack. The local inference documentation covers that path.

Two boundaries worth stating plainly:

  • The Memory Key is a memory boundary, not an identity layer. Application login stays yours, and the key gets resolved on your server from the authenticated user rather than chosen by the client. Callers sharing a key share a vault, so users and projects that must not mix get their own keys.
  • A compatible HTTP shape is not feature parity. Streaming, tool calls, and provider-specific request features each deserve a test with your real request shape before you build on the route.

The proxy quickstart has the request fields, the key suffix table, and the failure checklist. The proxy page walks the request path end to end. Create your MemoryRouter account, run the two-request test, and the second request comes back with the first one's context.