Persistent Memory for Open WebUI: Keep Native Memory, Add a Vault

John Rood··5 min read

Open WebUI earns its place. You tune the connections once, keep a local model for drafts, reach for a cloud model when the problem gets hard, and the workspace quietly becomes where the thinking happens. Somewhere in there is a decision the project depends on: the export format the downstream job expects, or the reason the Tuesday report moved to Thursday.

Native memory in Open WebUI is real, and it keeps doing its job inside the instance. The question this post answers is narrower. What does it take to have a decision like that available to the assistant you open somewhere else?

A vault is a second home for context, not a replacement

Open WebUI documents its own memory features, and they are worth keeping on. Nothing here replaces them. The gap an external vault covers is portability: a store that several supported clients can read, so a decision recorded while you work in one place can be recalled from another. The integrations directory lists which clients can point at the same vault.

That matters most for the person running a chat workspace and a coding agent side by side. Open WebUI learns the project in one window, the coding agent works the repo in another, and neither can see what the other learned. A shared vault is how one of them stops re-explaining the setup to the other.

The provider route, and exactly what it covers

The connection lives where your other providers live: Settings → Connections, or the Admin Panel when connections are managed centrally. Add an OpenAI-compatible connection with the URL https://api.memoryrouter.ai/v1 and a Memory Key, refresh the model list, and start a new chat with a model that belongs to that connection. The Open WebUI docs page walks through both the personal and the admin path, including the read, write, and off key suffix modes for controlling what a specific connection is allowed to do.

The part to internalize is that coverage follows the route. Requests that go through the MemoryRouter connection get capture and recall. Everything else behaves exactly as it does today. A direct Ollama connection you leave in place keeps running untouched, with no memory behavior added to it.

Old conversations are a separate job

Connecting the provider starts memory for new chats. It does not import the conversations you already have. Backfill is its own deliberate operation: a privileged Python function you review and install through Admin → Functions, then invoke as the user whose chats should be imported. It walks the active branch of that user's conversations and sends supported user and assistant text in batches.

Two things the function is not. It is not a per-user import ledger: the destination key and the completion flag are shared configuration, so on a multi-user instance one successful upload can suppress the next user's run, and a casual flag reset can send older content again. And it is not the approval-bound imports pipeline, which shows a preview and a price quote, then waits for your explicit approval and returns an itemized receipt.

One key is one vault

If other people use your instance, this is the trap that matters. A connection using one Memory Key shares one vault among every caller. Open WebUI accounts do not turn into separate MemoryRouter vaults on their own, and an administrator-wide key copied from a personal account is the wrong move on an instance you do not fully trust.

For a team, decide up front: one shared project vault, or a per-user mapping with the keys handled server-side. Either way, start with a single consenting test user and prove the boundary you actually intend to run. A second key is the cheapest isolation check. Ask it the same question and confirm the answer does not come from the first vault.

Prove it with a fact that could not ride along

The recall test takes a minute, and the interesting version ends outside Open WebUI.

  1. In a chat on the routed connection, write: "The synthetic demo marker for Project Halyard is TERN-4407."
  2. Let the exchange complete and give ingestion a moment.
  3. Open a fresh chat on the same connection and ask: "What is the synthetic demo marker for Project Halyard?" Keep the marker out of the question.
  4. Now connect a second supported client to the same vault and ask there.

Step four is the one people skip. If the fresh chat returns TERN-4407, the memory path is working. If the second client returns it too, the fact reached the vault itself, and that is the part native memory was never going to handle for you.

What this is not

  • Not offline. MemoryRouter is a hosted service. Selected conversation data leaves the machine running Open WebUI, so route the chats you are comfortable sending.
  • Not universal. Only requests through the MemoryRouter connection are covered, and model compatibility still matters. Check the supported models page before assuming every backend is covered.
  • Not a superset of your history. Recall returns relevant, bounded context rather than every conversation you have ever had.
  • Not an isolation layer. Shared credentials share data access, and removing the connection stops future routing without deleting what is already stored.

Keep the first run small on purpose. Connect one test user, store one synthetic fact, and watch a second tool read it back. If that is the setup your week is missing: create your MemoryRouter account and run the two-tool check on a real key.