Local Inference Does Not Mean Local Memory

John Rood··6 min read

Your model runs on hardware you own. The weights, the endpoint, and the trace of every call are inside your network. Then a new session opens with none of last week in it.

That is where most local stacks discover they solved one problem and skipped another. Where inference runs and where memory lives are two separate decisions. The first one gets made on purpose. The second one usually never gets made at all.

What the local stack already bought you

Control of the request path. Keys, routing, streaming, retries, evals, logging. You picked the runtime, the quantization, the upgrade schedule, and none of that is in question here. A memory service that asked you to give any of it up would be a nonstarter.

What the stack did not come with is a memory policy. A larger context window does not close that gap. The window is what a single request can carry. Memory is what remains when the request is gone. A runtime will serve a session that knows nothing about yesterday, every day, without complaint.

Where the context actually lives

Ask a local setup where last month's decisions went and the honest answer is file paths. Session logs under one tool's directory. A notes file kept in two places. The reasoning from a conversation in a different app, stored in a format only that app reads. The moment a second tool or a second machine touches the same project, the trail does not follow.

Maintaining your own durable store is a legitimate answer, and a genuine project. Deduplication, ranking, deletion, and per-user isolation each look small alone and grow teeth in combination. There is a lighter path if you accept one tradeoff. MemoryRouter's local inference mode leaves inference exactly where it is and adds retrieval and storage as a separate path over HTTP.

The loop: three calls

  • POST /v1/memory/prepare takes the messages you are about to send the model and returns this user's memory as text. One string, or null when nothing relevant matches. It does not call a provider, and it does not know which one you use next.
  • Your application injects that text into the prompt and calls the provider itself, with the same keys, routing, and streaming path as before. The provider call never passes through MemoryRouter.
  • POST /v1/memory/ingest takes the completed exchange. It accepts immediately and stores in the background, so it adds nothing to your user's latency. With streaming, call it once, after the final message is assembled.

Three round trips where proxy mode has one. The difference buys control: the model call stays in your code, so streaming, retries, evals, and logging stay there too. And because prepare returns plain text, the loop is provider-agnostic. The runtime behind the provider call, whether that is Ollama, LM Studio, llama.cpp, or a gateway you wrote yourself, is invisible to the memory path.

What leaves the machine

Plainly, because this is the part that decides the fit. The messages sent to prepare and the exchange sent to ingest leave your device. Retrieval and storage run on hosted infrastructure. Keeping inference local does not keep the submitted content local. Local inference mode is not an offline mode and not a zero-data-sharing mode.

If the requirement is that no conversation content leaves your network under any configuration, keep everything local and run that maintenance yourself. A hosted memory service is the wrong tool for that workload, and no setting turns it into the right one. The fit for local inference mode is narrower: tokens get generated in exactly one place, under your control, and memory is a deliberate, visible exception with its own key and its own boundary.

One scoping note. This is an API integration, not a plugin for your runtime, and the three calls above are the entire contract.

What your app still owns

Keys. One stable Memory Key per end user, resolved on your server from the authenticated user. Session grouping is not an isolation boundary, and it never was one.

Scope. Key suffixes split the loop: mk_xxx reaches both endpoints, mk_xxx:read suits retrieve-only apps, mk_xxx:write suits store-only ones.

Failure policy. When prepare is slow or unavailable, your application decides whether the reply goes out without memory. A 202 from ingest means queued, not searchable, so do not grade recall against a memory that is still in the pipeline.

Trust. Retrieved text is background data. Inject it as context, and never let it set application policy or authorize a tool call.

Prove it with a fake fact

Run the loop on synthetic data before any real history is involved.

  1. Through your application, complete one exchange: "For the demo, the Juniper workstation marks staging tasks with blue labels." Send it to ingest and check the acceptance response.
  2. Start a fresh conversation for the same user, with no old messages, and call prepare with only the new question: "Which color marks staging tasks on the Juniper workstation?"
  3. Read the returned context before you read the model's answer. The phrase has to appear in the retrieved text. That step separates retrieval from a model's talent for sounding right.
  4. Ask the same question with a second user's key and confirm the phrase does not come back.

Four steps on fake data, and the boundary stops being something you read and becomes something you watched.

Where this sits

If your application does not already own routing, streaming, and retries, proxy mode is the recommended default: one round trip, less code. Local inference mode is for stacks that will not hand over the call. Both paths write the same vault, and you can start on one and move to the other later without changing the user's memories.

The local inference documentation has the request fields, the key modes, and the failure checklist. The integrations directory shows where this path sits next to proxy mode. Create your MemoryRouter account, run the synthetic exchange through the loop once, and the next session can start with the last one's context.