Swap the Model, Keep the Memory

John Rood··4 min read

Run the same task through two models and a ranking writes itself. One model won. The trouble is that the two runs started from different places, so nobody can say whether the model or its missing context produced the difference.

That is the hidden variable in a model bakeoff, and it bites hardest when the stakes are real. Models get swapped for practical reasons, and none of them wait for a clean comparison. Every switch re-asks the same question: is this model better at the work, or does it just have less to go on?

Pin the context, change the model

What fixes this is moving the context out of the prompt entirely, into a vault every run reads from. Then the model is the only thing you changed.

MemoryRouter keeps memory separate from the model producing the answer. Any supported model retrieves from the same vault across sessions, and a switch does not mean re-importing the vault. There are two routes for this. On the proxy path, retrieval and storage happen inside the request path. With local inference, your application calls memory directly and keeps the model call on your own infrastructure.

The test: two models, one vault

Run this on synthetic data before real history is anywhere near it.

  1. Store one fact with the first model: "Remember this for the demo: the Larkspur launch sign-off code is amber-larkspur-27."
  2. Close that process. Open a new one with a different model, the same vault, and only the question: "What is the Larkspur sign-off code? Use retrieved memory, and say if it is unavailable."
  3. Count it as proof only when the code comes back with retrieval behind it. A model that says "I remember" without evidence is autocomplete being confident, not memory working.
  4. Ask the same question against a vault that never stored the fact. Whatever happens there is the noise floor. If a plausible code appears, you have watched the model improvise, and now a bare answer in the real run means nothing.

What you are looking for is specific: the second model returns the code from the shared vault, from a fresh process, with the code never entering its prompt. Both runs share the same memory key, because the key is the vault. The control uses a second key, which is what makes it a control.

The runnable version of this recipe, with the request shape for each provider, is in the supported models documentation.

What the catalog tells you, and what it does not

GET https://api.memoryrouter.ai/v1/models returns the providers and model IDs your configured dashboard keys expose, a suggested default, and the catalog's snapshot time. Read that timestamp before treating the list as current.

Two boundaries matter here. The endpoint reads a snapshot rather than a live entitlement check, so an empty list can simply mean no provider key is configured yet. And appearing in your catalog is not a promise that every provider capability survives the route. The compatibility overview shows coverage across providers, and the provider's own docs remain the source for what a specific model ID can do today.

What still changes when the model changes

Switching models does not switch everything. The following stays provider business:

  • Credentials and request formats. Native Anthropic and Gemini requests differ from the OpenAI-compatible path, and no provider key crosses over to another provider.
  • Capability checks per route. A chat model in the catalog does not imply support for tools, vision, streaming semantics, or embeddings on that route.
  • Billing, availability, and deprecations. Retrieval does not touch any of them.
  • Rate limits. They remain the provider's to enforce, and rotating memory keys is not a workaround.
  • Context spend. Retrieved memory still consumes input context, because retrieval does not enlarge the context window.

None of that argues against switching. It argues for holding memory constant while the rest moves, so a change stays legible six months later when you cannot remember which run had which problem.

Keep the baseline, switch the model

A model switch should be one config line plus the test above. The vault does not move and the memory does not get re-imported. The comparison you ran last quarter still means something next quarter.

Run it on a disposable vault: create your MemoryRouter account, store one synthetic fact, and ask for it from a second model. Every model switch after that one runs against a baseline that does not move.