Vector Memory vs RAG: Why Agents Need Both, and They Are Not the Same

John Rood··6 min read

RAG and vector memory get filed together because both use embeddings and both end in "the model knows more than it did before." That overlap hides the difference that matters: RAG retrieves from a corpus you wrote, and memory retrieves from work you did. Different writers, different lifecycles, different failure modes. Teams that blur them end up maintaining a document pipeline that cannot remember a decision.

If you want the category view first, what AI memory is covers where both fit across assistants, products, and agents. This post is the comparison: what each system is actually for, why one cannot do the other's job, and how they should compose in a real agent.

RAG, stated plainly

Retrieval-augmented generation is a read-mostly system. You author a corpus (documentation, tickets, policies, a knowledge base), chunk it, embed it, and retrieve passages at question time so the answer is grounded in something real.

The question RAG answers: what do the sources say?

The corpus changes when a human edits it. The retrieval target is a passage that supports an answer. Recency matters only if the source documents have a timestamp meaning. Grounding and citation matter a lot, because the whole point is that the model is not making things up.

Agent memory, stated plainly

Agent memory is a write-heavy system. The corpus is the work itself: turns, decisions, rejections, preferences, status, the reasoning behind a choice. It is written by the agent as a side effect of running, not by a human sitting down to curate.

The question memory answers: what did we decide, prefer, or try, and why?

The corpus changes every session. The retrieval target is context that reconstructs continuity. Recency matters enormously, because a preference stated last week should outweigh one from six months ago. Deletion matters enormously, because memory accumulates things people did not intend to keep.

Where they diverge

  • Who writes. RAG: humans, deliberately. Memory: the system, automatically, at the end of each turn. A memory layer that requires manual curation will not get curated.
  • Lifecycle. Documents get edited and versioned. Memories accrue, go stale, and get superseded. If nothing can expire or edit them, the vault degrades into noise the model still trusts.
  • Granularity. RAG wants chunks sized for citation. Memory wants the whole turn, cleaned but not compressed, because the reasoning is the part you cannot reconstruct later.
  • Ranking. RAG weighs authority and similarity. Memory weighs similarity plus recency, and has to resolve contradictions rather than return both sides.
  • Scoping. A document corpus is usually shared. Memory is usually personal or project-scoped, and getting that boundary wrong is a privacy incident, not a ranking bug.

The failure mode of RAG-as-memory

The tempting shortcut is to point the document pipeline at conversations: chunk the transcripts, embed them, retrieve. It demos fine and degrades in a specific way. A chunked transcript loses the shape of a decision. The question, the constraints, the rejected alternative, and the reason all live across chunk boundaries, and the retrieved fragment is "so we went with the queue" with none of the why. There is no notion of one decision superseding another, so both come back and the model picks one. And nobody wrote anything deliberately, so nothing in the pipeline ever cleans it up.

You can fix each of these one at a time. At the end, you have built an agent memory layer and you own it. That is a legitimate path if you want it. Go in clear-eyed about the eighteen months of maintenance.

The failure mode of memory-as-RAG

The mirror-image mistake is stuffing a knowledge base into a memory vault. Policies and handbooks are authored documents with authority; memory ranking is built for recency and relevance, not for "this is the canonical procedure." Edits to an authoritative document are awkward when the surrounding system treats every record as an accreted moment. Keep documents in your document pipeline. Keep memory for what happened.

How they compose

In a working agent, both systems answer different parts of the same request, and both feed the same prompt.

  • Memory supplies continuity: the project state, the constraints that were set, the approach that got rejected and why, the formats and preferences this person has established.
  • Retrieval supplies facts: the API contract, the internal policy, the runbook, the paragraph from the spec that governs the change.

The ordering that keeps a budget is worth the two minutes it takes to design. Inject memory first, because it is smaller and more request-specific, then retrieval if the task calls for sources. Keep them in separate blocks with clear provenance, so the model (and the human auditing the prompt) can tell context from source material. When they disagree, remember that memory records what you said and retrieval records what someone wrote; the disagreement itself is usually worth surfacing.

What to check on the memory side

If you are adding the memory half, the questions that separate a real implementation from a demo:

  • Does retrieval rank by meaning and time together, so a reversal two weeks ago outranks the original decision three months back?
  • Are stored memories full exchanges rather than bullet summaries, so reasons survive alongside outcomes?
  • Can you edit and delete individual memories, with deletion propagating to retrieval?
  • Is scoping enforced by credential, so one user's vault cannot be read with another's key?
  • What is the recall latency for a typical vault? In the MemoryRouter implementation, 300 to 500 ms is the average for a typical vault, and the number is measured before the model call rather than after.

The core concepts guide covers how vaults, memory keys, and retrieval fit together, including how raw memories and consolidated reflections are stratified for retrieval.

RAG was the right answer for a question that got solved: how do you ground a model in knowledge it was never trained on. Agent continuity is a different question, and it needs a system that writes, ranks, and forgets on purpose.

Create your MemoryRouter account, connect your agent, and let the document pipeline go back to documents.