Agent Memory

Ask an assistant something today, and unless something was specifically built to remember it, a new conversation tomorrow starts with no idea it ever happened. Agent memory is what makes that persistence possible: information a system deliberately stores and can bring back later, across separate sessions, instead of everything vanishing the moment one request ends.

Why the context window alone doesn't do this

A model only knows what's in its context window for the current request — think of that as what's currently on the desk. Once a session ends, that desk gets cleared. Memory is a separate system, built on top of the model, that decides what's worth keeping, stores it somewhere else, and brings the relevant parts back to the desk when a new request could actually use them. The context window doesn't persist on its own; memory is what makes something persist on purpose.

Short-term and long-term memory

Short-term: the conversation history within one ongoing session, resent as part of the context on every turn so the model has continuity within that single conversation. It disappears once the session ends unless something separately saves it.

Long-term: information deliberately saved and carried across sessions — a user's stated preferences, a fact mentioned weeks ago, a decision made in an earlier project. It has to be explicitly retrieved and reinjected into a future request's context; nothing about it survives automatically just because it happened before.

Deciding what to store, and retrieving it

Most memory systems do some version of three things: decide what's worth storing (often a summary or an extracted fact, not the entire raw conversation), store it somewhere searchable, and retrieve whatever's relevant to the current request before adding it to context. That last step usually looks a lot like RAG — searching stored memories by relevance and injecting the useful ones — the difference is what's being searched: a general document collection for RAG, a store of things this specific system has previously learned about this specific user or task for memory.

A real limitation worth knowing

Memory doesn't bypass the context window — a retrieved memory still has to fit inside it, and still competes for the same space as everything else in the request. Storing everything indiscriminately just relocates the over-retrieval problem RAG already has: too many saved memories crowding a request, most of them irrelevant to the current question. And because stored memory can include anything the system previously read, injected text can end up saved there too, and fire again later, in a session where the original source is no longer even present — the prompt injection page's own checklist already flags this: treat stored memory as untrusted on read, not as safe because it already made it into storage.

When it's worth building

Reach for memory when a system genuinely benefits from continuity — a recurring user whose preferences matter, a long-running task picked back up across sessions, an assistant meant to feel consistent over time rather than starting fresh each conversation. Skip it for one-off, stateless interactions where nothing about this request needs to inform a future one; building storage and retrieval for that adds a system that can fail for no real benefit.

In this guide
  1. Why the context window alone doesn't do this
  2. Short-term and long-term memory
  3. Deciding what to store, and retrieving it
  4. A real limitation worth knowing
  5. When it's worth building
  6. FAQ

FAQ

If a conversation is still within its context window, is that the same thing as memory?

No. That's just the current request still holding everything so far — nothing has been separately decided as worth keeping, and once the session ends, none of it persists unless a memory system explicitly saved it first.

Does adding memory make a system remember everything perfectly?

No — a memory system decides what to store, and that decision can miss things or save the wrong things, the same way a person's own notes can be incomplete. It's deliberate, selective persistence, not perfect recall.

Practice interview questions on agent memory →