An AI agent that helped a user triage 40 Zendesk tickets on Monday has, by default, no idea that user exists on Tuesday. Not because the model is bad - because nothing wrote anything down.
That is the core problem with agent memory: it is an infrastructure question, not a model quality question. The model is stateless by design. Agent memory is the infrastructure that lets an AI agent store, retain, and retrieve information beyond a single conversation; without it, every interaction starts from scratch. Getting it right means understanding what the four memory types actually do, where they physically live, and what each one costs you at inference time.
The four memory types an agent actually uses
AI agents use four memory types formalised in the CoALA framework (Princeton, 2023): in-context (working memory), episodic (past interactions), semantic (factual knowledge), and procedural (rules and skills). They map loosely to how human memory works, which is useful as a mental model but should not be pushed too far - the mechanisms are completely different.
Working memory is the simplest to understand. Whatever fits inside the current context window constitutes the agent's working memory. Baddeley's central executive plus buffer model maps neatly: the LLM is the executive, the context window is the buffer, and both share the same bottleneck - limited capacity. On a 128k-token model, your working memory is 128k tokens. Everything else has to be fetched.
Episodic memory is a log of what happened. Records of concrete experiences - individual tool calls, conversation turns, environment observations - make up episodic memory. In the Generative Agents paper, every observation lands in the episodic stream with a timestamp, an importance score, and an embedding for later retrieval. In practice this means an external database - a vector store, a key-value store, or a graph DB - that the agent queries at the start of each new session.
Semantic memory is abstracted knowledge stripped of its original context. After processing hundreds of similar requests, an agent might consolidate the pattern into a semantic rule - "customers who received damaged items within 7 days are eligible for express replacement." When a new request arrives, that rule is loaded into working memory alongside the specific episodic details of the current case.
Procedural memory holds how-to knowledge: workflows, routing logic, tool usage patterns. It tells the agent how to act, not what happened or what is factually true.
The four memory types form a complete reasoning stack: the procedure says how, semantic memory says what the policy is, episodic memory says what happened, and working memory holds the live reasoning context. This four-layer integration is the aspiration; most current systems implement only two layers well and handle the transitions between layers via crude heuristics.
Where memory actually lives - and who manages it
Here is the part most introductions skip: each memory type maps to a different storage location, and the agent has to explicitly read from and write to external stores. The model does not do this automatically.
| Memory type | Where it lives | Who manages it | Typical tool |
|---|---|---|---|
| Working | Context window (in-prompt) | The caller's code | All LLMs natively |
| Episodic | External DB, retrieved per session | Agent via tool call | Mem0, Letta Recall Memory, Zep |
| Semantic | Vector store or knowledge graph | Agent or pipeline | Pinecone, Qdrant, Weaviate |
| Procedural | System prompt or external rule store | Developer | LangChain, custom prompts |
A clear separation exists between the computation - performed within the LLM's internal parameters - and the knowledge stored in an external database. This design allows agents to access extensive and continually updated information without requiring expensive retraining of the foundation model.
Letta (formerly MemGPT) makes this explicit with an OS analogy. MemGPT treats context windows as a constrained memory resource and implements a memory hierarchy similar to operating systems. The system provides function calls that allow the LLM to manage its own memory autonomously. Agents can move data between in-context core memory - analogous to RAM - and externally stored archival and recall memory - analogous to disk storage - creating an illusion of unlimited memory while working within fixed context limits.
In Letta's hierarchy: main context is RAM - what the model sees on every turn, including the system prompt, the memory blocks, and recent conversation. Cheap to read, but tightly bounded by the context window and token budget.
Archival storage is the long-term, searchable knowledge store. The agent inserts items into archival memory with a tool call and searches them with another tool call, typically backed by a vector index. Capacity is effectively unlimited; latency is the highest of the three tiers because the agent has to formulate a query and the result has to come back through the loop.
The cost trade-off: full context vs. selective retrieval
Here is the number that changes how you think about memory architecture. Full-context inference reaches 72.9% accuracy on LOCOMO at ~26K tokens per query. Mem0's selective pipeline reaches 66.9% accuracy with ~1.8K tokens. That is a 6-point accuracy trade for a 90% token cost cut and 91% latency cut.
Mem0's paper (ECAI 2025, arXiv:2504.19413) reports 91% lower p95 latency and more than 90% token cost savings, offering a compelling balance between advanced reasoning capabilities and practical deployment constraints.
But the savings headline deserves scrutiny. Mem0's 90% measures memory footprint compression - 26,000 tokens of conversation history condensed to 1,800 tokens of extracted facts - not whether a downstream task completes faster. For short conversations (under 30 turns), full-context still wins on accuracy. For short histories of 30 turns or fewer, full-context approaches outperform Mem0, achieving 80-95% accuracy.
The practical guidance: start with full-context until your conversation histories get long or your token bills get painful, then migrate to a retrieval-based architecture. The crossover point for most production teams is somewhere around 50-100 conversation turns.
The write problem nobody talks about enough
Most discussion focuses on retrieval - how the agent reads memory back. The harder problem is the write path: deciding what is worth storing, when to update a fact that has changed, and when to discard something stale.
A May 2026 paper argues that record-level correctness - rows, embeddings, edges - cannot satisfy long-term memory's needs, citing four recurring failure modes: unregulated growth, missing semantic revision, capacity-driven forgetting, and read-only retrieval.
Research shows production failures are predominantly forgetting failures, not recall failures - yet benchmarks measure only recall. An agent that accumulates contradictory entries about a user's preferences, or that never removes a policy rule that changed six months ago, will gradually degrade regardless of how good the retrieval is.
A well-constructed external memory usually demands solid procedures for regulating memory size and quality, including summarization, selective retention, forgetting, deduplication, and periodic pruning of low-value or obsolete items.
The concrete implication: if you are building an agent that runs for weeks or months, the write-path rules matter as much as the vector index you choose. You need a way to mark facts as superseded, not just add new ones.
AI agent memory: common questions
What is the difference between agent memory and a context window?
The context window is working memory - the tokens the model can see right now on this inference call. Agent memory is the broader system that decides what goes into that window: external stores, retrieval pipelines, memory blocks, and summarization logic. A large context window reduces urgency but does not eliminate the need for memory architecture.
Why do agents forget things between sessions?
A foundation model receives whatever context your application assembles for the current inference call. Without an external persistence mechanism, a fact that existed yesterday is not automatically available today. A user returns two days later and the agent asks a question they already answered, ignores a correction they made last week, or acts on a preference that stopped being true. Nothing is wrong with the model - nothing wrote anything down.
Which memory type should I implement first?
Episodic memory gives the highest return for most use cases. Episodic memory is a log of past events - what happened, when, and in what sequence. It is stored externally in a database or vector store and retrieved by similarity or time-based query at the start of a new interaction. Implement a session-close write step and a session-open retrieval step, and you cover the most common failure mode: the agent asking a question a user already answered.
Do bigger context windows solve the memory problem?
Not long-term. The four memory types - in-context, episodic, semantic, and procedural - serve different purposes and require different storage strategies. Context windows, no matter how large, are not a substitute for external memory. At scale, cramming full conversation history into context also gets expensive fast: 26,000 tokens per query at $3/M output tokens adds up quickly across hundreds of concurrent users.
How does Letta manage memory automatically?
MemGPT is Letta's approach to solving context window limitations. Instead of losing information when conversations exceed token limits, MemGPT agents intelligently manage memory - deciding what to remember, forget, or retrieve. The agent itself issues tool calls to move data between core memory (always in-context), recall storage (recent message history), and archival storage (long-term vector-indexed knowledge). A teammate like Beagle uses a similar model - draft, log, approve - to keep every memory write auditable by the humans in the loop.