Most LLM agents shipped to production today are amnesiac by design. They answer a question, flush everything, and start the next request with a blank slate. That is fine for a search box. It is a serious architectural flaw for any agent that is supposed to learn from a conversation, respect a past decision, or avoid repeating a mistake it made last Tuesday.
The fix is not "add more context." It is understanding that agent memory is not one thing - it is four things, each with different storage, different cost, and different failure modes. Getting them confused is why most "persistent agent" builds feel brittle six weeks after launch.
What the four memory types actually are
AI agents use four memory types drawn from cognitive science - in-context (working) memory, episodic memory, semantic memory, and procedural memory - formalized for LLMs in the CoALA framework (Princeton, arXiv:2309.02427).
Each type stores a different class of information: the live context window, past events, factual knowledge, and behavioral rules respectively.
Here is what each one actually does at runtime:
Working memory is the context window. It is the agent's short-term working space - everything in the current context window. It is fast, perfectly accurate, and ephemeral: when the session ends, it is gone. Think of it as RAM. You do not need to build it; it exists by definition. What you do need to build is everything that survives past the session boundary.
Episodic memory is the raw log. Episodic memory records task-specific experiences
- full conversation turns, tool call outputs, ordered observations. It is valuable for audit trails, debugging, pattern analysis over conversations, and reflection loops that synthesize higher-level insights from experience. In practice, it usually looks like a message store: chunks of conversation history indexed by timestamp and retrieved via vector similarity when a new query arrives.
Semantic memory is the distilled version of episodic memory. The CoALA paper describes semantic memory as a repository of facts about the world; in agents today, it is most often used to personalize an application.
Practically, an LLM extracts information from the conversation or interactions the agent had. The exact shape of this information is application-specific. It is then retrieved in future conversations and inserted into the system prompt to influence the agent's responses. Example: an agent that handles engineering tickets notices the team always wants PR links included in summaries. That preference gets distilled into a semantic fact and prepended on the next session.
Procedural memory is the slowest-moving of the four. It represents knowledge of how to do things - the agent's skills, tool usage patterns, and stable workflows. For LLM agents, procedural memory is typically encoded in model weights, system prompts, and tool/function registries. It changes more slowly than other memory types and usually requires explicit updates: deployments, fine-tuning, tool additions.
Where each type lives and what it costs
The storage location determines the retrieval cost. That table is worth being explicit about:
| Memory type | Where it lives | Retrieval method | Latency | Per-session cost |
|---|---|---|---|---|
| Working | Context window (in-flight) | None - already in prompt | ~0 ms | Paid as input tokens |
| Episodic | Message store / vector DB | Embedding similarity search | 50-500 ms | Index + embedding write per turn |
| Semantic | Key-value store or vector DB | Exact key or similarity | 10-90 ms | LLM call on write to extract facts |
| Procedural | Model weights + system prompt | None - already compiled in | ~0 ms | Deployment / fine-tune cost upfront |
Systems requiring LLM-mediated entity extraction - including Mem0, Zep, and A-MEM - consume tokens and incur non-trivial latency at every write operation. Mem0 and Zep invoke two or more LLM calls per write, resulting in ingestion latencies of approximately 2 and 3 seconds respectively. For a customer support agent processing 1,000 messages a day, that adds up fast. It is the hidden "memory tax" most teams discover after launch.
For long-horizon workloads, full-history prefill cost is substantial and grows with accumulated history. Agent memory systems reduce this cost, but span two orders of magnitude in per-query serving latency on identical hardware, with comparable spread in construction costs. That spread is not random - it maps directly to which memory paradigm you chose.
The failure mode nobody ships around
The most common mistake: treating episodic memory as a substitute for semantic memory. A team stores raw conversation logs in a vector DB, retrieves the top-k chunks at each session start, and assumes the agent "remembers." It does - until the corpus grows. The problem surfaces when the developer expects the vector database to behave like memory, which involves more than retrieval. Memory implies knowing what is worth remembering, keeping information consistent over time, and updating beliefs when new information contradicts old ones.
Episodic memory accumulates contradiction. If a user says "I prefer Python" in March and "we've moved to TypeScript" in August, a raw vector retrieval might surface both - and the agent has no way to know which is current. Semantic memory solves this: a memory layer extracts and updates the fact ("language preference: TypeScript") rather than appending the raw exchange. Over time, raw experience becomes more useful when summarized into stable knowledge. Generative Agents popularized "reflection" mechanisms that periodically synthesize episodic memories into higher-level insights, which can then be stored as semantic memory and reused later.
The non-obvious consequence: episodic memory's real job is to be a source for semantic memory, not a permanent store. Agents can extract more reusable representations from episodic experiences. User preferences can be distilled as semantic memory, while navigation skills can be abstracted as procedural memory. Semantic and procedural memories are not isolated components - they are structured abstractions derived from episodic memory.
How prompt caching overlaps with working memory (and where it doesn't)
One thing that gets conflated: prompt caching and working memory. They operate at different layers.
Working memory is the current context the model reasons over. Prompt caching is a provider-side optimization that cuts the cost of processing the static parts of that context window. When an LLM processes your prompt, it generates key-value (KV) cache entries in its attention layers - mathematical representations of the relationships between tokens. Normally, the model recomputes this KV cache on every request. Prompt caching stores it so the model can skip that computation on subsequent requests that share the same prefix. The model still generates a fresh response every time; it is the redundant prefill work that gets cut.
It is exact-prefix matching, not similarity matching - that is semantic caching, a separate technique with separate economics. The practical consequence: your system prompt and static tool definitions are the prime caching targets. Dynamic content - conversation history, retrieved memories - has to come at the end, after the stable prefix, or you break the match and pay full price.
As of June 2026, all three major providers discount cached input by 90% on the GPT-5.x family, Anthropic's Claude models, and Gemini 2.5 and later. The fine print is where the math changes: Anthropic charges a cache-write surcharge (1.25× input for the 5-minute TTL, 2× for 1-hour), Gemini's explicit caching adds a per-hour storage fee, and each provider has a different minimum cacheable length and cache lifetime.
The connection to memory architecture: your procedural memory (system prompt, tool definitions) is exactly what prompt caching is designed to save you money on. Get the structure right - static content first, dynamic retrieved memories appended after - and you get both correct agent behavior and 90% off the largest part of your input bill.
AI agent memory: common questions
What is working memory in an AI agent?
Working memory is the agent's current context window - all tokens the model can see during a single inference call. It is fast and perfectly accurate, but ephemeral: nothing in working memory survives a session end unless your code explicitly writes it somewhere else. It is the foundation all other memory types feed into.
What is the difference between episodic and semantic memory in LLM agents?
Episodic memory stores raw experiences - full conversation turns, tool outputs, timestamps - in the order they happened. Semantic memory distills those experiences into stable, updateable facts. Episodic is good for audit trails and debugging; semantic is what lets an agent actually update its beliefs when a user contradicts something they said two months ago.
Do I need all four memory types in my agent?
Not always. A single-session task agent only needs working memory. A multi-turn assistant that personalizes over time needs episodic and semantic. An agent that needs to follow reliable workflows needs procedural memory in its system prompt. Most production agents need at least three of the four; the question is which three, and how they connect.
How does prompt caching relate to agent memory?
Prompt caching cuts the cost of reprocessing the static parts of your context window - system prompts, tool definitions, few-shot examples. That maps to procedural memory. It does not cache retrieved episodic or semantic facts, which change per request. Structure your prompts with static procedural context first and dynamic retrieved memory appended after to get the full discount.
What breaks when teams skip semantic memory?
The failure mode is expecting a vector database to behave like memory. Memory implies knowing what is worth remembering, keeping information consistent over time, and updating beliefs when new information contradicts old ones. Without a semantic extraction step, contradictions accumulate silently. The agent becomes less reliable as the episodic corpus grows, not more.