Understand AI Agent Memory Before You Build With It

AI agents have four distinct memory types, and most teams only wire up one. Here's how episodic, semantic, procedural, and working memory actually function under the hood.

Cover art for Understand AI Agent Memory Before You Build With It

Most LLM agents shipped to production today are amnesiac by design. They answer a question, flush everything, and start the next request with a blank slate. That is fine for a search box. It is a serious architectural flaw for any agent that is supposed to learn from a conversation, respect a past decision, or avoid repeating a mistake it made last Tuesday.

The fix is not "add more context." It is understanding that agent memory is not one thing - it is four things, each with different storage, different cost, and different failure modes. Getting them confused is why most "persistent agent" builds feel brittle six weeks after launch.

What the four memory types actually are

AI agents use four memory types drawn from cognitive science - in-context (working) memory, episodic memory, semantic memory, and procedural memory - formalized for LLMs in the CoALA framework (Princeton, arXiv:2309.02427).

Each type stores a different class of information: the live context window, past events, factual knowledge, and behavioral rules respectively.

Here is what each one actually does at runtime:

Working memory is the context window. It is the agent's short-term working space - everything in the current context window. It is fast, perfectly accurate, and ephemeral: when the session ends, it is gone. Think of it as RAM. You do not need to build it; it exists by definition. What you do need to build is everything that survives past the session boundary.

Episodic memory is the raw log. Episodic memory records task-specific experiences

  • full conversation turns, tool call outputs, ordered observations. It is valuable for audit trails, debugging, pattern analysis over conversations, and reflection loops that synthesize higher-level insights from experience. In practice, it usually looks like a message store: chunks of conversation history indexed by timestamp and retrieved via vector similarity when a new query arrives.

Semantic memory is the distilled version of episodic memory. The CoALA paper describes semantic memory as a repository of facts about the world; in agents today, it is most often used to personalize an application.

Practically, an LLM extracts information from the conversation or interactions the agent had. The exact shape of this information is application-specific. It is then retrieved in future conversations and inserted into the system prompt to influence the agent's responses. Example: an agent that handles engineering tickets notices the team always wants PR links included in summaries. That preference gets distilled into a semantic fact and prepended on the next session.

Procedural memory is the slowest-moving of the four. It represents knowledge of how to do things - the agent's skills, tool usage patterns, and stable workflows. For LLM agents, procedural memory is typically encoded in model weights, system prompts, and tool/function registries. It changes more slowly than other memory types and usually requires explicit updates: deployments, fine-tuning, tool additions.

Where each type lives and what it costs

The storage location determines the retrieval cost. That table is worth being explicit about:

Memory type Where it lives Retrieval method Latency Per-session cost
Working Context window (in-flight) None - already in prompt ~0 ms Paid as input tokens
Episodic Message store / vector DB Embedding similarity search 50-500 ms Index + embedding write per turn
Semantic Key-value store or vector DB Exact key or similarity 10-90 ms LLM call on write to extract facts
Procedural Model weights + system prompt None - already compiled in ~0 ms Deployment / fine-tune cost upfront

Systems requiring LLM-mediated entity extraction - including Mem0, Zep, and A-MEM - consume tokens and incur non-trivial latency at every write operation. Mem0 and Zep invoke two or more LLM calls per write, resulting in ingestion latencies of approximately 2 and 3 seconds respectively. For a customer support agent processing 1,000 messages a day, that adds up fast. It is the hidden "memory tax" most teams discover after launch.

For long-horizon workloads, full-history prefill cost is substantial and grows with accumulated history. Agent memory systems reduce this cost, but span two orders of magnitude in per-query serving latency on identical hardware, with comparable spread in construction costs. That spread is not random - it maps directly to which memory paradigm you chose.

The failure mode nobody ships around

The most common mistake: treating episodic memory as a substitute for semantic memory. A team stores raw conversation logs in a vector DB, retrieves the top-k chunks at each session start, and assumes the agent "remembers." It does - until the corpus grows. The problem surfaces when the developer expects the vector database to behave like memory, which involves more than retrieval. Memory implies knowing what is worth remembering, keeping information consistent over time, and updating beliefs when new information contradicts old ones.

Episodic memory accumulates contradiction. If a user says "I prefer Python" in March and "we've moved to TypeScript" in August, a raw vector retrieval might surface both - and the agent has no way to know which is current. Semantic memory solves this: a memory layer extracts and updates the fact ("language preference: TypeScript") rather than appending the raw exchange. Over time, raw experience becomes more useful when summarized into stable knowledge. Generative Agents popularized "reflection" mechanisms that periodically synthesize episodic memories into higher-level insights, which can then be stored as semantic memory and reused later.

The non-obvious consequence: episodic memory's real job is to be a source for semantic memory, not a permanent store. Agents can extract more reusable representations from episodic experiences. User preferences can be distilled as semantic memory, while navigation skills can be abstracted as procedural memory. Semantic and procedural memories are not isolated components - they are structured abstractions derived from episodic memory.

Beagle in action#product, Tuesday, 10:22am
The ask
'can you pull together the context from last sprint's retro before the planning call?'
Beagle drafts
searches episodic memory for last sprint's retro thread, extracts the key decisions as a semantic summary, drafts a brief with action items and owners
You approve
you approve; the summary posts into the planning thread with links to the source messages - no manual digging
Do this in your workspace →

How prompt caching overlaps with working memory (and where it doesn't)

One thing that gets conflated: prompt caching and working memory. They operate at different layers.

Working memory is the current context the model reasons over. Prompt caching is a provider-side optimization that cuts the cost of processing the static parts of that context window. When an LLM processes your prompt, it generates key-value (KV) cache entries in its attention layers - mathematical representations of the relationships between tokens. Normally, the model recomputes this KV cache on every request. Prompt caching stores it so the model can skip that computation on subsequent requests that share the same prefix. The model still generates a fresh response every time; it is the redundant prefill work that gets cut.

It is exact-prefix matching, not similarity matching - that is semantic caching, a separate technique with separate economics. The practical consequence: your system prompt and static tool definitions are the prime caching targets. Dynamic content - conversation history, retrieved memories - has to come at the end, after the stable prefix, or you break the match and pay full price.

As of June 2026, all three major providers discount cached input by 90% on the GPT-5.x family, Anthropic's Claude models, and Gemini 2.5 and later. The fine print is where the math changes: Anthropic charges a cache-write surcharge (1.25× input for the 5-minute TTL, 2× for 1-hour), Gemini's explicit caching adds a per-hour storage fee, and each provider has a different minimum cacheable length and cache lifetime.

The connection to memory architecture: your procedural memory (system prompt, tool definitions) is exactly what prompt caching is designed to save you money on. Get the structure right - static content first, dynamic retrieved memories appended after - and you get both correct agent behavior and 90% off the largest part of your input bill.

Agent memory before and after a deliberate architecture
Without Beagle
raw conversation logs dumped into a vector DB; retrieval surfaces contradictory facts; the agent re-introduces old preferences the user changed months ago
With Beagle
episodic logs feed a semantic extraction pass; the agent holds a clean, updateable fact store; working memory is topped with cached procedural context at ~0 extra cost per request
4distinct memory typesworking, episodic, semantic, procedural
2-3 singestion latencyfor graph-based memory systems like Zep per write
90%cached input discountall three major providers as of June 2026
2 orders of magnitudespread in per-query latencyacross agent memory systems on identical hardware

AI agent memory: common questions

What is working memory in an AI agent?

Working memory is the agent's current context window - all tokens the model can see during a single inference call. It is fast and perfectly accurate, but ephemeral: nothing in working memory survives a session end unless your code explicitly writes it somewhere else. It is the foundation all other memory types feed into.

What is the difference between episodic and semantic memory in LLM agents?

Episodic memory stores raw experiences - full conversation turns, tool outputs, timestamps - in the order they happened. Semantic memory distills those experiences into stable, updateable facts. Episodic is good for audit trails and debugging; semantic is what lets an agent actually update its beliefs when a user contradicts something they said two months ago.

Do I need all four memory types in my agent?

Not always. A single-session task agent only needs working memory. A multi-turn assistant that personalizes over time needs episodic and semantic. An agent that needs to follow reliable workflows needs procedural memory in its system prompt. Most production agents need at least three of the four; the question is which three, and how they connect.

How does prompt caching relate to agent memory?

Prompt caching cuts the cost of reprocessing the static parts of your context window - system prompts, tool definitions, few-shot examples. That maps to procedural memory. It does not cache retrieved episodic or semantic facts, which change per request. Structure your prompts with static procedural context first and dynamic retrieved memory appended after to get the full discount.

What breaks when teams skip semantic memory?

The failure mode is expecting a vector database to behave like memory. Memory implies knowing what is worth remembering, keeping information consistent over time, and updating beliefs when new information contradicts old ones. Without a semantic extraction step, contradictions accumulate silently. The agent becomes less reliable as the episodic corpus grows, not more.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle