AI Agent Memory Types, Mapped to What They Actually Store

AI agents have four distinct memory types - working, episodic, semantic, and procedural - that operate at different timescales and storage layers. Here's what each one does and where it breaks.

Cover art for AI Agent Memory Types, Mapped to What They Actually Store

An AI support agent closes a ticket at 2pm. At 9am the next day, the same customer writes back. The agent has no idea who they are. Not because it is dumb - because nobody told it how to remember.

A stateless AI agent has no memory of previous calls. Every request starts from scratch. That is fine for isolated tasks. It is a serious problem the moment an agent needs to track decisions, recall a user preference, or pick up a thread it dropped yesterday. The solution is not a single database - it is four distinct memory types, each storing a different class of information, each failing in a different way.

What the four AI agent memory types actually store

AI agents use four types of memory drawn from cognitive science - in-context (working) memory, episodic memory, semantic memory, and procedural memory - formalised for LLMs in the CoALA framework (Princeton, arXiv:2309.02427). Each type stores a different class of information: the live context window, past events, factual knowledge, and behavioural rules respectively.

Here is what that means in practice:

Memory type What it stores Where it lives Gone when?
Working Current prompt, tool outputs, reasoning steps Context window Session ends
Episodic Past interactions, outcomes, decisions Vector DB or log Never, unless evicted
Semantic Facts, definitions, entity relationships Vector DB / graph When updated
Procedural Skills, workflows, known failure modes System prompt / fine-tune When you change it

Working memory is fast (zero retrieval latency - it is already in context), volatile (lost when the context is cleared), and capacity-limited. Everything else requires an explicit write-and-retrieve cycle.

Working memory: what the model can actually see right now

To make a conversation work at all, your app re-sends the entire chat or last N messages every time the user says something new: the instructions you gave the model, all the earlier messages, any results from tools, and anything you looked up for it. That whole package of text you send in is working memory. It is the only thing the model can see while it writes its answer.

This is also where prompt caching intersects with agent memory in a way most architecture diagrams skip. When an LLM processes your prompt, it generates key-value (KV) cache entries in its attention layers - mathematical representations of the relationships between tokens. Normally, the model recomputes this KV cache on every request. Prompt caching stores it so the model can skip that computation on subsequent requests that share the same prefix. The model still generates a fresh response every time; it is the redundant prefill work that gets cut.

The practical consequence: a long, stable system prompt - your agent's persona, tool definitions, and operating rules - can be cached and reused. Anthropic's write cost is 1.25× input for a 5-minute TTL, but reads cost 0.10× input - a 90% discount. Minimum 1,024 tokens. For an agent running hundreds of requests a day against the same system prompt, that is not a footnote - it is the difference between a viable cost model and one that gets cancelled in a quarterly review.

Episodic memory: the record of what happened

Episodic memory captures what happened: specific events, tool calls, and their outcomes. Semantic memory captures what is true: facts and preferences extracted from experience. These two are routinely conflated, and the conflation causes bugs.

When your support agent sees "customer is frustrated about billing" in a new session, it should pull that from episodic memory - a timestamped record of the previous ticket, stored after that conversation ended. This memory type is often implemented using vector databases, which enable semantic retrieval across past episodes. Instead of exact matching, the agent can find experiences that are conceptually similar to the current situation, even if the details differ. In practice, episodic memory stores structured records of interactions: timestamps, user identifiers, actions taken, environmental conditions, and outcomes observed.

The problem is retrieval quality, and it is worse than vendor materials suggest. The EMGEB benchmark (arXiv:2501.13121, ICLR 2025) evaluated GPT-4o, Claude 3.5 Sonnet, Llama 3.1-405B and others on a synthetic 200-chapter narrative. On questions involving two or more events, no model matched the correct latest state more than 36% of the time, and no model recalled the full set of events in a chain more than 18% of the time. Even when a model retrieved the right events it often put them in the wrong order.

That is the real ceiling on episodic memory today. The storage layer is mostly solved. The temporal reasoning layer is not.

36%max correct on multi-event queriesacross all models tested at ICLR 2025
90%discount on Anthropic cache readsvs. fresh input token pricing
200-500msvector DB retrieval latencybefore embedding, reranking, and inference

Semantic and procedural memory: the parts teams skip

Semantic memory functions as the agent's knowledge repository. It stores factual information, concepts, rules, and general knowledge. This memory type provides the foundational understanding necessary for reasoning and inference, containing declarative knowledge such as definitions, relationships between concepts, and general principles that guide the agent's behavior.

Semantic memory for AI agents often stores the largest volume of agent knowledge in production. The volume of facts an agent accumulates over months of operation outgrows any context window long before it outgrows a properly indexed vector store. A support agent that has handled 10,000 tickets has a body of semantic knowledge - what "tier 2 escalation" means in your org, which errors map to which fixes - that no system prompt can hold.

Procedural memory is what most teams handle least consciously. Most frameworks bake procedural memory into the system prompt. More sophisticated setups extract patterns from accumulated episodic data and update procedural memory over time, so the agent gets better at routine tasks the more it runs them.

Procedural memory is often the most neglected. Most agents have no mechanism for encoding "approach A worked better than approach B for this class of problem" and using that in future decisions. The agent that fails at a given task on Monday will fail at it identically on Friday, unless someone updates its system prompt by hand.

Beagle in action#customer-success, 8:52am
The ask
'Priya just wrote back - she was frustrated about billing last month, can someone pull her history?'
Beagle drafts
reads the linked Zendesk thread and CRM note from the prior session, drafts a reply with the relevant context and prior resolution
You approve
you hit approve; the rep gets the history in-thread in under 30 seconds, with the source linked
Do this in your workspace →

"Just add a vector DB" is the wrong mental model

The most common agent memory mistake is treating this as a retrieval problem when it is a storage architecture problem. "Just add a vector DB" feels like an answer and then quietly fails because it mixes two different questions: what class of information you are storing (episodic vs. semantic vs. procedural) and where that information should live.

Appending every interaction to a vector store eventually produces retrieval noise, context dilution, and latency spikes. Memory consolidation is what prevents this. A well-designed agent memory layer writes selectively. An agent that writes every token of every interaction to memory produces noise at scale. Memory has to be selective.

Support agent between sessions
Without Beagle
agent starts fresh every ticket - customer repeats context, agent re-reads the same docs, the prior resolution is invisible
With Beagle
episodic store surfaces the prior ticket summary; semantic store provides current account facts; working memory holds the live conversation

The practical sequence for a team building this: start with working memory only (it is free, it is implicit). Add semantic memory when your agent starts answering factual questions wrong. Add episodic only when session continuity matters. Add procedural when you notice the agent repeating the same class of mistake across runs and you want it to stop.

AI agent memory types: common questions

What is working memory in an AI agent?

Working memory is the active context window - everything the model can see during a single inference call. It includes the system prompt, conversation history, tool outputs, and any retrieved documents. It is fast and zero-latency because it is already in the prompt, but it is volatile: once the call ends, nothing persists automatically.

What is the difference between episodic and semantic memory in an AI agent?

Episodic memory stores time-stamped records of specific past events - what happened, when, and what the outcome was. Semantic memory stores facts and relationships that persist regardless of when they were learned. An agent that recalls "this user prefers weekly summaries" is using semantic memory; one that recalls "this user complained about invoice 4471 last Tuesday" is using episodic memory.

Why do AI agents forget between sessions?

By default, LLMs are stateless - every API call starts from a blank slate. Session continuity requires an application layer that writes selected context to external storage after each session and retrieves it at the start of the next. Without that write-and-retrieve cycle, no memory survives the session boundary regardless of the model's size or capability.

Does a vector database solve the agent memory problem?

Partly. Vector databases solve a real problem in RAG architecture: inject relevant context into an LLM prompt at query time. For episodic memory lookup - retrieving recent conversation history, surface-level personalization, and document search - vector stores are fast and effective. But they struggle with temporal ordering, multi-hop relationship queries, and procedural knowledge, which require graph databases or structured stores. A production memory layer typically uses a vector store as one component, not the whole architecture.

How does prompt caching relate to agent memory?

Prompt caching is not a memory type but it is a key cost control for agents that use long, stable working memory. Prompt caching refers to the productized, provider-managed features that reuse KV tensors across API requests when prompts share common prefixes. By caching the KV tensors from the prefill phase, providers can skip redundant computation when subsequent requests begin with the same content, reducing both latency and cost. For an agent with a 2,000-token system prompt running 500 requests a day, the difference between cached and uncached reads is material at every provider's current pricing.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle