Harvey's legal-drafting agents were failing the same way, over and over. Filetype quirks forgotten. Tool workarounds rediscovered from scratch every session. According to Anthropic's launch announcement, Harvey's task completion rates rose roughly 6x in internal testing once Dreaming was turned on - not because of a model upgrade, but because the agents kept hitting the same small failures each run.
The feature that fixed it is called Dreaming. Anthropic shipped it alongside outcomes and multiagent orchestration for Claude Managed Agents on May 6, 2026. It is currently in research preview - gated behind a request form. But the architecture it introduces is worth understanding before it goes broad.
What Dreaming actually does (and what it does not)
Agents write to their memory stores as they work, but these writes are local and incremental: over many sessions a memory store accumulates duplicates, contradictions, and stale entries. Dreams let Claude clean that up. A dream reads an existing memory store alongside past session transcripts, then produces a new, reorganized memory store: duplicates merged, stale or contradicted entries replaced with the latest value, and new insights surfaced. The input store is never modified, so you can review the output and discard it if you don't like the result.
Supported models are Claude Opus 4.7 and Claude Sonnet 4.6; up to 100 past sessions can feed a single dream.
Dreaming is billed at standard API token rates and access is gated behind a request form, so it is not yet in the Claude consumer app.
The system does not change model weights. It curates the agent's persistent memory and context so the next session starts with a cleaner, more accurate knowledge base than the previous one left behind. That distinction matters. Dreaming is not self-training - it is scheduled janitorial work on a text file. It is a scheduled cron job that edits your agent's memory file between sessions - pruning duplicates, resolving contradictions, and surfacing patterns no single session could see.
Dreaming is an out-of-band asynchronous process which Anthropic says solves the in-band limitation, where agents must split effort between completing and executing tasks while also concurrently curating memory for their future selves. Dreaming spots recurring failure patterns where agents are consistently failing - wrong units, missing topics, broken tool configs - and proposes memory-store updates, again for human review.
The Harvey number is real, but it has a shape
Worth noting what that number does and doesn't prove. It comes from Harvey's internal testing on their specific use case. No external benchmark has verified it yet.
The 6x figure sounds extraordinary, but the underlying mechanism is mundane. Long-horizon legal work involves many repeated motifs: similar contract clauses, similar discovery queries, similar precedent searches across different matters. Without consolidation, each new session forces the agent to rediscover patterns it already encountered. The memory store grows but does not get smarter - it just gets longer, more contradictory, and harder for the model to navigate within its context window.
The Harvey 6x is an upper bound - your team's number will likely land between 1.5x and 3x on completion rate, with cost-per-completion reductions in the 30-60% band, on workloads that have repeated patterns and persistent memory.
Dreaming provides the most value for agents running the same task category repeatedly over many sessions - document review pipelines, customer support bots, code review systems, content generation agents. One-off or low-frequency agents do not accumulate enough session history to benefit significantly.
Google's Memory Bank: the other side of the same idea
Google's Memory Bank launched at I/O 2026 Day 1 (May 19) as part of the Gemini Enterprise Agent Platform, alongside ADK 2.0 GA. Thirteen days after Dreaming.
The architecture differs. Google's Memory Bank is an identity-scoped database primitive in the Gemini Enterprise Agent Platform.
Memory extraction pulls only the most meaningful information from source data to persist as memories. Memory consolidation then integrates newly extracted information with existing memories, allowing memories to evolve as new information is ingested.
Asynchronous generation means your agent doesn't have to wait for memory generation to complete. Continuous event ingestion streams and manages conversation events, automatically triggering memory generation based on batching rules you configure.
| Anthropic Dreaming | Google Memory Bank | |
|---|---|---|
| Launch | May 6, 2026 (research preview) | May 19, 2026 (GA with ADK 2.0) |
| Availability | Request-gated | Generally available |
| Trigger | Scheduled / /dream command |
Event-driven or manual API call |
| Input | Up to 100 past session transcripts | Streaming conversation events |
| Scope | Project / agent level | Identity-scoped per user |
| Storage | Client-controlled (local FS, S3, etc.) | Managed Google Cloud service |
| Pricing | Standard API token rates | Separate billing lines per Google Cloud |
Google announced the Gemini Enterprise Agent Platform on April 22, 2026 with Memory Bank, Agent Sessions, Agent Registry, and Agent Gateway named as components. That is 132 days between launch and the first metered day for agent state. The billing model is worth reading carefully before you wire this into production: Memory Bank and Sessions are now billed on three separate lines each, and none of them is a per-conversation price.
The security problem both vendors are underplaying
Persistent memory is a genuine improvement for high-frequency, repetitive agent workloads. It is also a wider attack surface than a context window. Cisco's security team discovered a method to compromise Claude Code's memory and maintain persistence beyond an immediate session into every project, every session, and even after reboots. They broke down how they were able to poison an AI coding agent's memory system, causing it to deliver insecure, manipulated guidance to the user.
After working with Anthropic's Application Security team on the issue, they pushed a change to Claude Code v2.1.50 that removes this capability from the system prompt.
That was before Dreaming existed. The consolidation pass introduces a subtler variant. Dreaming makes the risk subtler. It is one thing for a bad instruction to sit in a memory file. It is another for a dream process to encounter a set of poisoned, stale, or adversarial traces and clean them up into a more elegant, more authoritative-looking playbook. The danger is not only that the agent remembers the wrong thing. It is that it consolidates the wrong thing.
Researchers have proposed sleeper memory poisoning - a delayed attack in which adversarial external content causes an assistant to store a fabricated user memory that can later be retrieved and used after the original malicious context is gone. Results show that universal poisoning payloads can induce memory writes across proprietary models, including ChatGPT, Claude, and Gemini, and across memory-management regimes.
Anthropic's documentation flags the issue and recommends human review on memory updates for high-stakes workflows. The recommended mitigation is a layered store design. The recommended pattern is a three-store layout: use a read-only organization store for stable standards; use a read-only project store for verified architecture and domain facts; use a read-write working store for session lessons. Run dreams over the working store and recent verified sessions, then promote only reviewed outputs into the project store.
That is sound advice. It also means Dreaming is not a fire-and-forget feature. Someone on your team has to own the review gate.
Persistent agent memory: common questions
What is Anthropic Dreaming for Claude agents?
Claude Dreaming is a scheduled background process that reviews an AI agent's past sessions and rewrites its memory store, removing duplicates, replacing stale entries, and surfacing new patterns. Anthropic launched it on May 6, 2026 at Code with Claude as a research preview, alongside Outcomes and Multiagent Orchestration. It does not change model weights. It makes the context the next session sees cleaner and more accurate.
How is Dreaming different from a longer context window?
A model with a 200,000-token context sounds like infinite memory, but sending the entire conversation history on every call is expensive and inefficient. More importantly, context quality matters more than context volume - a 200,000-token window stuffed with irrelevant history produces worse outputs than a 10,000-token window with precisely the right information. Dreaming is the system that keeps the token budget lean by removing noise before the next session starts.
Is the Harvey 6x task completion gain credible?
Harvey describes the 6x task completion rate increase as driven by agents requiring dramatically fewer clarification rounds and fewer error corrections per task - the agent already knows what preferences apply and what mistakes to avoid. The 6x figure is task completion rate - tasks completed to user satisfaction per session - not raw speed. The number is plausible for a high-repetition workflow like legal drafting, but has not been verified by an external benchmark. Treat it as a ceiling, not a median.
Can agent memory be poisoned through Dreaming?
Yes. Giving agents structured persistent memory expands the attack surface for prompt-injection and memory-poisoning attacks. If a malicious input can convince an agent that the wrong instruction is the right one, Dreaming may consolidate that wrong instruction into the agent's long-term memory store, where it will be applied to future sessions automatically. The mitigation is separating writable working memory from reviewed, promoted project memory - and keeping a human on the promotion step.
How does Google Memory Bank compare to Anthropic Dreaming?
Memory Bank focuses on extraction and consolidation: it pulls meaningful information from conversation events and integrates newly extracted information with existing memories, allowing them to evolve as new information is ingested. Dreaming focuses on retrospective cleanup of what an agent already wrote. Both address memory decay; Memory Bank is generally available on Google Cloud while Dreaming remains in gated research preview as of September 2026.