One task at ProjectDiscovery ran 67.5 million input tokens across 1,225 steps at a 91.8% cache rate. A nearly identical task with a 3.2% cache rate cost roughly 60 times more. Same model. Same token volume. The only difference was prompt structure.
That gap is what prompt caching is about. Not a new model, not a smaller context window-just whether the API recomputes tokens it already computed sixty seconds ago.
What the KV cache actually stores
When an LLM processes your prompt, it generates key-value (KV) cache entries in its attention layers-mathematical representations of the relationships between tokens. Normally, the model recomputes this KV cache on every request. Prompt caching stores it so the model can skip that computation on subsequent requests that share the same prefix. The model still generates a fresh response every time; it's the redundant prefill work that gets cut.
This matters because of where the latency actually lives. Every LLM request goes through two latency phases: time to first token (TTFT), which measures how long the model takes to start responding, and time to last token (TTLT), which captures the full generation time. Both get worse as your prompts get longer. A long system prompt increases TTFT because the model processes every token through its attention mechanism before producing any output.
You send a request with a 2,000-token system prompt, and the model reprocesses every token. Send the next turn, and it reprocesses them again. Twenty turns in, you have paid to compute the same static instructions twenty times over. Prompt caching fixes this by storing the model's computed state for the parts of your prompt that do not change, so those tokens are reused across requests instead of recomputed.
One thing worth internalizing: prompt caching behaves differently from a standard cache in one crucial way. It does not return a stored answer. It reuses the internal computation for a shared prompt prefix, then the model still generates a new response token by token. If you expect it to return a saved reply, you're thinking of semantic caching-a different layer entirely.
How Anthropic and OpenAI implement it differently
OpenAI prompt caching is fully automatic-zero code changes, 50% cost discount on cached tokens, roughly 50% hit rate (best effort, not guaranteed). Anthropic prompt caching is manual-you set cache_control breakpoints, get a 90% cost discount on cache reads, and a 100% guaranteed hit rate when configured correctly. Both require a minimum of 1,024 tokens before a prefix is eligible.
| OpenAI | Anthropic | |
|---|---|---|
| Activation | Automatic | Manual cache_control breakpoints |
| Discount on cache read | ~50-90% (model-dependent) | 90% ($0.30/M vs $3.00/M) |
| Write premium | Free through GPT-5.5; 1.25× from GPT-5.6 | 1.25× (5-min TTL) or 2× (1-hr TTL) |
| Hit guarantee | Best-effort (~50%) | 100% on exact prefix match |
| Cache TTL | 5-10 min inactivity | 5 min default; 1 hr available |
| Min token threshold | 1,024 tokens | 1,024 (Sonnet/Opus); 2,048 (Haiku) |
Anthropic's design assumes you know your workload and want maximum savings on it. OpenAI's design assumes you'd rather not think about it. If your traffic is high-frequency and your prompts are large and stable, Anthropic's 90% discount will dominate the comparison; if your traffic is sporadic or your prompts evolve frequently, OpenAI's automatic 50% with no write premium is the cleaner choice.
On the TTL side, the cache's default minimum lifetime is 5 minutes, refreshed each time the cached content is used. If you find that 5 minutes is too short, Anthropic also offers a 1-hour cache TTL. The lifetime is measured from the start of the request that writes or reads the cache entry, not from the end of its response. Time spent generating a response counts against the lifetime, so the window for a follow-up request to reuse the cache is the lifetime minus the generation time.
That last point is non-obvious and costs teams money. A model that takes 90 seconds to generate a long response has consumed 90 seconds of a 5-minute TTL before the next turn even starts.
The prompt ordering rule that changes everything
The cache is a prefix cache. Cache matching is all or nothing: a changed character, timestamp, tool serialization, or rewritten earlier turn can cause a complete miss.
The practical design goal is to place everything stable-tool definitions, system instructions, retrieved documents, few-shot examples-as early in the prompt as possible, and to push everything that varies per request-the user's specific question, a session ID, a live timestamp-as late as possible, ideally after the last cache breakpoint.
Static function schemas are excellent caching candidates because they tend to be large and unchanged across calls. Place them in the cached prefix ahead of dynamic content, and they contribute to your prefix match on every request.
The most common way teams destroy their own hit rate is by letting dynamic content creep into the prefix. Timestamps and session IDs in the prefix destroy cache performance. Injecting something like "Today is March 6, 2026" into a system prompt invalidates the cache every day. Use a precise timestamp and it invalidates on every request.
At ProjectDiscovery, two tasks with nearly identical token volume showed a 91.8% cache rate versus 3.2%-roughly 60× the cost difference. The latter ran before the optimization rollout. Their fix: moving dynamic working memory out of the system prompt and into a trailing user message. One structural change. Cache hit rate jumped from 7% to 84%, overall LLM spend dropped 59%.
Where caching pays and where it quietly hurts
Prompt caching is not universally beneficial. A write premium makes reuse volume decisive; without one, caching has no downside from the first request.
For Anthropic's 5-minute TTL, the write cost is 1.25× standard. With the 5-minute TTL, you save the write premium back after one cache hit. Hit once, you've already broken even. Every hit after that is pure savings. For the 1-hour TTL at 2× write cost, you break even after two cache hits.
Break-even lands at 2.3 reuses of the same cached prefix within the one-hour TTL window. Any workload where the same system prompt or tool definitions are sent more than twice per hour is already in the money.
Where it hurts: one-shot queries that never repeat within the TTL window will pay the write premium with no offsetting reads. For most developers starting out, OpenAI's automatic caching is the right call-it requires no code changes and immediately applies to any prompt over 1,024 tokens. If you're at significant scale (500K+ tokens/day in static context), the jump to Anthropic's 90% discount justifies the implementation overhead.
Agentic workloads are the strongest case. Agentic workflows that loop through a plan, call a tool, read the result, and call the next tool burn through far more tokens than a single question-and-answer, because the model re-reads the accumulating conversation on every turn. A ten-turn agent session with a 5,000-token system prompt reprocesses that same prompt ten separate times if nothing is cached.
Without caching, a 20,000-token system prompt would be processed fully for every step in a 100-request workflow, resulting in 2 million processed tokens. With caching, the system pays the full input cost only for the first request. Later requests treat the static prefix as cached input, priced at a fraction of the standard rate. This approach results in cost reductions of approximately 89% for extended sessions, making autonomous operation economically viable.
One more non-obvious risk: cached data is never shared between businesses, even when two different customers send byte-identical prompts. Anthropic and OpenAI both certify their caching as zero data retention (ZDR) eligible-raw prompt text isn't stored at rest, only the computed key-value representations are held in memory, and they're destroyed once the TTL expires. That's worth knowing before security teams flag it.
A teammate like Beagle, which reads and replies inside Slack with a fixed system prompt on every request, is a natural fit for prefix caching-the tool definitions and behavioral instructions are identical across every channel message, and the only thing that changes is the user's text at the end.
Prompt caching LLM: common questions
What is prompt caching in LLMs?
Prompt caching stores the key-value tensors from an LLM's attention layers after processing a repeated prompt prefix. When the next request starts with the same prefix, the model skips that computation entirely. It still generates a fresh response-only the prefill work is saved. Both OpenAI and Anthropic support it natively via their APIs.
How much does prompt caching save?
Anthropic's prompt caching reduces costs by up to 90% and latency by up to 85% for long prompts. OpenAI achieves 50% cost reduction with automatic caching enabled by default. Real production numbers land between 41% and 80% cost reduction across providers, depending on how much of each prompt is stable prefix versus dynamic content.
Why is my prompt cache hit rate low?
The most common cause is dynamic content inside the cached prefix. A timestamp, session ID, or serialized tool result that changes per request will invalidate the entire prefix from that point forward. Cache matching is all or nothing: a changed character, timestamp, tool serialization, or rewritten earlier turn can cause a complete miss. Audit what lives before your cache breakpoint and move anything dynamic to the end.
Does prompt caching affect output quality?
No. The model still runs full autoregressive generation for every response. The static portion of every request bills at up to 90% off, with the model producing byte-identical output. No distillation, no quantization, no quality trade-off.
Is prompt caching safe for sensitive data?
Cached data is never shared between businesses, even on byte-identical prompts. Anthropic and OpenAI both certify caching as zero data retention (ZDR) eligible-raw prompt text isn't stored at rest, only the computed key-value representations are held in memory and destroyed when the TTL expires. Review your provider's current data-handling terms before caching anything regulated.