The average prompt token count in production applications grew nearly 4x between early 2024 and late 2025, from roughly 1,500 tokens to 6,000 per request. Every one of those extra tokens costs money to process-and if your system prompt is identical on every call, you are paying to recompute the same math thousands of times a day. LLM prompt caching is the mechanism that stops that. Not by returning a cached answer, but by skipping redundant GPU work inside the model itself.
What LLM prompt caching actually stores
Prompt caching works by reusing the key-value (KV) tensors that a transformer generates during attention computation-not by caching text or responses.
When an LLM processes your prompt, it generates key-value (KV) cache entries in its attention layers-mathematical representations of the relationships between tokens. Normally, the model recomputes this KV cache on every request. Prompt caching stores it so the model can skip that computation on subsequent requests that share the same prefix. The model still generates a fresh response every time; it is the redundant prefill work that gets cut.
That last point is worth dwelling on. It is worth separating this from semantic caching, which stores full prompt-response pairs and returns a saved answer for similar queries. Prompt caching does not reuse outputs. You still get a fresh, potentially different response each time, because only the intermediate computation is cached, not the final result.
The practical constraint is prefix matching. Inference providers check if a prefix of a prompt matches a cached entry upon receiving a request. If a match is found, the provider can reuse the cached KV attention states for the matching prefix and only compute the attention states for the unique tokens in the prompt.
A single character change anywhere in the cached prefix invalidates the match, so exact prefix matches matter more here than in most caching you have worked with.
Providers hold on to these matrices for each prompt for 5-10 minutes after the request is made, and if you send a new request that starts with the same prompt, they reuse the cached K and V rather than recalculating them.
How Anthropic and OpenAI implement it differently
Both providers offer prompt caching natively, but the mechanics-and the economics-differ enough that choosing the wrong one for your workload costs real money.
| OpenAI | Anthropic | |
|---|---|---|
| Activation | Automatic, no code change | Explicit cache_control breakpoints in your API payload |
| Minimum prefix | 1,024 tokens | 1,024 tokens (Sonnet/Opus), 2,048 tokens (Haiku) |
| Cache read price | 50% off input | 90% off input ($0.30/M vs $3.00/M on Claude Sonnet 4.6) |
| Cache write price | No premium | 1.25× base (5-min TTL) or 2.0× base (1-hour TTL) |
| TTL | 5-10 min (up to 1 hr off-peak) | 5 min standard, 1 hour extended |
| Hit rate guarantee | Best-effort ~50% | 100% when prefix matches exactly |
| Max breakpoints | 1 (automatic) | Up to 4 per request |
OpenAI prompt caching is fully automatic-zero code changes, 50% cost discount on cached tokens, roughly 50% hit rate (best effort, not guaranteed). Anthropic prompt caching is manual-you set cache_control breakpoints, get a 90% cost discount on cache reads, and a 100% guaranteed hit rate when configured correctly.
The Anthropic write premium trips people up. Anthropic prompt caching pricing comes down to three numbers: a 1.25× or 2× premium on cache writes depending on the TTL you pick, a 0.1× rate on cache reads, and the plain base input rate for everything after your last breakpoint. Those three multipliers, combined with per-model minimum cacheable token counts and a refresh rule that most teams misread, decide whether caching cuts your Claude bill by 80% or quietly makes it larger.
The break-even math is simpler than it looks. With the 5-minute TTL (1.25× write, 0.10× read), you save the write premium back after one cache hit. Hit once, you have already broken even. Every hit after that is pure savings.
The one prompt structure mistake that kills your hit rate
Automatic caching has no awareness of which parts of your prompt are stable versus dynamic. In production agentic systems, working memory, runtime context, and per-user variables sitting in the middle of the prompt change on every step, which leads to consistent cache misses on exactly the content that benefits most from caching.
The fix is ordering, not code. Order content most-to-least stable: tool definitions, then system prompt, then reference docs, then conversation history, then the live user query. Any change to a block invalidates that block and everything after it on Anthropic-so dynamic data must live at the very end.
The most instructive published case comes from ProjectDiscovery, whose security agent Neo runs 20 to 40-plus LLM steps per task on top of a 20,000-token system prompt.
Automatic caching is a smart default that Anthropic has made easy to adopt. Neo's scale and complexity required going further with explicit breakpoint placement and deliberate TTLs, which is what took them from single-digit hit rates to 84%.
One structural change-moving a single dynamic identifier from the middle of the prompt to the end-took the hit rate to 74% and cut the monthly inference bill 59%.
An agent like Beagle runs many repetitive inference calls against the same system prompt and tool schema. Prompt structure matters at that volume: a stable prefix that never varies by user or session is the whole game.
Where the hidden security risk lives
There is one non-obvious consequence of prompt caching that most teams never think about. Because cache hits are measurably faster than misses, they create a timing side-channel. Cache hits are measurably faster than cache misses. An attacker who can measure response latency could potentially infer whether a prompt prefix matches another user's cached prompt, frequency of certain system prompts in your deployment, and patterns that could enable prompt reconstruction through binary search. Research has shown attackers can distinguish hits vs. misses with statistical significance.
Prompt caching (both automatic and explicit) is ZDR eligible. Anthropic does not store the raw text of your prompts or Claude's responses. KV cache representations and cryptographic hashes of cached content are held in memory only and are not stored at rest. That is meaningful for compliance purposes, but the timing channel is separate from storage-and it is something shared-tenant deployments should factor into their threat model.
LLM prompt caching: common questions
How does prompt caching work technically?
Prompt caching stores the computational state from an LLM's attention layers so the model can skip redundant prefill work on repeated prompt prefixes. The result is lower time-to-first-token (TTFT) and cheaper input costs on every request that hits the cache for a shared prefix. The model still generates a new response-only the internal KV tensor computation is reused.
When does prompt caching not save money?
Prompt caching costs money when the cache write is never followed by a read-you paid the 1.25× write premium for nothing. A prefix you write once and read many times is enormously cheaper than re-sending it every call; a prefix you write but never read back (because it changes every request) is worse than not caching at all, since you paid the write premium for nothing. Low-traffic or highly dynamic prompts are the main failure modes.
What is the minimum prompt length to enable caching?
Automatic caching on OpenAI is enabled for prompts that are 1,024 tokens or longer, with cache hits occurring in increments of 128 tokens. When a request is made, the system checks if the initial portion (prefix) of your prompt is stored in the cache. Anthropic's minimum is also 1,024 tokens for most Claude models, with 2,048 tokens required for Haiku.
How do you check if your prompt cache is actually hitting?
The metric you want to track is the cached token fraction: the ratio of cache-read tokens to total input tokens, per request and as a rolling aggregate. Most LLM API responses include this data directly: Anthropic returns cache_read_input_tokens and cache_creation_input_tokens in the usage block; OpenAI returns prompt_tokens_details.cached_tokens in the response usage object.
What cache hit rate should I target?
A cache hit rate of 70%+ on stable-prompt workloads is achievable. Industry case studies show 84%+ is possible with disciplined prompt architecture. Under 30% on a workload with a fixed system prompt indicates a structural problem. If you are below that threshold with a stable system prompt, the first place to check is whether any dynamic content sits inside your cached prefix.