One real bug report from a developer's GitHub issue tells the story plainly: a 170,602-token system prompt was being written to Anthropic's cache on every single request, costing roughly $0.50 per message instead of the expected $0.05. The cause was a timestamp in the system prompt. Thirty-five dollars in one day when the budget said nine.
That is prompt caching working exactly as designed - and being defeated by one line.
This post explains the mechanism from the inside: what the model actually stores, how the math works, where it silently fails, and how to structure a prompt so the savings land.
What the model actually stores when it caches
When you send a prompt to an LLM, the model processes each token through its attention layers, generating key-value (KV) pairs - the model's internal representation of your input.
Prompt caching stores those KV tensors for stable prompt prefixes. When a subsequent request starts with the same prefix, the model skips re-computation and picks up where it left off.
The output is byte-identical. The model still generates fresh, non-deterministic responses. Caching only ever touches input-side tokens - output is never discounted.
Prompt caching is, precisely, a provider running prefix caching on their own hardware and charging a separate rate for the part you reuse. The hardware mechanism underneath is the same KV cache the GPU uses within a single request. The difference is that providers persist it across requests, key it by a hash of the token sequence, and charge you a fraction of the normal token price when a subsequent request hits.
When prefix caching is enabled, the framework caches fully computed KV blocks and indexes them by a chain hash. Each block's hash is computed over the parent block's hash, the block's token IDs, and optional extra keys. This chaining ensures that a cache hit requires an exact token-by-token match of the entire prefix up to and including the candidate block.
That last sentence is the thing most explainers skip past.
Why a single changed token kills the whole discount
A hit requires the prefix to be byte-identical to what was written. One changed character at position N invalidates everything after position N. No partial credit, no similarity matching.
This is not a quirk. It is a consequence of the hash chain. Each block is hashed together with its prefix history, so reuse is inherently prefix-sensitive. Once an early block differs, later blocks no longer line up with the cached chain even when the two prompts contain the same documents.
The practical consequence: the most common prompt caching failure is not a failure to cache - it is a failure to audit prompt stability first. Teams enable caching, then wonder why their hit rate stays below 10%. The cause is almost always the same: dynamic elements - timestamps, session metadata, rotating instructions - buried in the middle of the prompt, silently invalidating the cache on every request.
There is a subtler failure mode too. Because tools render first in Anthropic's request structure, changing a tool definition invalidates the system and messages caches behind it. If your tool list is built dynamically and the serialization order is non-deterministic (Go and Swift both randomize JSON key order by default), every request is a cold start.
How the pricing actually works - and what you pay on a cold start
The three providers structure this differently, and the differences matter at scale.
| Provider | Mechanism | Min tokens | Cache read discount | Write premium |
|---|---|---|---|---|
| Anthropic | Explicit cache_control breakpoints |
1,024 | 90% off (0.10×) | 1.25× (5 min TTL) or 2× (1 hr TTL) |
| OpenAI | Automatic, no code changes | 1,024 | 50% off | None |
| Google Gemini 2.5 | Implicit (automatic) | 1,024 | 75% off | None |
| DeepSeek | Automatic, 64-token chunks | None | 90% off | None |
Anthropic bills cached prompts in three token categories reported separately in the usage object: cache writes cost 1.25× the base input price for the default 5-minute TTL (2× for the 1-hour TTL), charged when content is first stored; cache reads cost 0.1× the base input price (a 90% discount); and uncached input - everything after your last breakpoint - bills at the standard 1× rate.
A common misreading is treating input_tokens as the whole prompt; it is only the uncached remainder, so a long-running agent can legitimately show 4,000 input tokens on a 200,000-token context.
The write premium means you do not save money on the very first request - you spend slightly more. The break-even math on Anthropic: 1 cached write + N cached reads = 1.25 + 0.1N, versus (N+1) uncached reads = N+1. Break-even is around 1.4 reads, so you need at least 2 cache reads per prefix to see net savings.
A hit rate above 30% on stable prompts means caching saves money. A hit rate above 60% means caching saves a significant amount of money.
On TTL: each breakpoint is ephemeral with a TTL of either 5 minutes (default) or 1 hour; the clock resets every time you read the cache, so a chatty conversation keeps its cache warm without re-paying the write cost. A sparse workload - support tickets arriving every 15 minutes - needs the 1-hour TTL or the cache expires between calls and you pay the write premium on every request.
Where prompt ordering becomes a structural decision
The placement rule is simple: put everything stable first, everything variable last. The order that produces the best hit rate is: system prompt (most stable), then tool definitions (stable per agent version). Retrieved documents and conversation history come after.
When OpenAI's engineering team outlined how they architected the Codex agent loop, they emphasized prompt structure as a first-class performance surface. In Codex CLI, system instructions, tool definitions, sandbox configuration, and environment context are kept identical and consistently ordered between requests to preserve long, stable prompt prefixes. The agent loop appends new messages rather than modifying earlier ones when runtime configurations change mid-conversation.
That last point is worth dwelling on. Most agent loops accumulate a conversation history. If you rewrite an earlier turn to compress it, you break the prefix and pay full price on the next call. Appending preserves the cache; editing destroys it.
Agentic workflows burn through far more tokens than a single question and answer, because the model re-reads the accumulating conversation on every turn of the loop. A ten-turn agent session with a 5k token system prompt reprocesses that same prompt ten separate times if nothing is cached. That is the use case where caching earns back its write premium fastest.
The average prompt token count grew nearly 4× between early 2024 and late 2025, from roughly 1,500 tokens to 6,000 per request. Longer prompts make caching more valuable, not less.
A Beagle-style teammate sitting in Slack handles a narrow version of this naturally - the system prompt and tool definitions it sends on every call stay identical across requests, so the stable prefix is large and the cache hit rate is high from the first week of real traffic.
Prompt caching: common questions
What is prompt caching in an LLM API?
Prompt caching stores the computational state from an LLM's attention layers so the model can skip redundant prefill work on repeated prompt prefixes. The result is lower time-to-first-token and cheaper input costs on every request that hits the cache for a shared prefix. It has no effect on output quality or token generation speed.
Does prompt caching change the model's output?
No. Cache reads bill at 0.10× base input on Anthropic and on OpenAI's newer models. The model recomputes nothing it has already seen, so output is unchanged. The model generates a fresh response from the point where the cached prefix ends.
Why is my cache hit rate still 0% after enabling caching?
The cause is almost always the same: dynamic elements - timestamps, session metadata, rotating instructions - buried in the middle of the prompt, silently invalidating the cache on every request. Also check that your prompt exceeds the minimum token threshold (1,024 for both Anthropic and OpenAI) and that tool definitions are serialized in a stable order.
How does Anthropic prompt caching differ from OpenAI?
Anthropic requires explicit cache_control breakpoints in your API request and charges a 1.25× write premium but gives a 90% discount on reads.
OpenAI has no explicit breakpoints; the system detects the longest matching prefix automatically. The Anthropic model is more controllable; the OpenAI model is less work to wire up. For long, stable system prompts and tool definitions, both produce similar effective discounts.
Is cached content stored securely?
Anthropic does not store the raw text of your prompts or Claude's responses. KV cache representations and cryptographic hashes of cached content are held in memory only and are not stored at rest. Caches are also isolated at the workspace level, not shared across accounts.