LLM Prompt Caching Cuts Input Costs Up to 90%-Here's the Mechanism

LLM prompt caching reuses transformer KV attention states across requests, cutting input token costs by 50-90% and time-to-first-token by up to 85%. Here's exactly how it works and what kills your hit rate.

Cover art for LLM Prompt Caching Cuts Input Costs Up to 90%-Here's the Mechanism

The average prompt token count in production applications grew nearly 4x between early 2024 and late 2025, from roughly 1,500 tokens to 6,000 per request. Every one of those extra tokens costs money to process-and if your system prompt is identical on every call, you are paying to recompute the same math thousands of times a day. LLM prompt caching is the mechanism that stops that. Not by returning a cached answer, but by skipping redundant GPU work inside the model itself.

What LLM prompt caching actually stores

Prompt caching works by reusing the key-value (KV) tensors that a transformer generates during attention computation-not by caching text or responses.

When an LLM processes your prompt, it generates key-value (KV) cache entries in its attention layers-mathematical representations of the relationships between tokens. Normally, the model recomputes this KV cache on every request. Prompt caching stores it so the model can skip that computation on subsequent requests that share the same prefix. The model still generates a fresh response every time; it is the redundant prefill work that gets cut.

That last point is worth dwelling on. It is worth separating this from semantic caching, which stores full prompt-response pairs and returns a saved answer for similar queries. Prompt caching does not reuse outputs. You still get a fresh, potentially different response each time, because only the intermediate computation is cached, not the final result.

The practical constraint is prefix matching. Inference providers check if a prefix of a prompt matches a cached entry upon receiving a request. If a match is found, the provider can reuse the cached KV attention states for the matching prefix and only compute the attention states for the unique tokens in the prompt.

A single character change anywhere in the cached prefix invalidates the match, so exact prefix matches matter more here than in most caching you have worked with.

Providers hold on to these matrices for each prompt for 5-10 minutes after the request is made, and if you send a new request that starts with the same prompt, they reuse the cached K and V rather than recalculating them.

How Anthropic and OpenAI implement it differently

Both providers offer prompt caching natively, but the mechanics-and the economics-differ enough that choosing the wrong one for your workload costs real money.

OpenAI Anthropic
Activation Automatic, no code change Explicit cache_control breakpoints in your API payload
Minimum prefix 1,024 tokens 1,024 tokens (Sonnet/Opus), 2,048 tokens (Haiku)
Cache read price 50% off input 90% off input ($0.30/M vs $3.00/M on Claude Sonnet 4.6)
Cache write price No premium 1.25× base (5-min TTL) or 2.0× base (1-hour TTL)
TTL 5-10 min (up to 1 hr off-peak) 5 min standard, 1 hour extended
Hit rate guarantee Best-effort ~50% 100% when prefix matches exactly
Max breakpoints 1 (automatic) Up to 4 per request

OpenAI prompt caching is fully automatic-zero code changes, 50% cost discount on cached tokens, roughly 50% hit rate (best effort, not guaranteed). Anthropic prompt caching is manual-you set cache_control breakpoints, get a 90% cost discount on cache reads, and a 100% guaranteed hit rate when configured correctly.

The Anthropic write premium trips people up. Anthropic prompt caching pricing comes down to three numbers: a 1.25× or 2× premium on cache writes depending on the TTL you pick, a 0.1× rate on cache reads, and the plain base input rate for everything after your last breakpoint. Those three multipliers, combined with per-model minimum cacheable token counts and a refresh rule that most teams misread, decide whether caching cuts your Claude bill by 80% or quietly makes it larger.

The break-even math is simpler than it looks. With the 5-minute TTL (1.25× write, 0.10× read), you save the write premium back after one cache hit. Hit once, you have already broken even. Every hit after that is pure savings.

90%cost reduction on cache readsAnthropic, vs standard input price
85%latency reductiontime-to-first-token on long prompts (Anthropic claim)
4×prompt length growthavg production request, early 2024 to late 2025
59%LLM cost cutProjectDiscovery, single architectural change

The one prompt structure mistake that kills your hit rate

Automatic caching has no awareness of which parts of your prompt are stable versus dynamic. In production agentic systems, working memory, runtime context, and per-user variables sitting in the middle of the prompt change on every step, which leads to consistent cache misses on exactly the content that benefits most from caching.

The fix is ordering, not code. Order content most-to-least stable: tool definitions, then system prompt, then reference docs, then conversation history, then the live user query. Any change to a block invalidates that block and everything after it on Anthropic-so dynamic data must live at the very end.

The most instructive published case comes from ProjectDiscovery, whose security agent Neo runs 20 to 40-plus LLM steps per task on top of a 20,000-token system prompt.

Automatic caching is a smart default that Anthropic has made easy to adopt. Neo's scale and complexity required going further with explicit breakpoint placement and deliberate TTLs, which is what took them from single-digit hit rates to 84%.

One structural change-moving a single dynamic identifier from the middle of the prompt to the end-took the hit rate to 74% and cut the monthly inference bill 59%.

An agent like Beagle runs many repetitive inference calls against the same system prompt and tool schema. Prompt structure matters at that volume: a stable prefix that never varies by user or session is the whole game.

Beagle in action#product-team, mid-afternoon
The ask
'can you summarize the feedback from the last three user interviews?'
Beagle drafts
loads a 6,000-token context doc (already cached from earlier calls that hour), runs inference only over the new question tokens
You approve
time-to-first-token is ~85% faster than a cold start; the cached prefix costs $0.30/M instead of $3.00/M
Do this in your workspace →

Where the hidden security risk lives

There is one non-obvious consequence of prompt caching that most teams never think about. Because cache hits are measurably faster than misses, they create a timing side-channel. Cache hits are measurably faster than cache misses. An attacker who can measure response latency could potentially infer whether a prompt prefix matches another user's cached prompt, frequency of certain system prompts in your deployment, and patterns that could enable prompt reconstruction through binary search. Research has shown attackers can distinguish hits vs. misses with statistical significance.

Prompt caching (both automatic and explicit) is ZDR eligible. Anthropic does not store the raw text of your prompts or Claude's responses. KV cache representations and cryptographic hashes of cached content are held in memory only and are not stored at rest. That is meaningful for compliance purposes, but the timing channel is separate from storage-and it is something shared-tenant deployments should factor into their threat model.

Running a support agent with a 6,000-token system prompt
Without Beagle
every request recomputes the full prompt KV cache; at 2,000 calls/day on Claude Sonnet 4.6, that is $36/day in input tokens for the system prompt alone
With Beagle
at 80% hit rate, the same prompt costs ~$8/day; the first call per session pays the write, all subsequent calls read at 10% price

LLM prompt caching: common questions

How does prompt caching work technically?

Prompt caching stores the computational state from an LLM's attention layers so the model can skip redundant prefill work on repeated prompt prefixes. The result is lower time-to-first-token (TTFT) and cheaper input costs on every request that hits the cache for a shared prefix. The model still generates a new response-only the internal KV tensor computation is reused.

When does prompt caching not save money?

Prompt caching costs money when the cache write is never followed by a read-you paid the 1.25× write premium for nothing. A prefix you write once and read many times is enormously cheaper than re-sending it every call; a prefix you write but never read back (because it changes every request) is worse than not caching at all, since you paid the write premium for nothing. Low-traffic or highly dynamic prompts are the main failure modes.

What is the minimum prompt length to enable caching?

Automatic caching on OpenAI is enabled for prompts that are 1,024 tokens or longer, with cache hits occurring in increments of 128 tokens. When a request is made, the system checks if the initial portion (prefix) of your prompt is stored in the cache. Anthropic's minimum is also 1,024 tokens for most Claude models, with 2,048 tokens required for Haiku.

How do you check if your prompt cache is actually hitting?

The metric you want to track is the cached token fraction: the ratio of cache-read tokens to total input tokens, per request and as a rolling aggregate. Most LLM API responses include this data directly: Anthropic returns cache_read_input_tokens and cache_creation_input_tokens in the usage block; OpenAI returns prompt_tokens_details.cached_tokens in the response usage object.

What cache hit rate should I target?

A cache hit rate of 70%+ on stable-prompt workloads is achievable. Industry case studies show 84%+ is possible with disciplined prompt architecture. Under 30% on a workload with a fixed system prompt indicates a structural problem. If you are below that threshold with a stable system prompt, the first place to check is whether any dynamic content sits inside your cached prefix.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle