Prompt Caching Cuts LLM Costs 90%. Here's Why Most Teams See 5%.

Prompt caching can reduce input token costs by up to 90% on Claude and OpenAI. But most production apps see 5-15%. The gap isn't a bug - it's a structural mistake in how prompts are built.

Cover art for Prompt Caching Cuts LLM Costs 90%. Here's Why Most Teams See 5%.

A support agent at a mid-sized SaaS company sends 200 Claude API calls an hour. Each call includes a 4,000-token system prompt describing tone, escalation rules, and a product FAQ. None of those tokens change between calls. But without prompt caching configured correctly, the model reprocesses every one of them, every time. That's 800,000 tokens of pure duplicate work - per hour.

Prompt caching exists specifically to eliminate this. It works by storing the model's computed state for the parts of your prompt that do not change, so those tokens are reused across requests instead of recomputed. The savings on paper are dramatic. Anthropic's prompt caching reduces costs by up to 90% and latency by up to 85% for long prompts. OpenAI achieves 50% cost reduction with automatic caching enabled by default.

In practice, most teams see a fraction of that. Developers see cache_creation_input_tokens in the response, but their bill barely moves. Anthropic says 70-90% savings. Most developers are seeing 5-15%. The gap is structural, not a bug.

This post explains what is actually happening inside the model, why the gap exists, and what changes close it.

What the model is actually storing

Prompt caching works by reusing the attention key-value (KV) cache across requests. A prompt is cached by storing the prompt's attention KV cache. If a subsequent prompt has a matching prefix with a cached prompt, the KV cache for the matching prefix can be retrieved from the cache.

What that means in practice: every token you send gets processed through every attention layer of the model, producing a key vector and a value vector at each layer. Each token produces a Key vector and a Value vector at every attention layer. For a model like Opus, that's hundreds of layers, each producing vectors of thousands of dimensions. A 100K-token prompt might produce a KV cache of 500MB-1GB per request.

The provider stores that enormous block of GPU-resident data so the next request can skip the prefill phase entirely if the prefix matches. That's why there's a 25% surcharge on cache writes - you're paying for VRAM allocation, not just compute. And it's why there's a minimum token threshold to trigger caching: 1,024 tokens for Sonnet and Haiku, up to 2,048-4,096 for Opus.

OpenAI offers automatic prompt caching on GPT-4o and newer models, where caching activates automatically for prompts exceeding a minimum token threshold, with cache hits occurring only for exact prefix matches. Anthropic provides developer-controlled caching through explicit cache breakpoints, allowing users to specify which portions of their prompt should be cached, with configurable TTL options. Google offers both implicit caching, which activates automatically with no guaranteed cost savings, and explicit context caching, where developers create and reference caches with guaranteed discounts.

The real cost math, by provider

Implementation details such as minimum token thresholds (typically 1,024-4,096 tokens depending on model), TTL durations (ranging from 5 minutes to 24 hours), and pricing structures vary across providers. Here is a direct comparison built from current pricing pages:

Provider Cache read discount Cache write surcharge Default TTL Minimum prefix
Anthropic (Claude Sonnet) 90% off input +25% on write 5 min (1hr available) 1,024 tokens
OpenAI (GPT-5.6+) ~50% off input +25% on write ~1 hour 1,024 tokens
Google (Gemini, explicit) ~75% off input Cache creation cost Configurable 32,768 tokens

Cache reads on Anthropic run at $0.30/M tokens vs $3.00/M fresh.

Because a write costs more than a plain input token but a read costs far less, caching pays off after a small number of hits. On the 5-minute TTL, two hits average a 32.5% saving, three hits 52%, and the asymptote is 90%.

One non-obvious number: caching shifts cost, it does not remove it - output tokens still bill at full rate, so an agent that emits long artifacts benefits less than the 90% headline implies. If half your spend is output tokens (a ratio of 5:1 per token is common on Claude Sonnet), the realistic ceiling on total-bill savings is closer to 45%, not 90%.

90%max cache read discountAnthropic, input tokens only
5-15%actual savings most devs seebefore fixing prompt structure
59-70%real reduction at ProjectDiscoveryafter one architectural change

Why the cache breaks - and why it's almost always your fault

When a request is processed, the system checks whether the prompt prefix matches previously cached content. A cache hit occurs when the entire prefix matches exactly, allowing the system to reuse previously computed KV tensors. A cache miss occurs when any token differs from the cached content, even at the very beginning, forcing complete recomputation of all tokens.

This is where teams get tripped up. The most common cache-breakers:

  • A timestamp or session ID near the top of the system prompt. Even adding a timestamp or a session ID at the beginning of the prompt breaks the cache for everything that follows.

  • Reordering tool definitions between calls. Tool definitions count as part of the prefix - adding or reordering a single tool invalidates the cache for every subsequent turn.

  • Context compaction. Summarization, compaction, or context truncation can change the prefix and reset cache reuse. One study found that shrinking tool output by 38.4% actually increased billed costs by 6.8% because the compaction invalidated cache hits and forced re-runs.

  • Dynamic content anywhere before the stable block. Put stable developer instructions and shared reference material first. If instructions contain timestamps, user-specific content, or other dynamic content, place those at the end rather than the beginning.

The fix is structural, not technical. Put stable content - system instructions, reference docs, tool schemas - at the top. Put dynamic content - user name, date, session-specific data - at the bottom or in the user turn.

ProjectDiscovery's Neo agent, which runs 20-40+ LLM steps per task over a 20,000-token system prompt, cut LLM costs 59% and lifted its cache hit rate from 7% to 74% in a single deployment simply by relocating dynamic working memory out of the cacheable prefix, ultimately reaching 84% over ten days.

Beagle in action#ai-ops, 10:32am
The ask
engineer notices Claude API spend doubled month-over-month despite no new features
Beagle drafts
pulls the last 30 days of token usage from the linked billing doc, drafts a breakdown showing 94% of spend is uncached input tokens, flags that the system prompt contains a per-request session_timestamp field at line 1
You approve
you approve the summary; it posts to the channel with the fix highlighted - move dynamic fields below the first cache breakpoint
Do this in your workspace →

How to actually verify your cache is working

Providers expose cache metrics in the API response. On Anthropic, look for cache_creation_input_tokens (the write) and cache_read_input_tokens (the hit) in the usage object. The cache's default minimum lifetime is 5 minutes. This lifetime is refreshed each time the cached content is used.

If you see cache_creation_input_tokens but no cache_read_input_tokens after the first call, you have a structure problem. Check:

  1. Is the shared prefix actually identical across calls? Diff two consecutive raw request bodies.
  2. Is the prefix above the 1,024-token minimum before your first breakpoint?
  3. Are requests arriving within the TTL window (5 minutes by default, 1 hour if you've enabled extended TTL)?
  4. Are tool definitions in a fixed order in every request?

Prompt caching, batching, and related features do not behave identically across the first-party Anthropic API, Amazon Bedrock, and Google Cloud Vertex AI. On Bedrock and Vertex AI, prompt caching supports explicit breakpoints but not automatic top-level caching. If you're routing through a cloud provider rather than hitting Anthropic directly, the behavior differs.

Multi-turn support agent, same 4,000-token system prompt each call
Without Beagle
every call reprocesses the full context from scratch - 200 calls/hour at $3/M tokens = $2.40/hr on input alone, $57/day
With Beagle
system prompt cached after turn 1, subsequent calls hit cache at $0.30/M - same volume costs $0.24/hr, $5.76/day

Prompt caching: common questions

What is prompt caching in LLMs?

Prompt caching stores the model's internal key-value computation state for a fixed prefix of your prompt. When the next request starts with the same prefix, the model skips reprocessing those tokens. The result is lower time-to-first-token and cheaper input billing - but only when the prefix matches exactly.

Does prompt caching work automatically or do I have to configure it?

It depends on the provider. OpenAI offers automatic prompt caching on GPT-4o and newer models, where caching activates automatically for prompts exceeding a minimum token threshold.

Anthropic provides developer-controlled caching through explicit cache breakpoints, allowing users to specify which portions of their prompt should be cached. With Anthropic you must add cache_control markers yourself.

Why is my cache hit rate low even though I enabled caching?

The most common cause is dynamic content near the start of your prompt - a timestamp, session ID, or user name placed before the stable system instructions. Any whitespace change invalidates the cached prefix. Tool definition order must be fixed in every request. Move all dynamic fields after the first cache breakpoint, or into the user turn.

Does prompt caching reduce output token costs too?

No. Caching shifts cost, it does not remove it - output tokens still bill at full rate. The discount applies only to the cached portion of input tokens. If your workload is output-heavy, the total bill reduction will be well below the 90% headline.

What breaks the cache in a multi-turn agentic workflow?

The cache is a strict prefix cache, so order and stability are everything. Tool definitions count as part of the prefix - adding or reordering a single tool invalidates the cache for every subsequent turn. The same holds for the system prompt, images, and any mid-conversation modification. In agentic systems, context compaction is a particularly common culprit because it rewrites earlier turns that were already cached.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle