How Does LLM Prompt Caching Actually Work?

Prompt caching cuts input token costs by up to 90% - but a single timestamp in your system prompt can collapse that saving to 6%. Here's the mechanics, plainly explained.

Cover art for How Does LLM Prompt Caching Actually Work?

Put a timestamp in your system prompt and Claude's 80% cost saving drops to 6%. That one fact tells you almost everything you need to know about how prompt caching actually works - and why most teams enabling it see smaller savings than the docs promise.

Here is what is happening under the hood, and what you need to do about it.

What prompt caching actually stores

Every API call to a large language model normally reprocesses your whole prompt from scratch. That prompt usually includes system instructions, tool definitions, and your full chat history - and in most production apps, most of that content stays identical from one request to the next.

Prompt caching operates on the key-value (KV) cache - the same data structure that makes transformer decoding tractable. During the prefill phase, every prompt token is processed and the key (K) and value (V) matrix projections from each attention layer are computed and stored. During decode, each newly generated token attends to those cached K/V matrices rather than recomputing the projections for every previous token.

Provider-level prompt caching extends that mechanism across requests. When a new request shares an identical prefix with a recent one, the provider reuses the already-computed KV state for that prefix instead of running prefill again. You pay a reduced "cache read" rate for those tokens, and time-to-first-token drops because the expensive prefill is skipped.

One thing worth being clear about: the model still generates a fresh response every time - it's the redundant prefill work that gets cut. Output tokens are never discounted. Only input tokens that fall inside the cached prefix get the cheaper rate.

The main constraint is prefix matching. Prompt caching works by comparing the beginning of your current prompt against what is already cached. If the cached prefix and your new prompt are exactly identical - token-for-token - up to a certain point, the model reuses the cached computation for that portion and only processes new tokens from where the match ends. A single token change anywhere in the prefix breaks the match from that point forward.

That last sentence is why the timestamp kills you.

90%input token discountAnthropic cache reads (0.10× base)
50%input token discountOpenAI automatic caching
6%actual savingwhen a timestamp sits in the system prompt (Claude)

How the three major providers differ

Major LLM providers have implemented prompt caching with varying approaches. OpenAI offers automatic prompt caching on GPT-4o and newer models, where caching activates automatically for prompts exceeding a minimum token threshold, with cache hits occurring only for exact prefix matches. Anthropic provides developer-controlled caching through explicit cache breakpoints, allowing users to specify which portions of their prompt should be cached, with configurable time-to-live (TTL) options. Google offers both implicit caching, which activates automatically with no guaranteed cost savings, and explicit context caching, where developers create and reference caches with guaranteed discounts.

Implementation details such as minimum token thresholds (typically 1,024-4,096 tokens depending on model) and TTL durations (ranging from 5 minutes to 24 hours) vary across providers and are subject to change.

Here is how the pricing stacks up across the two APIs most teams use:

Anthropic (Claude Sonnet) OpenAI (GPT-4o)
Base input cost $3.00 / 1M tokens $2.50 / 1M tokens
Cache write cost $3.75 / 1M tokens $2.50 / 1M tokens
Cache read cost $0.30 / 1M tokens $1.25 / 1M tokens
Read discount 90% 50%
Opt-in required? Yes - explicit cache_control markers No - automatic
Min prefix length 1,024 tokens 1,024 tokens
Default TTL 5 minutes ~60 minutes

Anthropic offers deeper savings (90% vs 50%) but requires explicit implementation. For agents with stable, structured prompts, Anthropic's explicit caching typically delivers 2x more savings than OpenAI's automatic approach.

The trade-off is operational complexity. On OpenAI, caching is fully automatic. On Anthropic, you opt in by marking which blocks of your prompt you want cached via cache_control breakpoints.

Why agents are where this matters most

A single-turn chatbot with a 500-token system prompt sees modest benefit from caching. An agent loop with 10 tool calls is where the economics get dramatic.

Agentic workflows - the kind that loop through a plan, call a tool, read the result and call the next tool - burn through far more tokens than a single question and answer, because the model re-reads the accumulating conversation on every turn of the loop. A ten-turn agent session with a 5k-token system prompt reprocesses that same prompt ten separate times if nothing is cached.

In Anthropic's own measurements, prompt caching was the largest cost lever by a wide margin: it cut agent-loop cost by a factor of 2.7 to 5.3 on benchmark tasks, and cut a small triage agent's bill by 83%, or 88% with input trimming added.

The structure of an agent prompt is also naturally suited to caching. In a typical agentic workflow, the system prompt and tool definitions can easily exceed 10,000 tokens - the instructions on how to behave, the documentation for the APIs it can call, and the examples of how to format its output. Without caching, every single turn of the reasoning loop requires the model to re-process those 10,000 tokens.

On turn N, everything except the latest message is identical to turn N−1. In an agent loop with 10-15 internal tool calls before responding to the user, you send 10-15 requests where almost the entire prompt is a cache hit.

Beagle in action#engineering-ai-costs, 2:47pm
The ask
'our Claude spend tripled this sprint - agent loop is hammering the API'
Beagle drafts
checks the thread, finds the team's system prompt includes a dynamic {timestamp} field injected at each turn, drafts a reply explaining the cache invalidation pattern and a fix
You approve
you approve; the context posts with a one-line code change and a rough estimate of the savings it unlocks
Do this in your workspace →

The antipatterns that silently kill your cache hit rate

This is where most teams leave money on the table. The provider dashboard shows caching is "enabled." The bill barely moves.

On a 30-turn agent loop, prompt caching saves 80-89% depending on the provider - not the full advertised 90%. Put a timestamp in the system prompt and Claude's saving falls to 6%. A volatile tool definition can make caching cost more than not caching at all.

The mechanics are unforgiving: any modification to the cached portion of your prompt - even injecting a timestamp - causes a complete cache miss.

The fix is structural. The most important rule is to prioritize static material over dynamic content. Put the system prompt, tool definitions, and reference material at the head of the request, and the section that changes with each call - the user's latest message - at the very end.

The canonical layer order for an agent:

  • Layer 1 - tool definitions (rarely change within a session) → cache
  • Layer 2 - static agent instructions (never change) → cache
  • Layer 3 - background documents or knowledge base → cache if long
  • Layer 4 - session context or user profile → no cache (changes per user)
  • Layer 5 - conversation history → no cache (grows each turn)
  • Layer 6 - current user message → always fresh, always last

Research across 500+ agent sessions with 10,000-token system prompts found that prompt caching reduces API costs by 45-80% and improves time-to-first-token by 13-31% across providers. Strategic prompt cache block control - placing dynamic content at the end of the system prompt, avoiding dynamic function calling, and excluding dynamic tool results - provides more consistent benefits than naive full-context caching, which can paradoxically increase latency.

One more non-obvious risk: the timing side-channel. Your agent's 10,000-token system prompt, containing business logic, tool credentials, and RAG instructions, sits in a shared cache. A co-tenant on the same API endpoint can probe it token by token using nothing but response latency. It is a real attack surface, documented in peer-reviewed research accepted at ICML 2025. For most internal tooling it is a theoretical concern; for anything customer-facing that embeds credentials or proprietary logic in the system prompt, it is worth knowing about.

Agent loop, 10 turns, 5k-token system prompt
Without Beagle
model reprocesses 5,000 system prompt tokens on every turn - 50,000 input tokens just for context re-reads, billed at full rate each time
With Beagle
system prompt cached after turn 1; turns 2-10 pay cache-read rate on those tokens, cutting the context-read bill by up to 90% and shaving latency on every reply

Prompt caching: common questions

What is prompt caching in LLMs?

Prompt caching stores the key-value attention tensors computed during the prefill phase of an LLM request. When a subsequent request shares an identical prefix, the provider reuses those tensors instead of recomputing them. The result is a steep discount on cached input tokens - up to 90% on Anthropic, 50% on OpenAI - and lower time-to-first-token.

Does prompt caching change the model's output?

No. The model recomputes nothing it has already seen, so output is unchanged. The cache only touches input-side tokens during prefill. The decode phase - where the model generates its response - runs identically whether or not the input was cached.

Why isn't my prompt caching saving anything?

The most common cause is a dynamic value - a timestamp, a request ID, a randomly-ordered tool list - sitting inside the part of your prompt you expect to be cached. If the cached prefix and your new prompt are exactly identical token-for-token up to a certain point, the model reuses the cached computation. A single token change anywhere in the prefix breaks the match from that point forward. Move all dynamic content to the end of the prompt, after your static system instructions and tool definitions.

How long does a prompt cache last?

Anthropic's default cache lasts 5 minutes and refreshes for free every time it is read, so a steady stream of requests keeps it alive indefinitely. A 1-hour cache is available for double the write cost, suited to workloads with gaps longer than 5 minutes - slower agent loops, human-in-the-loop approval steps, or users who take their time responding. OpenAI's automatic caching has a longer default window of around 60 minutes.

Is prompt caching worth it for short prompts?

A 2,000-token prompt producing a 1,500-token answer spends most of its cost on output, not input. Caching the whole prompt perfectly saves under 20% of that call. The ROI is front-loaded toward long system prompts, large tool definitions, and multi-turn agent loops where the same stable prefix gets sent many times. For short, one-off requests, the write cost on Anthropic can actually exceed the read savings if the cache never gets hit a second time.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle