Understand Prompt Caching Before Your Agent Bill Doubles

Prompt caching stores the KV tensors behind repeated prompt prefixes so your LLM skips redundant computation. Here's how it works under the hood-and what quietly kills your hit rate.

Cover art for Understand Prompt Caching Before Your Agent Bill Doubles

One task at ProjectDiscovery ran 67.5 million input tokens across 1,225 steps at a 91.8% cache rate. A nearly identical task with a 3.2% cache rate cost roughly 60 times more. Same model. Same token volume. The only difference was prompt structure.

That gap is what prompt caching is about. Not a new model, not a smaller context window-just whether the API recomputes tokens it already computed sixty seconds ago.

What the KV cache actually stores

When an LLM processes your prompt, it generates key-value (KV) cache entries in its attention layers-mathematical representations of the relationships between tokens. Normally, the model recomputes this KV cache on every request. Prompt caching stores it so the model can skip that computation on subsequent requests that share the same prefix. The model still generates a fresh response every time; it's the redundant prefill work that gets cut.

This matters because of where the latency actually lives. Every LLM request goes through two latency phases: time to first token (TTFT), which measures how long the model takes to start responding, and time to last token (TTLT), which captures the full generation time. Both get worse as your prompts get longer. A long system prompt increases TTFT because the model processes every token through its attention mechanism before producing any output.

You send a request with a 2,000-token system prompt, and the model reprocesses every token. Send the next turn, and it reprocesses them again. Twenty turns in, you have paid to compute the same static instructions twenty times over. Prompt caching fixes this by storing the model's computed state for the parts of your prompt that do not change, so those tokens are reused across requests instead of recomputed.

One thing worth internalizing: prompt caching behaves differently from a standard cache in one crucial way. It does not return a stored answer. It reuses the internal computation for a shared prompt prefix, then the model still generates a new response token by token. If you expect it to return a saved reply, you're thinking of semantic caching-a different layer entirely.

How Anthropic and OpenAI implement it differently

OpenAI prompt caching is fully automatic-zero code changes, 50% cost discount on cached tokens, roughly 50% hit rate (best effort, not guaranteed). Anthropic prompt caching is manual-you set cache_control breakpoints, get a 90% cost discount on cache reads, and a 100% guaranteed hit rate when configured correctly. Both require a minimum of 1,024 tokens before a prefix is eligible.

OpenAI Anthropic
Activation Automatic Manual cache_control breakpoints
Discount on cache read ~50-90% (model-dependent) 90% ($0.30/M vs $3.00/M)
Write premium Free through GPT-5.5; 1.25× from GPT-5.6 1.25× (5-min TTL) or 2× (1-hr TTL)
Hit guarantee Best-effort (~50%) 100% on exact prefix match
Cache TTL 5-10 min inactivity 5 min default; 1 hr available
Min token threshold 1,024 tokens 1,024 (Sonnet/Opus); 2,048 (Haiku)

Anthropic's design assumes you know your workload and want maximum savings on it. OpenAI's design assumes you'd rather not think about it. If your traffic is high-frequency and your prompts are large and stable, Anthropic's 90% discount will dominate the comparison; if your traffic is sporadic or your prompts evolve frequently, OpenAI's automatic 50% with no write premium is the cleaner choice.

On the TTL side, the cache's default minimum lifetime is 5 minutes, refreshed each time the cached content is used. If you find that 5 minutes is too short, Anthropic also offers a 1-hour cache TTL. The lifetime is measured from the start of the request that writes or reads the cache entry, not from the end of its response. Time spent generating a response counts against the lifetime, so the window for a follow-up request to reuse the cache is the lifetime minus the generation time.

That last point is non-obvious and costs teams money. A model that takes 90 seconds to generate a long response has consumed 90 seconds of a 5-minute TTL before the next turn even starts.

91.8%cache hit rate (ProjectDiscovery, post-fix)up from 3.2% before prompt restructure
90%cost reduction on cached reads (Anthropic)$0.30/M vs $3.00/M standard rate
5 mindefault cache TTLcounts from request start, not response end
1,024minimum tokens before caching activatesboth OpenAI and Anthropic (most models)

The prompt ordering rule that changes everything

The cache is a prefix cache. Cache matching is all or nothing: a changed character, timestamp, tool serialization, or rewritten earlier turn can cause a complete miss.

The practical design goal is to place everything stable-tool definitions, system instructions, retrieved documents, few-shot examples-as early in the prompt as possible, and to push everything that varies per request-the user's specific question, a session ID, a live timestamp-as late as possible, ideally after the last cache breakpoint.

Static function schemas are excellent caching candidates because they tend to be large and unchanged across calls. Place them in the cached prefix ahead of dynamic content, and they contribute to your prefix match on every request.

The most common way teams destroy their own hit rate is by letting dynamic content creep into the prefix. Timestamps and session IDs in the prefix destroy cache performance. Injecting something like "Today is March 6, 2026" into a system prompt invalidates the cache every day. Use a precise timestamp and it invalidates on every request.

At ProjectDiscovery, two tasks with nearly identical token volume showed a 91.8% cache rate versus 3.2%-roughly 60× the cost difference. The latter ran before the optimization rollout. Their fix: moving dynamic working memory out of the system prompt and into a trailing user message. One structural change. Cache hit rate jumped from 7% to 84%, overall LLM spend dropped 59%.

Beagle in action#eng-ai, Tuesday morning
The ask
'why is our Claude bill up 40% this month? system prompt hasn't changed'
Beagle drafts
pulls the last 7 days of usage logs from the linked cost dashboard, spots that cache_read_input_tokens dropped to near zero after a deploy on Thursday
You approve
you approve the reply-"timestamp injection in the new deploy broke the prefix; rolling back the system prompt order should restore cache hits"-and the team has a root cause in two minutes, not two hours
Do this in your workspace →

Where caching pays and where it quietly hurts

Prompt caching is not universally beneficial. A write premium makes reuse volume decisive; without one, caching has no downside from the first request.

For Anthropic's 5-minute TTL, the write cost is 1.25× standard. With the 5-minute TTL, you save the write premium back after one cache hit. Hit once, you've already broken even. Every hit after that is pure savings. For the 1-hour TTL at 2× write cost, you break even after two cache hits.

Break-even lands at 2.3 reuses of the same cached prefix within the one-hour TTL window. Any workload where the same system prompt or tool definitions are sent more than twice per hour is already in the money.

Where it hurts: one-shot queries that never repeat within the TTL window will pay the write premium with no offsetting reads. For most developers starting out, OpenAI's automatic caching is the right call-it requires no code changes and immediately applies to any prompt over 1,024 tokens. If you're at significant scale (500K+ tokens/day in static context), the jump to Anthropic's 90% discount justifies the implementation overhead.

Agentic workloads are the strongest case. Agentic workflows that loop through a plan, call a tool, read the result, and call the next tool burn through far more tokens than a single question-and-answer, because the model re-reads the accumulating conversation on every turn. A ten-turn agent session with a 5,000-token system prompt reprocesses that same prompt ten separate times if nothing is cached.

Without caching, a 20,000-token system prompt would be processed fully for every step in a 100-request workflow, resulting in 2 million processed tokens. With caching, the system pays the full input cost only for the first request. Later requests treat the static prefix as cached input, priced at a fraction of the standard rate. This approach results in cost reductions of approximately 89% for extended sessions, making autonomous operation economically viable.

Running a 50-step agent task with a 10,000-token system prompt
Without Beagle
every step reprocesses all 10,000 tokens at full price-500,000 input tokens billed at standard rates before the agent produces a single output
With Beagle
after the first step writes the cache, 49 steps read from it at 10% the cost-effective input bill drops ~89%, and time-to-first-token shrinks by hundreds of milliseconds per step

One more non-obvious risk: cached data is never shared between businesses, even when two different customers send byte-identical prompts. Anthropic and OpenAI both certify their caching as zero data retention (ZDR) eligible-raw prompt text isn't stored at rest, only the computed key-value representations are held in memory, and they're destroyed once the TTL expires. That's worth knowing before security teams flag it.

A teammate like Beagle, which reads and replies inside Slack with a fixed system prompt on every request, is a natural fit for prefix caching-the tool definitions and behavioral instructions are identical across every channel message, and the only thing that changes is the user's text at the end.

Prompt caching LLM: common questions

What is prompt caching in LLMs?

Prompt caching stores the key-value tensors from an LLM's attention layers after processing a repeated prompt prefix. When the next request starts with the same prefix, the model skips that computation entirely. It still generates a fresh response-only the prefill work is saved. Both OpenAI and Anthropic support it natively via their APIs.

How much does prompt caching save?

Anthropic's prompt caching reduces costs by up to 90% and latency by up to 85% for long prompts. OpenAI achieves 50% cost reduction with automatic caching enabled by default. Real production numbers land between 41% and 80% cost reduction across providers, depending on how much of each prompt is stable prefix versus dynamic content.

Why is my prompt cache hit rate low?

The most common cause is dynamic content inside the cached prefix. A timestamp, session ID, or serialized tool result that changes per request will invalidate the entire prefix from that point forward. Cache matching is all or nothing: a changed character, timestamp, tool serialization, or rewritten earlier turn can cause a complete miss. Audit what lives before your cache breakpoint and move anything dynamic to the end.

Does prompt caching affect output quality?

No. The model still runs full autoregressive generation for every response. The static portion of every request bills at up to 90% off, with the model producing byte-identical output. No distillation, no quantization, no quality trade-off.

Is prompt caching safe for sensitive data?

Cached data is never shared between businesses, even on byte-identical prompts. Anthropic and OpenAI both certify caching as zero data retention (ZDR) eligible-raw prompt text isn't stored at rest, only the computed key-value representations are held in memory and destroyed when the TTL expires. Review your provider's current data-handling terms before caching anything regulated.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle