Why Does the Same LLM Call Cost 90% Less the Second Time?

Prompt caching cuts LLM token costs by up to 90% with no code changes to outputs - but only if you understand KV tensors, prefix rules, and TTL traps. Here's exactly how it works.

Cover art for Why Does the Same LLM Call Cost 90% Less the Second Time?

A ten-turn agent session with a 5,000-token system prompt reprocesses that same prompt ten separate times if nothing is cached. That is not a bug. It is the default behaviour of every major LLM API - and it is also the most fixable waste in production AI today.

The fix is prompt caching. It is not a new model, a new architecture, or a vendor-specific magic trick. It is a structural property of how transformers compute attention, and once you understand the mechanism, the 90% discount stops looking like marketing and starts looking like arithmetic.

What the model is actually doing when it reads your prompt

Prompt caching stores the computational state from an LLM's attention layers so the model can skip redundant prefill work on repeated prompt prefixes. To see why that matters, you need to understand what "prefill work" means.

LLMs generate text one token at a time. At every layer of the transformer, each token gets projected into three vectors: a Query, a Key, and a Value. The Key and Value vectors for each token are computed from everything that came before it. In a 10,000-token system prompt, that is 10,000 sets of KV computations - and every single API call reprocesses your whole prompt from scratch. That prompt usually includes system instructions, tool definitions, and your full chat history. In most production apps, most of that content stays identical from one request to the next, yet without prompt caching the model works through that repeated prefix before it writes a single new word, and you pay full price for it every time.

Providers hold on to these KV matrices for each prompt for 5-10 minutes after the request is made, and if you send a new request that starts with the same prompt, they reuse the cached K and V rather than recalculating them.

What's particularly useful is that you can partially match a cache entry and still use the bit that matched, not the whole thing.

This is the whole mechanism. No distillation. No quality trade-off. Just skipping math you already paid to do.

How Anthropic, OpenAI, and Google each price it

The three providers landed on caching at different times with structurally different approaches.

Anthropic launched prompt caching in public beta on August 14, 2024 and reached general availability on December 17, 2024; OpenAI shipped automatic caching on October 1, 2024; and Google introduced explicit context caching at Google I/O in May 2024, then added zero-setup implicit caching for Gemini 2.5 models on May 8, 2025.

Provider Mechanism Cache read price Min prefix TTL
Anthropic Explicit cache_control breakpoints ~10% of input rate 1,024 tokens 5 min (1 hr available)
OpenAI Automatic, no opt-in ~50% of input rate 1,024 tokens 5-10 min
Google Gemini 2.5 Implicit server-side ~10% of input rate Varies by model Session-scoped

Anthropic uses explicit cache_control breakpoints. Writes cost 1.25× the normal input rate; reads cost 10% of normal input rate - a 90% discount. Cache TTL is 5 minutes by default, with a 1-hour option at a higher write rate.

Compared to Anthropic and Google, OpenAI's prompt caching is the most frictionless, requiring no code changes. Automatic caching is enabled for prompts that are 1,024 tokens or longer, with cache hits occurring in increments of 128 tokens.

Prompt caching with Anthropic models is explicitly controlled by the developer. Developers mark sections of the prompt as cacheable using the cache_control parameter. That extra control also means more ways to get it wrong - more on that below.

90%discount on Anthropic cache readsvs. full input token price
50%discount on OpenAI cached tokensfully automatic, no code changes
10 ×the same session replays the system prompton a 10-turn agent loop without caching

The concrete example: a Slack AI assistant on 2,000 requests per day

Picture a team AI assistant running on Claude. Its system prompt contains tone guidelines, a tool schema for searching Notion, and a few examples - call it 10,000 tokens. Users send about 2,000 queries per day. Without caching, the model processes 20 million system-prompt tokens daily that are byte-for-byte identical.

With Anthropic's current pricing on Claude Sonnet, a 10,000-token system prompt at 2,000 requests per day with a 75% cache hit rate saves roughly $1,215 per month. The user turn - the actual new question - might be 50 tokens. That is what you should be paying for.

Here is where the non-obvious insight lives: context accumulation in naive agent loops follows a quadratic cost curve, because the entire history is re-serialised and re-injected into the LLM's context window at every step. While the message history grows linearly, total billed input tokens grow quadratically, because each call re-sends prior context. Prompt caching cuts the flat system-prompt cost, but the growing conversation history is the dominant cost driver in multi-step loops, and each new tool output or reasoning trace is unique per iteration, so it cannot be cached. Caching reduces the system-prompt term in the formula but leaves the quadratic accumulation term untouched.

That is the limit most posts on this topic skip. Caching is not a complete solution to agentic token costs - it is a solution to the cheapest and most obvious slice of them.

Beagle in action#ops-ai, 11:03am
The ask
'can you pull the latest deployment status from Linear?'
Beagle drafts
reads the cached system prompt and tool schema (cached prefix, ~8k tokens), calls the Linear tool, drafts a reply with the current status
You approve
the 8,000-token system prompt costs 10% of normal on every request; only the new tool output and user turn hit full price
Do this in your workspace →

The TTL trap: why your cache silently goes cold

The discount only lands if the cache is warm when the next request arrives. The cache's default minimum lifetime is 5 minutes, and this lifetime is refreshed each time the cached content is used. That sounds fine until you map it to real workloads.

A Slack bot that handles a few messages per day: each call is effectively cold. An agent loop that takes 8-15 minutes between human inputs: writes the cache, walks past the 5-minute TTL, then writes it again on the next turn.

When the cache goes cold, the next request that would have been a hit is billed as a fresh write - at 1.25× the input rate. There is no error, no warning, just a slightly higher bill on that call.

There is a documented production incident that makes the stakes concrete. Anthropic silently dropped Claude Code's prompt-cache TTL from 1 hour to 5 minutes around early March 2026. Without explicit awareness, idle gaps of 5 minutes or more between messages evaporate the cache and force a full cold cache-write on the next message - priced at 1.25× base input on the entire conversation prefix.

On a 200K-token Opus session, that is roughly $1.25 per resume; across a working day this can raise per-session cost by 30-60%.

The fix is straightforward: dynamic content such as timestamps, user IDs, session identifiers, or request-specific metadata should go at the end of the prompt, after all shared content. Even a date injected into the system prompt ("Today's date is March 15") will prevent caching if it changes between requests, because the token at that position will have a different KV representation.

One more counterintuitive finding from production: in one study, shrinking tool output by 38.4% increased billed costs by 6.8% because the compaction invalidated cache hits and forced re-runs. Smaller prompts are not always cheaper when caching is in play.

Running a 10-turn AI agent loop
Without Beagle
system prompt re-processed 10 times at full rate; 10,000-token prompt costs the same on turn 1 and turn 10
With Beagle
system prompt cached on turn 1, read at 10% cost for turns 2-10; only new tool output and user messages hit full price

What good cache structure actually looks like

In a real production Claude workload, the input side of the bill is almost always dominated by content that does not change between calls: a system prompt with policies and tone instructions, a tool schema with a dozen function definitions, few-shot examples, a retrieved document the user is asking follow-up questions about.

Put those in order - static first, dynamic last - and mark the boundary with cache_control. You place up to four cache_control markers in a request. Everything from the start of the prompt up to each marker becomes a cacheable prefix block. On the next request, if that prefix matches byte-for-byte, those tokens are billed as cache reads at ~10% instead of fresh input.

One research paper on tool-heavy agent architectures found that placing stable tool summaries before the user message in a 30-turn session yielded a cache hit rate of 84%, versus 22% for naive full-schema injection that invalidates on every tool-list update.

Keep tool ordering deterministic - if tool definitions arrive in a different order each call, the prefix changes and the cache is never reused.

A teammate like Beagle, which holds a fixed system prompt and a stable tool schema across every call, is a natural fit for this pattern: the cacheable prefix is identical for every user in the same workspace, so a single cache write amortises across the whole team.

Prompt caching: common questions

What exactly gets stored in the cache?

The cache stores the Key-Value tensors computed for each token in the matching prefix. A prompt is cached by storing the attention KV cache. If a subsequent prompt has a matching prefix with a cached prompt, the KV cache for the matching prefix can be retrieved from the cache. The model produces byte-identical output whether it reads from cache or recomputes - no quality change.

Does prompt caching require code changes?

OpenAI's prompt caching requires no code changes - it is automatic above a 1,024-token threshold.

Anthropic's prompt caching offers more control but requires code changes to cache specific parts of a prompt using the cache_control parameter. Google Gemini 2.5 caches implicitly with no explicit setup on newer models.

What invalidates the cache?

Two things: time and content. Caching is prefix-based. If your prompt's first character changes, the entire cache becomes invalid. On the time side, Anthropic's default TTL is 5 minutes, refreshed on each read. The 1-hour TTL is available for slower workloads but costs 2× the normal write rate instead of 1.25×.

Is prompt caching safe - does it leak data between users?

Cache entries are scoped to the same account: as long as the prefix is byte-identical and the calls are made under the same account, a shared system prompt benefits from a single cache entry regardless of which user triggered the write. Cross-account cache sharing is not how any major provider implements this.

When does prompt caching not help?

When your prompt has no stable prefix. If a tool or integration appends a timestamp, session ID, or user-specific variable to the system prompt on every request, that system prompt is never identical across calls. No cache hit is possible. Every single turn processes the full system prompt from scratch. Also: caching only addresses the static portion of agent cost. The growing conversation history in a multi-step loop is unique per iteration and cannot be cached - caching leaves the quadratic accumulation cost of a long agent run untouched.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle