The AI Inference Cost Paradox Teams Keep Walking Into

Token prices have fallen 280x since 2022. Enterprise AI bills are rising anyway. Here's the exact math on why-and how agents, reasoning models, and the Jevons Paradox explain the gap.

Cover art for The AI Inference Cost Paradox Teams Keep Walking Into

Querying a model at GPT-3.5-level performance cost $20 per million tokens in November 2022. By October 2024, Gemini 1.5 Flash-8B hit the same quality bar for $0.07-a 280-fold price drop measured by Stanford's 2025 AI Index. And yet, enterprise AI budgets are not shrinking. They're growing. The teams spending hardest on optimization are often the ones whose bills climb fastest.

This is not a billing glitch. It has a name: the Jevons Paradox. And understanding it is probably more valuable than any prompt engineering trick you'll find this week.

Why cheaper tokens produce bigger bills

William Stanley Jevons identified the mechanism in 1865: when technological improvement reduces the cost of using a resource per unit of output, total consumption tends to increase, not decrease, because cheaper use induces greater demand. Watt's fuel-efficient steam engine did not conserve coal-it made coal economically useful in industries that had never used it, and England's coal consumption soared.

Cheaper tokens do not replace existing AI workloads at lower cost. They unlock new workloads that were previously too expensive to build. Those workloads-agent loops, multi-step reasoning chains, multi-model pipelines-consume dramatically more tokens per task than the simpler systems they extended or replaced.

The OpenRouter data makes the pattern concrete. OpenRouter published an analysis of over 100 trillion tokens processed across its platform. Programming went from roughly 11% of total token volume in early 2025 to over 50% by March 2026.

This didn't happen because existing programmers started using more tokens. It happened because an entirely new use case-agentic coding-was created by cheaper, faster inference. Tools like Claude Code, Cursor, and Codex turned programming from a human activity into a human-directed AI activity.

Token consumption per developer has risen an estimated 100× to 1,000× from 2023 levels, from tens of thousands to tens of millions of tokens per month. Tokens are roughly 48× cheaper, but developers consume 100× to 1,000× more of them. Net spend per developer rises 2× to 20× despite the price collapse.

280×price drop since Nov 2022GPT-3.5-level performance per million tokens
5-30×more tokens per taskagentic workflows vs. a single chat query
50%+share of platform tokensprogramming/agent use cases by March 2026

The two multipliers that break pilot economics

There are two distinct ways a reasonable-looking bill turns into a shock in production.

The agentic loop multiplier. A simple chatbot query triggers one inference call, but an agentic workflow-where a large language model reasons iteratively, calls external tools, verifies outputs, and self-corrects-can trigger 10 to 20 model calls for a single user-initiated task. Illustrative math from production benchmarks is unforgiving: at a blended rate of roughly $1 per million input tokens and $4 per million output, a RAG query with retrieved context costs around $0.008 per task; an agentic task with 12 chained calls costs around $0.12-about 15× more for the same question answered at higher quality.

Enterprises that scaled past the pilot phase discovered this multiplier only after their production bills arrived. The pilot economics bore no relationship to the production economics of multi-step agentic loops running thousands of times per day.

The reasoning token multiplier. Reasoning models like o3 and DeepSeek R1 generate internal chain-of-thought tokens before producing a visible answer. These models generate internal chain-of-thought tokens before producing the final answer. Models like DeepSeek-R1 and o3 can produce 2,000-30,000 thinking tokens per query, resulting in 5-50× more tokens per task than standard models. The pricing reflects this: reasoning models cost 5-10× more than standard models because they use much more compute per query during extended thinking time.

What makes this hard to catch: reasoning token usage varies dramatically across models. GPT-5-nano uses reasoning tokens on 71% of queries, while some chat-optimized models use them on only 2%. This hidden cost substantially inflates cheaper models' per-query expense beyond what their ultra-low per-token price would suggest. A model with a $0.05/M input price can end up costing more per finished task than one priced at $2.50/M, depending on how many reasoning tokens it burns.

Beagle in action#ai-budget, quarterly review
The ask
'our inference bill tripled even though we moved to cheaper models'
Beagle drafts
pulls token logs from the past 30 days, drafts a breakdown by workflow type-showing that the new agentic summarization pipeline is driving 73% of spend at 40× the token count of the old chat feature
You approve
you approve; the team sees exactly which workflow to optimize before the next billing cycle
Do this in your workspace →

The two levers that actually move the number

Model routing: match the model to the task

Most requests-simple questions, categorization, summarization-do not need the most powerful model. GPT-4o Mini, Gemini 2.5 Flash, or DeepSeek V3 handle them at a fraction of the cost. Routing is the practice of automatically sending each request to the cheapest model that can handle it, rather than calling a frontier model for everything.

The math compounds quickly. Google's Gemini Flash-Lite leads at $0.075 per million input tokens for capable summarization-class work. Sending 80% of a team's requests there instead of to a $3/M mid-tier model cuts blended input cost by roughly 85% on that volume-before any other optimization.

The non-obvious catch: you have to measure task-level quality, not just benchmark scores. A model that scores well on MMLU may still fail on your specific document format. Routing decisions need production telemetry, not just spec sheets.

Prompt caching: stop paying for the same tokens twice

Prompt caching is a provider-side feature that stores the processed state of the beginning of your prompt-the system prompt, tool definitions, and prior conversation-so repeated requests with the same prefix don't reprocess those tokens. Cached reads cost about a tenth of the base input price on Anthropic, OpenAI, and Gemini. For agents that resend the same context every turn, it's the largest single cost lever available.

The implementations differ in a way that matters:

Provider Caching mechanism Cached input discount Setup required
Anthropic Explicit cache_control marker 90% off reads One extra API field per content block
OpenAI Automatic, prefix-based 50% off None-applies to prompts ≥ 1,024 tokens
Google Gemini Configurable TTL Down to $0.03/M (from $0.30/M) Configurable per-request

Most teams treat each API call as stateless. Prompt caching requires you to recognize reusable context patterns-document analysis, system instructions, retrieval results-and design your application architecture around them. This is an architectural decision, not a parameter tweak: it changes how you structure request batching, session management, and cache invalidation logic.

One gotcha worth knowing: on Anthropic's system, the first time a prompt is processed Anthropic stores the cache and you pay 25% more than the normal input price for that write. The savings only materialize on subsequent hits. At fewer than two or three requests reusing the same prefix, you're better off not caching.

Agentic summarization pipeline, 100K tasks/month
Without Beagle
every call re-sends the full 8,000-token system prompt and tool schema; billed at $3/M; input cost alone reaches ~$2,400/month
With Beagle
system prompt cached; 90% discount on reads after the first call; same workload costs ~$300/month in cached input tokens

What this means for teams shipping AI features

The pricing deflation is real- the price for a given level of benchmark performance has decreased around 5× to 10× per year for frontier models on knowledge, reasoning, math, and software engineering benchmarks. These reductions are due to economic forces, hardware efficiency improvements, and algorithmic efficiency improvements. That trend almost certainly continues.

But the Jevons dynamic is also real: Uber burned through its entire 2026 AI coding budget in four months, driven by Claude Code adoption. Goldman Sachs projects a 24× increase in token consumption by 2030.

The practical upshot is that cost modeling for AI features needs a different shape than SaaS cost modeling. Teams that piloted with chat and shipped with agents watch consumption grow 10× with zero change in pricing or headcount. Budget for the agentic multiplier from the start, build routing into the architecture rather than retrofitting it, and treat caching as a structural choice rather than an optimization you'll get to later.

The teams whose AI bills make sense aren't the ones that found a cheaper model. They're the ones that built cost instrumentation before they needed it.

Beagle in action#engineering, sprint planning
The ask
'before we add more agent steps, can someone sanity-check the token cost?'
Beagle drafts
queries the last sprint's token telemetry, drafts a per-feature cost breakdown with a 10× agentic multiplier applied to the proposed new steps
You approve
you approve the draft; the team ships with a cost ceiling baked in, not discovered after the fact
Do this in your workspace →

AI inference cost: common questions

Why is my AI bill rising if token prices keep falling?

The Jevons Paradox: cheaper tokens unlock workloads that were previously uneconomical-agentic loops, reasoning chains, multi-model pipelines-and these new workloads consume far more tokens per task than the chat interactions they replaced. The per-token price drops; the tokens-per-task multiplier rises faster. Total spend goes up.

How much more do agentic workflows cost than a simple chat query?

Significantly more. A single chat query runs around 800 tokens. An agentic task with tool calls, retries, and context reloading runs 10,000-50,000 tokens, with 10-20 model calls per user action. Gartner's March 2026 analysis put the range at 5-30× more tokens per task than a standard chatbot query.

What is prompt caching and how much does it save?

Prompt caching stores the processed state of your system prompt and static context so the model skips recomputing those tokens on every call. Anthropic charges 10% of the base input rate on cache reads; OpenAI charges 50% of base automatically. For agents that resend a large system prompt every turn, caching is typically the highest-leverage single cost reduction available-41-80% savings documented in multi-turn agent sessions.

Do reasoning models cost more per query than standard models?

Yes, substantially. Reasoning models like o3 and DeepSeek R1 generate internal thinking tokens before producing a visible answer-2,000 to 30,000 extra tokens per query depending on task complexity. This makes them 5-10× more expensive per query than standard models at equivalent published per-token prices, and the cost is easy to miss because thinking tokens don't appear in the output you see.

How should teams budget for AI inference costs?

Model cost-per-task, not cost-per-token. Multiply tokens per request by requests per user by users, at your blended routing mix's rate, then stress-test with a 10-20× agentic multiplier for any workflow that chains model calls. Build in routing-cheap models for classification and summarization, frontier models for reasoning-heavy tasks only-and implement prompt caching before you scale, not after your first big bill.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle