The price of a fixed level of AI performance has fallen roughly 1,000x in three years
- GPT-3-quality output that cost $60 per million tokens in 2021 now runs at fractions of a cent. And yet the average enterprise AI budget grew from $1.2 million per year in 2024 to $7.0 million in 2026 , a near 5x jump. 73% of those organizations exceeded their AI cost projections in the past year. The number that explains this contradiction is not the token price. It is the loop count.
What agentic loops actually cost at the token level
The chatbot math is easy: one prompt, one response, one bill. Agent math is different. A Stanford Digital Economy Lab paper co-authored by Erik Brynjolfsson found that agentic tasks are "uniquely expensive, consuming 1,000x more tokens than code reasoning and code chat," and that the high cost is in input tokens rather than output. That last part surprises most teams. You expect to pay for what the agent writes. You actually pay mostly for what it reads - tool outputs, retrieved context, prior loop state fed back in on every step.
Input tokens, not output tokens, dominate the overall cost in agentic coding, even when token caching is enabled. Token usage is also highly variable: while more complex tasks tend to consume more tokens on average, usage varies substantially across runs, with some runs using up to 30× more tokens than others on the same task.
That variance is the budget killer teams miss. You model the happy path - agent resolves the request in 3-4 loops. On the happy path, your agent resolves a request in 3-4 loops. On the unhappy path - ambiguous queries, failed tool calls, contradictory results - it can run 10-14 loops. Production data shows the unhappy path costs up to 500% more than the happy path.
The reasoning model tax nobody budgets for
Switching to a reasoning model to improve agent accuracy looks like a quality upgrade. It is also a cost multiplier that does not show up in the per-token headline price.
Reasoning models - OpenAI o1/o3, Anthropic extended-thinking, Gemini thinking - bill for internal reasoning tokens not returned in the visible output; these can be $5-50× the visible volume and dominate cost.
A simple query that returns 7 tokens on a basic model can consume 603 tokens on a reasoning model. Same answer, 86x the cost.
The OpenTelemetry GenAI spec is starting to expose this.
It defines token usage attributes including gen_ai.usage.reasoning.output_tokens for reasoning spend on its own - the one that tells you if a reasoning model is running away with your budget.
Most teams are not instrumenting this yet.
Reasoning models can consume 100x more tokens internally than they output, creating a cost paradox where cheaper per-token pricing leads to higher total bills.
A teammate like Beagle - running inside Slack with a draft-and-approve model - avoids one common compounding factor: retries caused by ambiguous approval signals. When a human sees the draft and approves or redirects in a single message, the agent loop terminates cleanly. No multi-turn clarification spiral, no repeated tool calls burning context.
Where the inference market actually split
This is the central fact of AI economics in 2026: per-token prices are down dramatically from 2023 across the budget tier, but total enterprise AI spending is up sharply.
The market has split into two tracks that do not move together:
| Track | Direction since mid-2025 | Why |
|---|---|---|
| Budget-tier API pricing | Down ~35% YoY | Competition, MoE efficiency, vLLM throughput gains |
| Frontier model pricing | Up ~100% since Jan 2026 | Reasoning compute, larger context, new capability tiers |
| Enterprise AI budgets | Up 483% in two years | More agents, more loops, more tokens per task |
| Teams limiting AI use due to cost | 1 in 5 | Loop costs surfacing as budgets scale |
API prices have dropped 40-60% since mid-2025, yet self-hosted inference has become even cheaper thanks to more efficient models and better inference engines. The gap between API and dedicated GPU hosting has actually widened in favour of self-hosting for high-volume workloads.
The non-obvious part: as organizations expand deployment, AI operating costs including for tokens are becoming a meaningful consideration - but not yet a widespread constraint. One in five respondents says their organization is limiting AI use because of operating costs. That 1-in-5 figure is the leading indicator. It is not a ceiling yet, but it is where the growth friction starts.
A lower rate can increase total spend if it attracts more traffic, causes more retries, or reduces accepted outcomes. That is the second-order effect most procurement decisions ignore: cheaper tokens do not cap your bill if agents run more loops.
What teams can measure this week
The practical answer to agentic cost control is not a different model - it is observability before optimization.
Count loops, not just tokens. Your billing dashboard shows token volume. It does not show how many agent loops fired to produce that volume. Instrument loop count per workflow as a separate metric.
Separate reasoning tokens from output tokens. The OpenTelemetry GenAI spec's
gen_ai.usage.reasoning.output_tokensattribute gives you this breakdownbut only if you emit it. Add it now, before you scale.
Set per-workflow token budgets. Token budgets per agent and workflow are the most direct cost control, with suggested ceilings of 10,000 tokens per customer service conversation and 50,000 per code generation session.
Model the unhappy path explicitly. Build your cost estimate at the 90th percentile of loop count, not the median. If that number is acceptable, deploy. If not, add a hard loop cap before you hit production.
Route by task complexity. Model routing based on task complexity reduces costs without sacrificing quality. Reserve reasoning models for tasks where the accuracy delta is measurable, not as the default.
Optimize cost per accepted outcome, not cost per generated token. That reframe changes what you instrument, what you cap, and what you route.
AI agent inference costs: common questions
Why are AI agent costs higher than chatbot costs?
Agents run in loops - they call tools, receive outputs, feed results back into context, and repeat. Each loop adds input tokens. A single agentic task can consume 1,000x more tokens than a single chat turn, driven mostly by repeated context injection rather than model output length.
Do reasoning models cost more for agent workloads?
Yes, significantly. Reasoning models bill for internal "thinking" tokens that never appear in the visible output. These can run $5-50x the volume of the returned text. A response that looks like 7 output tokens may have consumed 600 reasoning tokens. Instrument gen_ai.usage.reasoning.output_tokens to see this directly.
Why is my AI bill growing even though token prices keep falling?
Falling per-token prices lower the unit cost, but agents generate far more units per task than chatbots. More agents, more loops, and more reasoning model use means total token volume grows faster than the price falls. Enterprise AI budgets grew 483% in two years against per-token price drops of 40-60%.
What is the most useful cost metric for agent workflows?
Cost per accepted outcome - a verified, human-approved result - is more actionable than cost per token. A workflow that generates cheap tokens but requires 12 retries is more expensive than one that uses a pricier model and closes in one loop. Track loop count, reasoning token share, and retry rate alongside token volume.
How do I cap runaway agent spend without breaking workflows?
Set hard token budgets per workflow type, instrument loop count as a first-class metric, and route reasoning models only to tasks with a measurable quality benefit. Escalate to human review when an agent approaches its token ceiling rather than letting it loop into an expensive dead end.