Your finance team opened the cloud bill in April. The per-token price had fallen again. The total owed had climbed again. Both things were true at the same time.
This is the defining tension in AI infrastructure right now, and it is not a bug in the accounting. The inference market has split: the floor is collapsing while the ceiling is rising. Most coverage picks one half of that sentence and runs with it. This post covers both, because understanding only one side is what causes AI budgets to blow up.
How AI inference costs actually fell
GPT-4-class inference fell from $30 to under $0.50 per million tokens - roughly 95% in two years, roughly 1,000x in three. To make that concrete: when GPT-3 became publicly accessible in November 2021, it was the only model achieving an MMLU score of 42, at a cost of $60 per million tokens. By March 2026, multiple models exceeded that benchmark at $0.06 per million tokens or less.
Three forces compounded on each other. Open-source competition forced API pricing down. Mixture-of-experts architectures reduced the compute required per token. Hardware improved. API prices dropped 40-60% since mid-2025, and self-hosted inference became even cheaper through more efficient models and better inference engines.
The current spread across providers is striking. Per million tokens in 2026: Google Gemini Flash runs at $0.10 input / $0.40 output; GPT-4.1 Nano at $0.10 / $0.40; DeepSeek V3 at $0.27 / $1.10; while Claude Opus 4 sits at $15.00 / $75.00 and OpenAI's o1 at $15.00 / $60.00.
The cheapest production models in 2026 cost around $0.04 per million tokens. The most expensive frontier reasoning models run upward of $180 per million tokens. That is a 4,500x pricing spread between the low end and the high end.
If you are routing every query to the top of that range, you are not buying better answers for most of them. You are buying the same answer at 4,500x the cost.
Why your bill keeps climbing anyway
Here is the part that the "tokens are cheap" coverage misses entirely.
A simple chatbot query triggers one LLM inference call. An agentic workflow - where an autonomous AI agent reasons iteratively, breaks down a task, calls tools, verifies outputs, and self-corrects - may trigger 10 to 20 LLM calls to complete a single user-initiated task. When you deploy AI into a Slack channel or a support queue, you are not paying for one call per message. You are paying for the whole loop.
Reasoning models compound this further: they can consume 100x more tokens internally than they output, creating a cost paradox where cheaper per-token pricing leads to higher total bills. The reasoning tokens are billed as output - the expensive side of every invoice. DeepSeek V4 runs in thinking mode unless you explicitly disable it; Gemini 3.5 Flash defaults to medium thinking level; Claude Opus 4.8 and GPT-5.5 default to high or medium reasoning effort. Most teams never change it.
Reasoning models burn many times more tokens than conventional ones to reach the same answer, and the overspend concentrates on easy problems. On one evaluation benchmark, some reasoning models burned over 900 tokens answering "2+3=?"
The second structural driver is the shift from seat pricing to consumption pricing. Seat-based software has fixed costs per user. AI has variable costs per interaction. One employee using AI for basic email summarization might consume 10,000 tokens per day. Another using the same tool for code generation, document analysis, or complex reasoning tasks might consume 10 million tokens per day. Same seat. 1,000x cost difference. No visibility until the bill arrives.
The routing fix most teams are not using
The fix is not switching to a cheaper model across the board. It is routing each request to the right model for that request.
Intelligent model routing is the key lever. Most tasks can run on budget-tier models ($0.10-$1/M tokens) without quality loss. Reserve frontier models ($15-$30+/M tokens) for complex reasoning and agentic workflows. Enterprises using intelligent routing typically cut costs by 60-80% without impacting user experience.
What does this look like in practice? A few concrete patterns:
Classify before you call. Route a simple intent-detection step to a cheap model first. If the query is a lookup, a status check, or a templated reply, the $0.10/M model handles it. Only novel reasoning gets escalated.
Turn down reasoning effort per task. Every major vendor gives you a dial: GPT-5.5
reasoning_effort(none → xhigh), Claude effort (low → max), Geminithinking_level(minimal → high), DeepSeek thinking type. Setting it per task is a one-line change.Cache aggressively. Route easy queries to cheaper models, persist what your agents learn so they don't re-derive it next session, and instrument token usage so you know which lever is working.
Separate your agent loops. Planning steps, tool calls, and verification steps have different quality requirements. A verification pass that checks whether a draft reply answers the original question does not need the same model that wrote the draft.
The leading organisations in enterprise AI deployment in 2026 are converging on a three-tier infrastructure model. Public cloud serves elastic training workloads, experimentation, and frontier model access. Private or colocation infrastructure serves predictable, high-volume inference with known latency requirements. Edge infrastructure serves the subset of AI applications with latency requirements so tight that even low-latency colocation cannot serve them adequately. That three-tier shape is overkill for a 50-person team, but the underlying logic - match infrastructure to workload characteristics - applies at any scale.
What the split market means for teams building on AI
The useful frame here is not "tokens are cheap" or "tokens are expensive." Both things are true simultaneously, and confusing the two is what causes most enterprise AI budgets to blow up.
The Price of Progress, a 2026 study of benchmark-level pricing, found both halves of the paradox at once: the price of reaching a given benchmark score fell 5x to 10x per year, while the cost of running the frontier itself rose 3x to 18x, because each marginal capability gain demands disproportionately more inference.
The practical read for a team building AI features into their Slack workflows, their support queue, or their internal tooling: the commodity tier is genuinely cheap now and getting cheaper. A teammate like Beagle can route most workplace lookups, thread summaries, and draft replies through the budget tier without any quality degradation - because most of those tasks are not hard. The frontier is for the tasks that actually need it.
The 2026 response to the AI inference cost crisis has produced a new discipline: FinOps for AI. The same framework that enterprise IT applied to cloud cost management in 2018-2022 is now being applied to AI inference spend - with token budgets, model routing policies, and inference optimisation teams becoming standard features of mature enterprise AI programmes.
If you are not yet doing this, the bill is the reminder. The good news is the commodity tier has never been this capable.
AI inference costs for teams: common questions
Why did my AI bill go up if token prices are dropping?
Token prices dropped, but consumption multiplied faster. An agentic workflow triggers 10-20 model calls per user request. Reasoning models burn up to 100x more tokens internally than they output. If you added agentic features or upgraded to a reasoning model without adjusting routing, total spend climbs even as per-token rates fall.
What is the actual price range for LLM inference in 2026?
Budget models like Gemini Flash and GPT-4.1 Nano run at roughly $0.10-$0.40 per million tokens. Mid-tier models like DeepSeek V3 and Mistral Small sit at $0.20-$1.10. Frontier reasoning models like Claude Opus 4 and o1 run $15-$75 per million output tokens. The spread between cheapest and most expensive production models is approximately 4,500x.
What is model routing and how much does it actually save?
Model routing means classifying each query before calling a model, then sending it to the cheapest model capable of answering correctly. Simple lookups, status checks, and template fills go to budget models; complex multi-step reasoning goes to frontier models. Enterprises using intelligent routing typically cut costs 60-80% without degrading output quality.
Should I turn off reasoning effort on my AI models?
For most tasks, yes. Reasoning tokens are billed as expensive output tokens, and research published in 2026 found that higher thinking budgets sometimes reduce accuracy rather than improving it. Every major provider exposes a reasoning effort parameter. Setting it to low or minimal for routine tasks is a one-line change that can cut a significant share of your inference bill.
Is self-hosting cheaper than API calls in 2026?
It depends on volume. API prices dropped 40-60% since mid-2025, which narrowed the self-hosting advantage. But for high-volume, predictable workloads, self-hosted inference on efficient open-weight models is still materially cheaper. The hidden cost is engineering time: a dedicated ML engineer to maintain inference infrastructure runs $180,000-$240,000 annually in the US - a real hurdle for smaller teams.