Uber's CTO confirmed earlier this year that the company burned through its entire 2026 AI budget in four months, as Claude Code adoption jumped from 32% to 84% of its 5,000-engineer organisation, with individual engineers racking up $500 to $2,000 per month.
The CTO himself spent $1,200 in a single two-hour demo session. This is not a story about a company being reckless. It is a story about a pricing model that enterprise finance - and most team leads - have not learned to read yet.
The uncomfortable part: the price to produce GPT-4-quality output fell from about $30 per million input tokens in 2023 to under $0.50 in 2026. Token prices are genuinely collapsing. Total bills are genuinely rising. Both things are true, and the gap between them is the thing worth understanding if you are deciding how to deploy AI features for your team right now.
Why AI inference costs are falling and rising at the same time
AI inference costs are falling because accelerators, model compression, serving software, routing, and scale can reduce the compute needed for a useful response - but that does not guarantee a cheaper product. Context growth, retries, multimodal inputs, higher traffic, and latency requirements can raise the delivered cost per completed interaction.
The structural version of this: inference now accounts for 85% of the enterprise AI budget and roughly two-thirds of all global AI compute spend. The "Inference Flip" - the point where cumulative global spending on running AI models officially surpassed training - occurred in early 2026, and this shift arrives at the same time that agentic workloads are multiplying inference demand in a non-linear way.
A single chatbot API call might cost $0.001. A multi-step agent that plans, retrieves context, invokes tools, reflects on output, and self-corrects can cost $0.10 to $1.00 per task completion - a 100× to 1,000× multiplier.
Gartner's March 2026 analysis confirmed that agentic AI models require 5-30× more tokens per task than standard chatbots.
So the per-token headline rate keeps falling, while the per-task token count keeps climbing. Your invoice is the product of those two numbers. Right now, the second one is winning.
The Grok 4.7 pricing cliff is a live example
Grok 4.7, released September 21, 2026, costs $2.00 per million input tokens, $0.50 per million cached input tokens, and $6.00 per million output tokens for prompts below 200K tokens. That is a reasonable rate for a frontier model. It is also half the story.
Once a prompt reaches 200K tokens, xAI bills the entire request at $4.00/M input, $1.00/M cached input, and $12.00/M output. The long-context rate is double - and it applies to the whole request, not just the overflow. The model carries a 500,000-token context window overall, so the long-context tier applies to a meaningful chunk of real coding and document-review workloads, not just edge cases.
The tool calls are a separate meter on top. Web Search is $5 per 1,000 calls, Code Execution at $5 per 1,000, and File Attachments at $10 per 1,000.
Since the agent autonomously decides how many tools to call, costs scale with query complexity. A 20-turn debugging session that calls the code interpreter six times costs more than the token rate suggests at first glance - and that math is not hypothetical. A 20-turn agent that repeatedly sends 30K tokens of context can process roughly 600K input tokens across the session before counting its generated output.
Against the field, Grok 4.7 matches GPT-6 Sol on input ($2) but is 40% cheaper on output ($6 vs $10), and it is a fifth of Fable 5.1's rate. That is a genuine advantage. But the 200K cliff, the tool call meters, and a configurable reasoning loop mean the sticker rate is not what a production agent actually costs.
What this means for teams shipping AI features
The most confusing aspect of the AI inference cost situation for enterprise finance teams is the simultaneous reality of falling unit costs and rising total bills. Epoch AI's analysis confirms that per-token inference prices have fallen between 9× and 900× per year for various performance milestones - yet the same enterprises watching token prices collapse are seeing their monthly AI bills multiply.
The Uber story is the cleanest illustration of the mechanism. Uber maintained internal leaderboards that ranked engineers according to their Claude Code usage. That encouraged employees to consume more AI resources even though the teams promoting adoption were not responsible for controlling the budget.
Five-to-twenty-fold increases in per-developer consumption are now documented in agentic mode, and no public benchmark shows a matching multiplier on output value.
The governance gap is wide. Only 43% of organisations have formal AI governance policies, and only 21% have mature agentic governance.
Three things that actually help:
- Measure per-task cost, not per-token cost. A single token rate tells you nothing useful about a looping agent. Track cost-per-completed-task, and set thresholds before you roll out.
- Treat context length as a cost dial. Most agents carry more history than they need. Trimming context or summarising mid-session is often cheaper than switching models.
- Separate tool-call meters from token meters in your monitoring. On Grok 4.7 and similar APIs, web search and code execution are their own line items. An agent that searches on every reasoning step can easily spend more on tool calls than on tokens.
A note on what falling prices actually give you. The frontier still costs money - but yesterday's frontier is nearly free.
Inference costs are dropping roughly 10× per year for the same capability level. The practical consequence: a model that was frontier-quality 12 months ago now runs at a fraction of its original price. Teams that route non-critical tasks to last-generation models - rather than always calling the newest flagship - can keep total bills stable even as agentic usage grows.
A teammate like Beagle handling routine lookups in Slack - draft-and-approve answers to repeated internal questions - is a good example of where a cheaper, faster model is the right call, not the newest one.
The question to ask before your next AI feature ships is not "what does it cost per million tokens?" It is "what does it cost per task completion, at the 90th percentile of a real agent session?" Those two numbers can differ by a factor of 100. Right now, most teams only know the first one.
AI inference costs for teams: common questions
What is the difference between token cost and inference cost?
Token cost is the per-million-token rate a provider publishes. Inference cost is what you actually pay to complete a task - which includes context length, number of model calls, tool invocations, retries, and caching behaviour. For simple chat, the two numbers are close. For agentic loops, they can differ by two orders of magnitude.
Why are my AI bills going up when prices keep falling?
Per-token prices are falling, but agents consume far more tokens per task than a single chat turn does. Gartner's March 2026 data puts the multiplier at 5-30× more tokens per agentic task. If your team has shifted from chat to agents, usage volume is growing faster than unit prices are dropping.
How do I calculate what an AI agent actually costs per task?
Take a sample of 50-100 real agent sessions. For each, record total input tokens, output tokens, cached tokens, and any tool-call counts. Price each component at the applicable tier (including long-context repricing if any session crosses your provider's threshold), then divide by completed tasks. That per-task number is your planning unit.
Does a longer context window always cost more?
Not per token, but yes in practice. Longer context windows encourage developers to pass more history per call. On Grok 4.7, any prompt over 200K tokens reprices the entire request at 2× - so a longer window can double the cost of a session that crosses that line, even if average sessions are shorter.
What is the cheapest way to run an AI agent at work?
Route to the cheapest model that reliably completes the task. Use prompt caching for repeated system-prompt content. Summarise conversation history rather than passing full transcripts. Set reasoning effort explicitly rather than defaulting to high. Monitor tool-call counts separately from token counts. Each of these is independent; used together they can cut per-task cost by 50-80% without changing the model.