Cheaper Tokens Are Making Your AI Agent Bills Worse

Gartner's August 2026 forecast says agentic workflow costs will rise more than fivefold by 2028, even as token prices keep falling. Here's what that means for teams shipping AI features today.

Cover art for Cheaper Tokens Are Making Your AI Agent Bills Worse

On August 17, Gartner published a number that should sit next to every agent roadmap: AI inference costs per agentic workflow will increase more than fivefold through 2028. The uncomfortable part is what it sits alongside. In the same breath, Gartner expects the per-token price of frontier inference to fall by more than 90 percent by 2030. They call the combination the inference paradox: the visible unit price goes down, and the cost of the thing you actually buy - a completed workflow - goes up.

If you are still measuring your AI spend in dollars per million tokens, you are measuring the wrong thing.

Why cheap tokens do not make agents cheap

The short version: an agent is not one model call; it is a loop. A simple chatbot interaction may involve one request and one response. An agentic workflow involves a sequence of model calls - planning a task, selecting tools, retrieving information, interpreting results, checking work, correcting mistakes, and deciding what to do next.

And here is the structural problem: an agentic workflow does not answer one question and stop. It reads relevant files, forms a plan, executes a step, validates the output, revises based on the result, queries additional context, and loops until the task is complete. Each of those steps is a separate API call. Each call resends the full accumulated context window as input. The model does not remember the previous call. It is told everything again, every time.

Gartner's March 2026 analysis found agentic workflows burn 5 to 30 times more tokens per task than a simple chatbot query. Stanford researchers studying coding agents found the gap can go further: up to 1,000x versus simple code chat, driven almost entirely by input tokens the model has to re-read, not output tokens it generates.

Where a simple chatbot must read and interpret a query and quickly respond with a probabilistically reasonable answer, an AI agent must constantly reason, negotiate, and question itself. That reasoning costs tokens on every cycle.

5-30xmore tokens per agentic taskvs. a standard chatbot query (Gartner, 2026)
5x+projected rise in per-workflow inference costthrough 2028 despite falling token prices
73%of enterprises exceeded AI cost projectionsin the past year (FinOps Foundation 2026)

The market-level picture makes this concrete. The AI inference market has split: the floor is collapsing while the ceiling is rising. Both are true simultaneously, and confusing the two is what causes most enterprise AI budgets to blow up.

The median LLM API price across 130 models tracked by BenchLM as of July 24, 2026 is $1.00 per million input tokens and $4.00 per million output tokens

  • that is the floor. But frontier pricing has doubled since January 2026 alone , as newer generations command higher prices for expanded capability. Teams who route everything to the latest frontier model are buying the ceiling and paying for it per step.

The pilot-to-production gap is where budgets break

Most agents feel cheap in pilots. The reason is simple: pilots run a handful of tasks manually, with a human watching each one. Production runs thousands of tasks autonomously. The pattern is consistent across enterprise deployments: the pilot consumed a fraction of what production consumes - not because the pilot was poorly designed, but because a chatbot and an agent are fundamentally different cost models, and almost nobody models that difference before the deployment decision is finalized.

A real number to anchor this: a growth-stage SaaS company with 35 engineers had been running Claude Code, Cursor, and a custom autonomous bug-triage agent for four months. Their April 2026 bill was $87,000.

After routing routine triage to a smaller model and escalating to frontier only on hard cases - plus adding context pruning and per-developer spend caps - their May 2026 bill was $24,000. Annual savings: $756,000.

Nothing changed about what the agent did. What changed was which model handled which step.

Running a ticket-triage agent in production
Without Beagle
every step routed to the same frontier model - planning, tool calls, validation, and reply all at $15/MTok output; 10,000 tickets/month costs ~$16,000 in LLM spend alone
With Beagle
a tiered routing layer sends routine classification to a sub-$1/MTok model and escalates only ambiguous or high-stakes cases to frontier - same throughput, 60-80% lower LLM bill

The ROI models that justified most enterprise agentic deployments were built on chatbot-level token assumptions. The production numbers are an order of magnitude higher.

Gartner's analyst Will Sommer put it plainly: "Defaulting to generic autonomous intelligence will result in unbounded costs orders of magnitude higher than those of optimized product ecosystems."

What inference tiering actually means for a team

Gartner's analyst said it directly: "Product leaders cannot rely on more efficient token economics to rationalize AI costs. Each successive generation of AI capability will necessitate more, and often more expensive, tokens. There is no reliable, economical one-size-fits-all model on the horizon. Producing competitive AI products will require developing and maintaining complex multimodel ecosystems."

The practical shape of that multimodel ecosystem is a routing layer: classify each step in a workflow, and match it to the cheapest model that can handle it reliably.

What that looks like in practice:

  • Triage and classification steps are pattern-matching jobs. A fast 8B or small MoE model - DeepSeek V4 Flash costs fractions of a cent per thousand tokens - handles these without frontier overhead.

  • Tool calling and result parsing are structured tasks. Mid-tier models are adequate for most of these, and they are far cheaper than the frontier reasoning tier.

  • Final synthesis or hard edge cases are where a frontier model earns its cost. A task routed to a frontier reasoning model may cost 190x more than the same task handled by a fast small model.

  • Context management is the other lever. Dynamic tool loading can reduce context overhead by 85%

  • which matters because the agent re-reads that context on every step.

  • Prompt caching is available on Anthropic, OpenAI, and AWS Bedrock, with cached input charged at 10-25% of normal input cost. For an enterprise running 5,000 agent loops per day, that is $2,000+ saved per day on system prompt caching alone.

A teammate like Beagle lives inside this constraint by design. Every triggered response is a draft that waits for approval - which means the agent does not spiral into multi-step loops on unreviewed tasks. The draft-and-approve model is a natural token budget: the workflow stops until a human nods.

Beagle in action#ops-alerts, 2:47pm
The ask
'can someone pull which customers were affected by last night's timeout?'
Beagle drafts
queries the linked incident doc and status page, drafts a reply listing affected accounts with timestamps and a source link
You approve
you review and approve in one click - one workflow, one bounded token budget, logged with its reason
Do this in your workspace →

Stop tracking cost per token. Track cost per completed task.

Ask most enterprise leaders about AI costs, and they will cite a figure per million tokens. Ask what that spend delivers, and the answer becomes far less clear. That gap is not a reporting issue - it is a measurement flaw. Cost per token shows the price of raw input, not the value of what is created.

The unit that matters is cost per completed task, measured against success rate and escalation rate. What companies need to verify is not whether they have lowered the per-token price for each model. It is whether they can track, workflow by workflow, the cost per accepted outcome, the success rate, the number of retries, and the burden of human review.

Artificial Analysis publishes an absolute dollar cost per task on its Intelligence Index. At max reasoning effort, as of August 2026: GPT-5.6 Sol costs $1.04 per task, Claude Opus 4.8 costs $1.80, and Claude Opus 5 costs $2.03. Those are useful numbers - not because they tell you which model to pick, but because they give you a real denominator. If your agent completes a task worth $50 in human time at $1.04, that is a strong trade. If it costs $2.03 and succeeds only 60% of the time, the math stops working.

The inference paradox is real, but it is not a reason to avoid agents. It is a reason to be precise about which tasks agents run, on which model tiers, and how many steps you let them take unsupervised before a human reviews the work.


Agentic AI inference cost: common questions

Why are AI agent costs rising even though token prices are falling?

Because agents use many more tokens per task than chatbots. A chatbot exchange might use 200-2,000 tokens. An agentic workflow re-reads the full accumulated context on every step, and a multi-step task can easily exceed 100,000 tokens. Falling price per token is outpaced by rising token volume per workflow - Gartner calls this the inference paradox.

What is inference tiering and how does it cut agent costs?

Inference tiering means routing each step of an agentic workflow to the cheapest model that can handle it reliably, rather than using one frontier model for everything. Classification and triage steps go to small, fast models costing under $1 per million tokens. Hard edge cases and final synthesis go to frontier models. Documented enterprise deployments show 60-80% cost reductions from this pattern alone.

How many tokens does an AI agent actually use per task?

It depends on workflow complexity, but Gartner's 2026 analysis puts agentic systems at 5-30x more tokens than a standard chatbot query for the same job. A simple tool-calling agent uses roughly 5,000-15,000 tokens per task. A complex multi-agent orchestration system can exceed one million tokens per task.

Should I measure AI cost by tokens or by task?

Track cost per completed task, not cost per million tokens. Token price shows what you paid for raw compute; it says nothing about what you got. The useful denominator is cost per accepted outcome, alongside the success rate and escalation rate for each workflow. That framing tells you whether a workflow is profitable - token price alone does not.

When does it make sense to run a reasoning model for every agent step?

Rarely. Reasoning models earn their cost on ambiguous, high-stakes steps where a smaller model fails often enough that the retry cost or human escalation cost exceeds the frontier premium. For structured steps - parsing, classification, tool selection - a fast small model is almost always the better trade. The 190x price gap between a frontier reasoning model and a capable small model makes blanket frontier routing hard to justify at any real task volume.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle