One team audited 11 popular MCP servers and found 137 tool definitions injecting 22,945 tokens into their context window before their model read a single word of the user's actual message. That is not a niche problem. It is a direct consequence of how tool calling works under the hood - and most teams only discover it on their first billing statement.
What "tool calling" actually means
The LLM does not execute tool calls directly. Instead it creates a data structure that describes the call, passing that to a separate program for execution and further processing. That distinction matters more than it sounds.
The LLM isn't calling the function itself. The model outputs a description of which function to call and what parameters to use. Your code does the actual execution. This separation is crucial: it means the model can't accidentally delete your database or make unexpected API calls.
You'll see three names for the same idea: function calling (OpenAI's original term), tool use (Anthropic's term for Claude), and tools (the API parameter both now use). They mean the same mechanism: you describe functions, the model decides when to call one, you run it, you return the result.
Here is the full loop, step by step:
- You describe your available functions using JSON Schema - names, parameter types, descriptions.
The model receives the user prompt and the tool catalog, decides whether to call a tool, then emits a structured tool call with typed arguments.
3. Once you receive this response from the LLM, the calling client is responsible for executing the function and then returning the result back to the LLM as part of an updated prompt.
4. The LLM then returns a response by either specifying another function to call or returning a response that can be sent to the user.
After executing the real function, you must send the result back to the model in a new message with role='tool' and the tool_call_id from the original response, otherwise the model cannot reason over the result.
Skip that step and the model answers as if the tool never ran.
The schema is literally inside the prompt
This is the part most explanations skip. When you send a tool schema to an LLM, you're not 'registering' a function. You're injecting a JSON blob into the model's context window, formatted as a system message.
The model is then fine-tuned to output a special token sequence that signals a tool call. This is why the schema counts against your token budget - it's literally part of the prompt.
That has real budget consequences. In OpenAI function-calling format, a single minimal tool costs roughly 60 tokens - the function name, description, parameter names, types, and the JSON scaffolding.
Add a complex tool - multiple parameters, longer descriptions, nested types - and individual tools run 150-300 tokens. A modestly equipped agent with 20-30 tools can easily spend 3,000-6,000 tokens on definitions alone.
Scale that up and the numbers get uncomfortable. MCP tool definitions consume 5-15× more tokens than the simplest possible schema for the same tool. In a typical Claude Code session with 20-30 registered MCP tools, the tool schema alone occupies 15-30 KB of context window before a single user message is sent.
One measured deployment running 31 tools found that tool definitions alone consumed 8,759 tokens - 46% of each API call - before any conversation content was processed. The system prompt added another 27%. That means the actual user message was a minority of the bill.
How the model picks the right tool - and stays honest about JSON
Through training on large numbers of examples, the model learns three core capabilities: intent recognition (whether a user's request requires tool invocation); tool selection (choosing the most appropriate tool from the available list); and parameter generation (producing valid parameter objects according to the tool's JSON Schema definition).
That third capability - parameter generation - is where things used to break. Before constrained decoding, prompting a model to "output JSON" produced valid output 80-95% of the time, which sounds fine until you do the math. A pipeline with five LLM calls, each at 97% parse success, has an 86% end-to-end success rate. For agentic workflows with dozens of tool-calling steps, this compounds into unacceptable failure rates.
Modern providers solve this with constrained decoding.
When a model decides to invoke a tool, the inference engine switches to constrained decoding mode. In this mode, token sampling is constrained by predefined JSON Schema - the model can only generate token sequences that conform to the schema structure. For example, if the schema defines a parameter type as "integer", the decoder masks the probability of all non-integer tokens, ensuring output validity.
This fundamentally solves the most troublesome problem of early prompt-based tool invocation - malformed JSON output. In production environments, constrained decoding reduces JSON parsing error rates from 15-25% with prompt-based methods to nearly 0%.
Prompt-only JSON extraction - no constrained decoding, just instructions to "output JSON" - fails at 8-15% of calls in production systems processing millions of requests.
At production volume, a 2% parse failure rate on 50,000 daily requests is 1,000 broken downstream operations per day, and the retry logic that handles them doubles the inference cost for those requests.
But there's a catch constrained decoding does not fix. Research published at EMNLP 2024 found that strict format constraints degraded reasoning accuracy by up to 27 percentage points on math benchmarks. The mechanism: JSON output forces models to emit the answer field before completing chain-of-thought reasoning, short-circuiting the deliberation that produces correct results. For tasks requiring multi-step reasoning, forcing a schema on the output can substantially hurt accuracy. Constrained decoding guarantees the shape of the output. It says nothing about whether the values inside are right.
What parallel tool calls change
A single function call is one round trip. An agent repeats it - call a tool, read the result, decide the next call - until the task is done. Every agent relies on function calling under the hood. The sequential version of that loop has obvious latency costs.
Parallel tool calls let the model issue multiple requests at once when the answers don't depend on each other. The model emits two function call items simultaneously; the caller runs both in parallel; sends both results back in one batched message; and the model writes the final answer.
Insufficient parallel tool invocation is an important source of token inefficiency in multi-step tool-use agents: weaker parallel tool-calling capability directly increases token usage and therefore leads to higher monetary cost under real API pricing. Two agents solving the same task can show dramatically different token bills just based on how aggressively they batch calls.
A teammate like Beagle works inside this loop - getting a trigger, drafting the structured request, handing back the result - without stepping outside the model's draft-and-approve flow.
How LLM tool calling works: common questions
What is the difference between function calling and tool calling?
Nothing, functionally. OpenAI introduced the term "function calling" in June 2023, then renamed the API parameter to tools later that year. Anthropic uses "tool use." They all describe the same mechanism: you declare what your code can do, the model decides when to ask for it, your application runs it, and the result goes back to the model.
Does the model actually run my function?
No. The model outputs a JSON object describing which function to call and with what arguments. Your application reads that object, decides whether to execute it, runs the function, and passes the result back as a new message. The model is a suggestion engine; your code is the executor.
Why do tool definitions cost tokens?
When you send a tool schema to an LLM, you're not registering a function. You're injecting a JSON blob into the model's context window, formatted as a system message. The model is fine-tuned to output a special token sequence that signals a tool call. This is why the schema counts against your token budget - it's literally part of the prompt. A 30-tool setup can consume 3,000-6,000 tokens per request before the conversation starts.
What is constrained decoding in tool calling?
Constrained decoding eliminates parse failure by making invalid output impossible at the token level rather than detecting it after generation. At each decoding step, the grammar engine determines which tokens can legally continue the output and masks the logits for all others. The result is a JSON object guaranteed to be schema-valid - though not necessarily semantically correct.
Can tool calling work without fine-tuning?
Yes, but poorly. The model is reading your descriptions and matching them to what the user is asking. That selection logic is just prompting under the hood. Models that have been specifically fine-tuned on tool-use examples pick the right tool far more reliably and generate cleaner argument JSON. The fine-tuning teaches the model the tool-call token format; the descriptions teach it which tool to pick.