How LLM Tool Calling Actually Works Under the Hood

LLM tool calling is not what it sounds like. The model never runs your code - it outputs structured JSON and your app does the work. Here's the full loop, explained plainly.

Cover art for How LLM Tool Calling Actually Works Under the Hood

One team audited 11 popular MCP servers and found 137 tool definitions injecting 22,945 tokens into their context window before their model read a single word of the user's actual message. That is not a niche problem. It is a direct consequence of how tool calling works under the hood - and most teams only discover it on their first billing statement.

What "tool calling" actually means

The LLM does not execute tool calls directly. Instead it creates a data structure that describes the call, passing that to a separate program for execution and further processing. That distinction matters more than it sounds.

The LLM isn't calling the function itself. The model outputs a description of which function to call and what parameters to use. Your code does the actual execution. This separation is crucial: it means the model can't accidentally delete your database or make unexpected API calls.

You'll see three names for the same idea: function calling (OpenAI's original term), tool use (Anthropic's term for Claude), and tools (the API parameter both now use). They mean the same mechanism: you describe functions, the model decides when to call one, you run it, you return the result.

Here is the full loop, step by step:

  1. You describe your available functions using JSON Schema - names, parameter types, descriptions.

The model receives the user prompt and the tool catalog, decides whether to call a tool, then emits a structured tool call with typed arguments.

3. Once you receive this response from the LLM, the calling client is responsible for executing the function and then returning the result back to the LLM as part of an updated prompt.

4. The LLM then returns a response by either specifying another function to call or returning a response that can be sent to the user.

After executing the real function, you must send the result back to the model in a new message with role='tool' and the tool_call_id from the original response, otherwise the model cannot reason over the result. Skip that step and the model answers as if the tool never ran.

The schema is literally inside the prompt

This is the part most explanations skip. When you send a tool schema to an LLM, you're not 'registering' a function. You're injecting a JSON blob into the model's context window, formatted as a system message.

The model is then fine-tuned to output a special token sequence that signals a tool call. This is why the schema counts against your token budget - it's literally part of the prompt.

That has real budget consequences. In OpenAI function-calling format, a single minimal tool costs roughly 60 tokens - the function name, description, parameter names, types, and the JSON scaffolding.

Add a complex tool - multiple parameters, longer descriptions, nested types - and individual tools run 150-300 tokens. A modestly equipped agent with 20-30 tools can easily spend 3,000-6,000 tokens on definitions alone.

Scale that up and the numbers get uncomfortable. MCP tool definitions consume 5-15× more tokens than the simplest possible schema for the same tool. In a typical Claude Code session with 20-30 registered MCP tools, the tool schema alone occupies 15-30 KB of context window before a single user message is sent.

22,945 tokensinjected before first messageacross 11 real MCP servers (137 tools)
~1,000 tokensper heavy tool definitionmeasured via Anthropic's token-counting API
3,000-6,000tokens for 20-30 toolsbefore any conversation content

One measured deployment running 31 tools found that tool definitions alone consumed 8,759 tokens - 46% of each API call - before any conversation content was processed. The system prompt added another 27%. That means the actual user message was a minority of the bill.

How the model picks the right tool - and stays honest about JSON

Through training on large numbers of examples, the model learns three core capabilities: intent recognition (whether a user's request requires tool invocation); tool selection (choosing the most appropriate tool from the available list); and parameter generation (producing valid parameter objects according to the tool's JSON Schema definition).

That third capability - parameter generation - is where things used to break. Before constrained decoding, prompting a model to "output JSON" produced valid output 80-95% of the time, which sounds fine until you do the math. A pipeline with five LLM calls, each at 97% parse success, has an 86% end-to-end success rate. For agentic workflows with dozens of tool-calling steps, this compounds into unacceptable failure rates.

Modern providers solve this with constrained decoding. When a model decides to invoke a tool, the inference engine switches to constrained decoding mode. In this mode, token sampling is constrained by predefined JSON Schema - the model can only generate token sequences that conform to the schema structure. For example, if the schema defines a parameter type as "integer", the decoder masks the probability of all non-integer tokens, ensuring output validity.

This fundamentally solves the most troublesome problem of early prompt-based tool invocation - malformed JSON output. In production environments, constrained decoding reduces JSON parsing error rates from 15-25% with prompt-based methods to nearly 0%.

Prompt-only JSON extraction - no constrained decoding, just instructions to "output JSON" - fails at 8-15% of calls in production systems processing millions of requests.

At production volume, a 2% parse failure rate on 50,000 daily requests is 1,000 broken downstream operations per day, and the retry logic that handles them doubles the inference cost for those requests.

But there's a catch constrained decoding does not fix. Research published at EMNLP 2024 found that strict format constraints degraded reasoning accuracy by up to 27 percentage points on math benchmarks. The mechanism: JSON output forces models to emit the answer field before completing chain-of-thought reasoning, short-circuiting the deliberation that produces correct results. For tasks requiring multi-step reasoning, forcing a schema on the output can substantially hurt accuracy. Constrained decoding guarantees the shape of the output. It says nothing about whether the values inside are right.

Beagle in action#product-team, 2:47pm
The ask
'can someone pull the open P1 count from Linear before the standup?'
Beagle drafts
recognizes the lookup intent, calls a tool against the Linear API with the correct status filter, gets the count back, and drafts a reply with the figure and a link to the filtered view
You approve
you approve; the answer posts in the thread with a traceable source - no copy-paste, no tab-switching
Do this in your workspace →

What parallel tool calls change

A single function call is one round trip. An agent repeats it - call a tool, read the result, decide the next call - until the task is done. Every agent relies on function calling under the hood. The sequential version of that loop has obvious latency costs.

Parallel tool calls let the model issue multiple requests at once when the answers don't depend on each other. The model emits two function call items simultaneously; the caller runs both in parallel; sends both results back in one batched message; and the model writes the final answer.

Insufficient parallel tool invocation is an important source of token inefficiency in multi-step tool-use agents: weaker parallel tool-calling capability directly increases token usage and therefore leads to higher monetary cost under real API pricing. Two agents solving the same task can show dramatically different token bills just based on how aggressively they batch calls.

A teammate like Beagle works inside this loop - getting a trigger, drafting the structured request, handing back the result - without stepping outside the model's draft-and-approve flow.

Answering a live data question in Slack
Without Beagle
someone tabs to a different tool, runs a query or search, pastes the result back into the channel, and the source is immediately lost
With Beagle
a tool call fetches the answer in-thread, the result posts with its source linked, and the action is logged

How LLM tool calling works: common questions

What is the difference between function calling and tool calling?

Nothing, functionally. OpenAI introduced the term "function calling" in June 2023, then renamed the API parameter to tools later that year. Anthropic uses "tool use." They all describe the same mechanism: you declare what your code can do, the model decides when to ask for it, your application runs it, and the result goes back to the model.

Does the model actually run my function?

No. The model outputs a JSON object describing which function to call and with what arguments. Your application reads that object, decides whether to execute it, runs the function, and passes the result back as a new message. The model is a suggestion engine; your code is the executor.

Why do tool definitions cost tokens?

When you send a tool schema to an LLM, you're not registering a function. You're injecting a JSON blob into the model's context window, formatted as a system message. The model is fine-tuned to output a special token sequence that signals a tool call. This is why the schema counts against your token budget - it's literally part of the prompt. A 30-tool setup can consume 3,000-6,000 tokens per request before the conversation starts.

What is constrained decoding in tool calling?

Constrained decoding eliminates parse failure by making invalid output impossible at the token level rather than detecting it after generation. At each decoding step, the grammar engine determines which tokens can legally continue the output and masks the logits for all others. The result is a JSON object guaranteed to be schema-valid - though not necessarily semantically correct.

Can tool calling work without fine-tuning?

Yes, but poorly. The model is reading your descriptions and matching them to what the user is asking. That selection logic is just prompting under the hood. Models that have been specifically fine-tuned on tool-use examples pick the right tool far more reliably and generate cleaner argument JSON. The fine-tuning teaches the model the tool-call token format; the descriptions teach it which tool to pick.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle