Someone on your team types "what's our current MRR?" into Slack. An AI assistant replies in seconds with a real number, sourced from the right dashboard. No hallucination, no stale training data. What just happened between that question and the answer is worth understanding - because once you see it, you also see where it breaks.
The mechanism is called tool calling (or function calling - the terms are mostly interchangeable now). It is the reason AI assistants can do anything beyond reciting what they already know. Here is what actually happens.
What the model actually does when it calls a tool
Tool calling is a structured handshake, not magic. Tool use is a contract between your application and the model: you specify what operations are available and what shape their inputs and outputs take; the model determines when and how to call them. The model never executes anything on its own - it emits a structured request, your code runs the operation, and the result flows back into the conversation.
That last part is the thing most explanations skip. The model is not browsing the internet or calling your database directly. It is producing text - specifically, a JSON blob - that describes what it wants to call and with what arguments. Your application does the actual work.
The model still generates tokens one by one. It doesn't "natively" return a Python dictionary.
When you use function calling, the API applies constrained decoding, sometimes called "guided generation": it constrains token generation so that only tokens producing valid JSON matching your schema are allowed at each step. This is fundamentally different from just asking the model to "reply in JSON" inside a prompt - the output is structurally enforced at inference time.
Here is the sequence in full:
- You send a user message alongside a list of available tools (each described as a JSON schema)
- The model reads both and decides whether a tool is needed
- If yes, it emits a
tool_callsblock - a structured description of which function to call and what arguments to pass - Your application receives that block, executes the actual code, and sends the result back
- The model reads the result and generates a final response to the user
When a model calls one of your tools, the API response contains a tool_use block with the tool name and a JSON object of arguments. Your application extracts those arguments, runs the operation - a database query, an HTTP call, a file write - and sends the output back in a tool_result block on the next request. The model never sees your implementation; it only sees the schema you provided and the result you returned.
The hidden cost: your tool schema is part of the prompt
Here is the thing that catches teams off guard when their inference bill arrives.
Under the hood, functions are injected into the system message in a syntax the model has been trained on. This means callable function definitions count against the model's context limit and are billed as input tokens.
When you send a tool schema to an LLM, you're not "registering" a function - you're injecting a JSON blob into the model's context window, formatted as a system message. The model is then fine-tuned to output a special token sequence that signals a tool call. This is why the schema counts against your token budget: it's literally part of the prompt.
The numbers get real fast. Tools are sent with every request. A moderately complex tool schema - five tools with detailed descriptions and parameters - adds 800-1,200 tokens per request. At 10,000 requests/day, that is 8-12 million tokens per day of pure schema overhead.
Scale that up: a production agent with 20 tools and standard schema definitions carries 5,000 to 8,000 tokens of schema overhead per call. At current Opus pricing, that works out to roughly $1.4 million per year in schema tokens alone on 100,000 daily calls.
There is a compounding issue on top of that. Function calling schemas are not cached by prompt caching - only the message history is cached. Tools are resent on every request, making schema consolidation your only optimization path for high-volume applications.
OpenAI and Claude do this differently under the hood
The mechanism is the same across providers. The wire format is not.
OpenAI's implementation appends tool definitions to the system message before tokenization. The model then generates a tool_calls field in the response. Under the hood, the model outputs a JSON string inside the arguments field. The API then parses this JSON for you - but if the model outputs malformed JSON, the API returns an error.
When you call the Claude API with the tools parameter, the API constructs a special system prompt from the tool definitions, tool configuration, and any user-specified system prompt. That constructed prompt is designed to instruct the model to use the specified tool(s) and provide the necessary context.
Claude responds with stop_reason: "tool_use" and one or more tool_use blocks. Your code executes the operation and sends back a tool_result.
Server tools - such as web_search, web_fetch, code_execution - run on Anthropic's infrastructure; you see the results directly without handling execution yourself.
| OpenAI | Anthropic Claude | |
|---|---|---|
| Tool call signal | finish_reason: tool_calls |
stop_reason: tool_use |
| Result field | tool_calls[].function.arguments |
tool_use content block |
| Server-side tools | Limited (code interpreter, web) | web_search, bash, browser, text_editor |
| Strict schema mode | Yes (strict: true) |
Schema validation enforced |
| Parallel calls | Yes, since GPT-4 Turbo (late 2023) | Yes |
The practical difference matters when you are routing across providers or upgrading models: if you are routing across multiple providers or upgrading between model versions, you cannot assume the parallelism behavior will be stable. Your orchestration layer needs to handle multi-tool responses regardless of whether you requested them.
When the model fires multiple tools at once
One non-obvious thing: a single user message can result in several tool calls happening simultaneously.
OpenAI introduced parallel function calling with GPT-4 Turbo in late 2023. When the model determines that multiple tool calls are independent, it returns all of them in a single response.
In many agents, the dominant latency comes from I/O rather than LLM inference. Parallel tool calls let the model request multiple external functions simultaneously, typically reducing total latency to the slowest single tool plus the inference cycles needed for planning and synthesis. Benchmarks like LLMCompiler show roughly 1.4x to 2.4x latency speedups on many tasks, with some scenarios reaching up to 3.7x.
But parallel execution has a trap. Parallel tool execution is a decision the model makes, not one your orchestration layer makes. When a model emits multiple tool_use blocks in a single response, your runner is expected to invoke all of them and return their results together before the next inference step. The model does not see intermediate results - it sees everything at once.
The moment you enable parallel execution, every hidden assumption baked into your tool design becomes visible. Tools that work reliably in sequential order silently break when they run concurrently. The behavior that was stable turns unpredictable, and often the failure produces no error - just a wrong answer returned with full confidence.
The other thing most tutorials skip: tool description quality determines which tool gets called. Models pick tools by reading their descriptions. Two tools with overlapping descriptions confuse the model. A description that names parameters but doesn't explain when to use the tool often gets ignored.
A teammate like Beagle runs on this same loop - user message arrives, relevant tool is identified from the schema, a draft response is assembled from the live result, and a human approves before anything posts.
LLM tool calling: common questions
What is the difference between function calling and tool calling?
The terms describe the same mechanism. They are used interchangeably in foundation model documentation, vendor marketing, and engineering Slack channels, and describe the same underlying mechanism: a model emits a structured request, an external system executes it, and the result feeds back into the model's context. "Tool calling" became the dominant term once the pattern expanded beyond simple local functions to include APIs, databases, and web services running in production.
Does the model actually run any code?
No. The model never executes anything on its own. It emits a structured request, your code (or the provider's servers) runs the operation, and the result flows back into the conversation. The model's only job is to decide what to call and with what arguments.
Why do my tool calls sometimes use the wrong arguments?
The JSON schema you define is not just documentation - it is the only thing the model sees. If your descriptions are vague, the model will hallucinate arguments. Write descriptions that explain when to use the tool, not just what it does. Include concrete parameter examples where the shape is non-obvious.
Can a model call multiple tools in one turn?
Yes. Parallel function calling lets the model generate multiple function calls simultaneously for a single query, determining the appropriate number of invocations. The benefit is latency: three independent API calls that each take 200ms run in 200ms total rather than 600ms. The risk is that tools designed for sequential use can silently conflict when run concurrently.
Do tool schemas count toward my token limit and cost?
Yes, and this surprises most teams. Functions are injected into the system message and count against the model's context limit, billed as input tokens. If you run into token limits, limit the number of functions loaded up front, shorten descriptions where possible, or use tool search so deferred tools are loaded only when needed.