How LLM Tool Calling Works, One Slack Message at a Time

LLM tool calling lets a model reach outside itself to fetch live data or trigger real actions. Here is exactly what happens between your message and the reply.

Cover art for How LLM Tool Calling Works, One Slack Message at a Time

Someone on your team types "what's our current MRR?" into Slack. An AI assistant replies in seconds with a real number, sourced from the right dashboard. No hallucination, no stale training data. What just happened between that question and the answer is worth understanding - because once you see it, you also see where it breaks.

The mechanism is called tool calling (or function calling - the terms are mostly interchangeable now). It is the reason AI assistants can do anything beyond reciting what they already know. Here is what actually happens.

What the model actually does when it calls a tool

Tool calling is a structured handshake, not magic. Tool use is a contract between your application and the model: you specify what operations are available and what shape their inputs and outputs take; the model determines when and how to call them. The model never executes anything on its own - it emits a structured request, your code runs the operation, and the result flows back into the conversation.

That last part is the thing most explanations skip. The model is not browsing the internet or calling your database directly. It is producing text - specifically, a JSON blob - that describes what it wants to call and with what arguments. Your application does the actual work.

The model still generates tokens one by one. It doesn't "natively" return a Python dictionary.

When you use function calling, the API applies constrained decoding, sometimes called "guided generation": it constrains token generation so that only tokens producing valid JSON matching your schema are allowed at each step. This is fundamentally different from just asking the model to "reply in JSON" inside a prompt - the output is structurally enforced at inference time.

Here is the sequence in full:

  1. You send a user message alongside a list of available tools (each described as a JSON schema)
  2. The model reads both and decides whether a tool is needed
  3. If yes, it emits a tool_calls block - a structured description of which function to call and what arguments to pass
  4. Your application receives that block, executes the actual code, and sends the result back
  5. The model reads the result and generates a final response to the user

When a model calls one of your tools, the API response contains a tool_use block with the tool name and a JSON object of arguments. Your application extracts those arguments, runs the operation - a database query, an HTTP call, a file write - and sends the output back in a tool_result block on the next request. The model never sees your implementation; it only sees the schema you provided and the result you returned.

Beagle in action#analytics, 10:52am
The ask
'can someone pull current MRR from Chartmogul?'
Beagle drafts
identifies the get_mrr tool in its schema, emits a structured call with the correct date parameter, receives the live figure, drafts a reply with the number and source link
You approve
you approve; the answer posts in the thread with a traceable source, not a guess
Do this in your workspace →

The hidden cost: your tool schema is part of the prompt

Here is the thing that catches teams off guard when their inference bill arrives.

Under the hood, functions are injected into the system message in a syntax the model has been trained on. This means callable function definitions count against the model's context limit and are billed as input tokens.

When you send a tool schema to an LLM, you're not "registering" a function - you're injecting a JSON blob into the model's context window, formatted as a system message. The model is then fine-tuned to output a special token sequence that signals a tool call. This is why the schema counts against your token budget: it's literally part of the prompt.

The numbers get real fast. Tools are sent with every request. A moderately complex tool schema - five tools with detailed descriptions and parameters - adds 800-1,200 tokens per request. At 10,000 requests/day, that is 8-12 million tokens per day of pure schema overhead.

Scale that up: a production agent with 20 tools and standard schema definitions carries 5,000 to 8,000 tokens of schema overhead per call. At current Opus pricing, that works out to roughly $1.4 million per year in schema tokens alone on 100,000 daily calls.

There is a compounding issue on top of that. Function calling schemas are not cached by prompt caching - only the message history is cached. Tools are resent on every request, making schema consolidation your only optimization path for high-volume applications.

800-1,200tokens per requestadded by a 5-tool schema
8-12Mtokens per dayschema overhead at 10k daily requests
5,000-8,000tokens per callfor a 20-tool production agent

OpenAI and Claude do this differently under the hood

The mechanism is the same across providers. The wire format is not.

OpenAI's implementation appends tool definitions to the system message before tokenization. The model then generates a tool_calls field in the response. Under the hood, the model outputs a JSON string inside the arguments field. The API then parses this JSON for you - but if the model outputs malformed JSON, the API returns an error.

When you call the Claude API with the tools parameter, the API constructs a special system prompt from the tool definitions, tool configuration, and any user-specified system prompt. That constructed prompt is designed to instruct the model to use the specified tool(s) and provide the necessary context.

Claude responds with stop_reason: "tool_use" and one or more tool_use blocks. Your code executes the operation and sends back a tool_result.

Server tools - such as web_search, web_fetch, code_execution - run on Anthropic's infrastructure; you see the results directly without handling execution yourself.

OpenAI Anthropic Claude
Tool call signal finish_reason: tool_calls stop_reason: tool_use
Result field tool_calls[].function.arguments tool_use content block
Server-side tools Limited (code interpreter, web) web_search, bash, browser, text_editor
Strict schema mode Yes (strict: true) Schema validation enforced
Parallel calls Yes, since GPT-4 Turbo (late 2023) Yes

The practical difference matters when you are routing across providers or upgrading models: if you are routing across multiple providers or upgrading between model versions, you cannot assume the parallelism behavior will be stable. Your orchestration layer needs to handle multi-tool responses regardless of whether you requested them.

When the model fires multiple tools at once

One non-obvious thing: a single user message can result in several tool calls happening simultaneously.

OpenAI introduced parallel function calling with GPT-4 Turbo in late 2023. When the model determines that multiple tool calls are independent, it returns all of them in a single response.

In many agents, the dominant latency comes from I/O rather than LLM inference. Parallel tool calls let the model request multiple external functions simultaneously, typically reducing total latency to the slowest single tool plus the inference cycles needed for planning and synthesis. Benchmarks like LLMCompiler show roughly 1.4x to 2.4x latency speedups on many tasks, with some scenarios reaching up to 3.7x.

But parallel execution has a trap. Parallel tool execution is a decision the model makes, not one your orchestration layer makes. When a model emits multiple tool_use blocks in a single response, your runner is expected to invoke all of them and return their results together before the next inference step. The model does not see intermediate results - it sees everything at once.

The moment you enable parallel execution, every hidden assumption baked into your tool design becomes visible. Tools that work reliably in sequential order silently break when they run concurrently. The behavior that was stable turns unpredictable, and often the failure produces no error - just a wrong answer returned with full confidence.

The other thing most tutorials skip: tool description quality determines which tool gets called. Models pick tools by reading their descriptions. Two tools with overlapping descriptions confuse the model. A description that names parameters but doesn't explain when to use the tool often gets ignored.

Asking an AI for live data
Without Beagle
model answers from training data - the number is stale, possibly by months; no source, no way to verify
With Beagle
model emits a tool call to the right API, your app fetches the live figure, the model's reply includes the number and a traceable call_id logged with its arguments

A teammate like Beagle runs on this same loop - user message arrives, relevant tool is identified from the schema, a draft response is assembled from the live result, and a human approves before anything posts.

LLM tool calling: common questions

What is the difference between function calling and tool calling?

The terms describe the same mechanism. They are used interchangeably in foundation model documentation, vendor marketing, and engineering Slack channels, and describe the same underlying mechanism: a model emits a structured request, an external system executes it, and the result feeds back into the model's context. "Tool calling" became the dominant term once the pattern expanded beyond simple local functions to include APIs, databases, and web services running in production.

Does the model actually run any code?

No. The model never executes anything on its own. It emits a structured request, your code (or the provider's servers) runs the operation, and the result flows back into the conversation. The model's only job is to decide what to call and with what arguments.

Why do my tool calls sometimes use the wrong arguments?

The JSON schema you define is not just documentation - it is the only thing the model sees. If your descriptions are vague, the model will hallucinate arguments. Write descriptions that explain when to use the tool, not just what it does. Include concrete parameter examples where the shape is non-obvious.

Can a model call multiple tools in one turn?

Yes. Parallel function calling lets the model generate multiple function calls simultaneously for a single query, determining the appropriate number of invocations. The benefit is latency: three independent API calls that each take 200ms run in 200ms total rather than 600ms. The risk is that tools designed for sequential use can silently conflict when run concurrently.

Do tool schemas count toward my token limit and cost?

Yes, and this surprises most teams. Functions are injected into the system message and count against the model's context limit, billed as input tokens. If you run into token limits, limit the number of functions loaded up front, shorten descriptions where possible, or use tool search so deferred tools are loaded only when needed.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle