The model does not run your code. That one fact, once it clicks, makes everything about tool calling make sense. When GPT-4o or Claude decides it needs to check an inventory level or post a Slack message, it stops generating prose and instead outputs a structured JSON object that says: "call this function, with these arguments." Your application catches that output, runs the actual function, and hands the result back. The model never touched your database.
The key mechanic: the model only decides which tool to call and with what arguments. The actual execution happens in your code. You then feed the tool's result back to the model, which uses it to generate a final response.
That loop - prompt → model decides → your code executes → result returns → model continues - is the whole thing. Everything else is detail on top of it.
What the model actually sees
Before a model can call a tool, you have to describe every available tool to it - in the same request. Each tool gets a JSON schema: a name, a description of what it does, and a typed definition of its parameters. The model needs to know what tools are available, what parameters they accept, and what they return - and all of that lives in the system prompt.
Here is what a minimal tool definition looks like in the OpenAI format: