Tool calling is the mechanism that lets a model ask an application to run a function on its behalf, then use the result to keep working. It is one of the main building blocks for agents. Without tools, the model cannot fetch fresh data or take actions by itself. The application can still place information in the prompt, but the model has no direct reach beyond that context.
For example, if a user asks “What is the weather in Paris right now?”, the model cannot know this from training alone. Tool calling lets it request a get_weather(city) function, receive the answer, and use it to reply.
How a tool call works
A tool call usually follows this sequence:
- The application describes the available tools to the model — their names, purposes, and inputs.
- The model decides a tool is needed and the API returns a structured tool-call object, usually with a call ID, function name, and JSON arguments.
- The application executes the actual function or API call.
- The result is sent back to the model as part of the conversation.
- The model reads the result and continues, either answering the user or calling another tool.
The model never runs code itself. It only produces a request. The application is responsible for executing it safely and returning the result.
Describing tools to the model
Each tool needs a clear definition so the model knows when and how to use it. A typical tool definition includes:
- Name: a short, unambiguous identifier, like
search_orders. - Description: what the tool does and when to use it.
- Parameters: the expected inputs, usually as a structured schema (name, type, required or optional).
For example:
Tool: search_orders
Description: Find a customer's orders by email or order ID.
Parameters:
- query (string, required): email address or order ID
Vague descriptions lead to misuse. If two tools look similar, the model may call the wrong one or pass bad arguments.
Good tool design makes the next step obvious:
- Use clear names and descriptions that say when to use the tool and when not to use it.
- Use JSON Schema with required fields, types, ranges, and enums where choices are limited.
- Keep tools small and focused. A
refund_ordertool is easier to secure than a broadrun_admin_actiontool. - Return concise results with the facts the model needs next, not a full raw API dump.
- Return helpful errors the model can act on, such as
order_not_foundormissing_email. - Make write actions idempotent when possible, so retries do not duplicate work.
- Set timeouts and retry only safe failures.
Structured output for calls
Tool calls are usually returned as structured data rather than free text. Modern APIs expose them as tool-call objects. The application should not scrape prose to guess what the model meant.
{
"id": "call_42",
"name": "search_orders",
"arguments": { "query": "user@example.com" }
}
Structured calls reduce ambiguity. The application can read the call ID, function name, and arguments directly, validate them, and send the matching result back under the same ID. This is related to prompt engineering and structured output, but the important point is that tool calls are machine-readable requests, not ordinary assistant text.
Handling the result
Once a tool runs, the host appends its output to the conversation as tool-result context. The model reads that result on the next turn. It has not learned the information into its weights; it is simply using new context for this interaction. The result should be clear, relevant, and shaped for the next model step.
If a tool call fails — for example, the order was not found — the error should also be returned to the model in a clear form. A silent failure can cause the model to guess or hallucinate an answer instead of reporting the problem.
Parallel and multi-step calls
Some tasks need more than one tool call. If the calls are independent, the model may request them in one assistant turn and the host can run them in parallel. For example, it can fetch weather and hotel availability for the same city at the same time.
Dependent calls must happen in sequence. If the second call needs an ID returned by the first call, the host should wait, append the first result, and let the model choose the next call. Each result is matched to its call ID so the model can tell which output belongs to which request.
Why tool calling needs guardrails
Tool calling gives a model real-world reach, so it also introduces risk. A model could call a tool with the wrong arguments, call a destructive tool by mistake, or be tricked into misusing a tool through malicious input in the conversation.
Common safeguards include:
- Narrow, single-purpose tools instead of broad, powerful ones.
- Authorizing every call as the end user, not as “the model.”
- Validating arguments server-side before execution.
- Requiring human confirmation for sensitive actions, like deleting data or sending money.
- Treating tool output as untrusted data that may contain prompt injection, especially when it came from web pages, files, tickets, or emails.
- Logging every call and its result for audit and debugging.
- Limiting call rates, budgets, and how many calls can happen in a row.
The host or harness enforces these safeguards. The model can request a call, but it should not be the authority that decides whether the user is allowed to run it.
Tool calling vs MCP
Tool calling is the general mechanism: a model requests a function, the application runs it, and the result comes back. MCP is a standard protocol for exposing tools (along with resources and prompts) in a consistent format across many applications. You can have tool calling without MCP, but MCP builds on the same core idea.
The key idea
Tool calling lets a model request actions instead of only generating text. The model proposes a call, the application executes it safely, and the result flows back into the conversation. Clear tool definitions, structured requests, and safety checks are what make tool calling dependable rather than risky.