Tool Calling
Tool calling, also called function calling or tool use, is the capability that lets a large language model (LLM) request an action from external code. The developer describes the available tools with a name, a description and an input schema. The model decides when a tool is needed and returns a structured call with arguments, and the application runs it and sends the result back.
How Tool Calling Works
The flow is similar across providers. OpenAI's documentation describes five steps, and Anthropic's follows the same round trip:
- Send the request with tool definitions. Each tool has a name, a description of what it does and when to use it, and parameters defined in JSON Schema.
- Receive a tool call. Instead of plain text, the model returns a structured request naming the tool and the arguments it wants to pass. In Anthropic's API this arrives as a
tool_useblock with a stop reason oftool_use. - Execute the tool. The application, not the model, runs the function, calls the API or queries the database.
- Return the result. The application sends the output back to the model, in Anthropic's API as a
tool_resultblock linked to the original call. - Get the final answer or another call. The model uses the result to answer, or requests further tools.
A few controls shape this behavior. By default the model chooses whether to call a tool; developers can force a specific tool or any tool when needed. Some providers can return several calls at once for independent lookups. Strict modes make the model's arguments conform exactly to the schema. Some providers also run certain tools themselves, such as web search, so the application sees only the results.
Why Tool Calling Matters
On its own, a language model can only produce text from what it already knows and what is in its context. Tool calling connects it to live data and real actions: today's numbers, a customer record, a file in the repository, a test run. It is the mechanism underneath every AI agent, which is essentially a loop of deciding, calling tools and reading results.
The quality of tool use depends heavily on tool design. Clear names, precise descriptions and tight schemas reduce wrong calls. Safety stays with the application: Anthropic's documentation notes that when required details are missing, a model may infer a reasonable-looking value rather than ask, so applications should validate arguments, limit what each tool can touch and require confirmation for actions that are destructive or costly.
Tool Calling Example
A product analytics assistant has a tool called get_feature_usage that takes a feature ID and a date range. A product manager asks, "How many teams used bulk export last month?" The model returns a call with feature_id: "bulk_export" and last month's dates. The application queries the analytics warehouse and returns "142 teams." The model answers with that figure and offers to compare it with the previous month to track feature adoption. The model never touched the warehouse; it only asked.
Tool Calling vs. MCP
Tool calling is the model-level capability and API pattern. The Model Context Protocol (MCP) is a standard for how applications discover and connect to tools, data and prompts offered by external servers. An application can define tools directly in its own code, or collect them from MCP servers and pass them to the model in the same way. Either way, the model's side of the exchange is still a tool call.