Function Calling

A model can't run your code directly. Function calling gives it a way to ask for it to run: you describe a function — a name, a description, and the exact arguments it takes — and the model can decide to call it. That description is a schema: a structured, written definition of the function that the model reads instead of your actual code. Function calling isn't a separate mechanism from tool calling — the request-execute-respond loop that makes it work is the same one. What "function calling" names specifically is the function itself: the custom tool you wrote.

In this guide
  1. Why two names for close to the same thing
  2. Writing a function definition that actually works
  3. Write a function, or use a built-in tool?

Why two names for close to the same thing

"Function calling" is the older term, and it originally named the whole capability. Vendor-run capabilities like web search or code execution came along later, using that same request-and-response pattern. They needed a name too, and "tool calling" took over as the broader umbrella term. Today, OpenAI's own documentation defines a function as one specific type of tool: the custom, schema-defined kind, distinct from a built-in tool the vendor supplies. The distinction only matters when you need to be precise about which kind of tool you're talking about — in casual conversation, the two terms are still used interchangeably.

Writing a function definition that actually works

A model never reads your code. The name, description, and argument schema you write are the entire interface it has to decide whether and how to call your function. Treat them as documentation written for the model, not comments written for a coworker.

  • Name it by what it does. get_weather tells the model something; handle_request doesn't.
  • Write the description for a decision, not a reference manual. State plainly when this function should be called and any real constraint on using it — that's the information the model weighs before deciding to call it at all.
  • Constrain arguments where you can. An enum of allowed values rules out a whole category of wrong input; an open-ended string invites the model to guess at a format you never specified.
  • Keep required arguments genuinely required. Every optional argument is a value the model might have to invent on its own rather than ask for.

            // Open-ended — the model can send anything here
            { "unit": { "type": "string" } }

            // Constrained — only these two values are ever possible
            { "unit": { "type": "string", "enum": ["celsius", "fahrenheit"] } }
            

The first version leaves room for the model to send "Celsius", "C", or "degrees F" — anything that looks plausible — and your code has to handle whatever shows up. The second version makes those variations impossible before the call ever reaches your code.

Both major vendors also offer a stricter mode for exactly this reason. Anthropic's strict: true and OpenAI's strict field both force the model's generated arguments to match your schema exactly. That removes an entire category of malformed calls before your code ever runs them.

Write a function, or use a built-in tool?

If the vendor already offers a built-in for what you need — web search, code execution — reach for that first. It runs on the vendor's own infrastructure, and there's no schema or execution code for you to maintain. Write a custom function when the capability is specific to your own systems: your database, your internal API, anything no vendor could plausibly offer as a built-in. This is usually exactly what an AI agent needs in order to act on the systems that are actually yours.