Key facts
| Mechanism | The model returns tool_calls; your code executes them |
| Schema | JSON Schema parameters with names and descriptions written for the model |
| Loop | Execute, append the result, call again until finish_reason is stop |
| Validation | Validate arguments and enforce permissions before execution |
| Parallel calls | Modern models can request several tool calls in a single turn |
| Structured output | JSON mode constrains responses to valid JSON |
| Models | Function calling is live across modern models in the 30+ catalogue |
| Roadmap | Assistants-style managed endpoints are coming soon |
TL;DR
- Function calling is structured output: the model proposes, your code executes.
- Write tool descriptions for the model, not for a human developer.
- Validate arguments and enforce permissions before running anything.
- Append tool results as messages and loop until the model stops asking.
- Handle parallel tool calls and tool errors without breaking the loop.
How it works, step by step
- Identify two to five actions the agent must perform and give each a precise name.
- Describe each tool for the model: what it does, when to use it and what it returns.
- Define JSON Schema parameters with types, enums and required fields.
- Send the tools array with your messages and read the returned tool_calls.
- Validate arguments, authorize the call, execute it and append a tool message per call.
- Return tool errors as readable text so the model can adapt instead of crashing.
- Cap turns, log every call and evaluate the trajectory before shipping.
Try it yourself
Open the function calling tester →
What function calling actually is
Function calling is not the model executing code. It is a structured output mode: you provide a list of available functions in the request, and the model may respond with a machine-readable call naming a function and arguments. Nothing happens until your code decides to run it. That separation is the whole safety story — the model can propose, but only your application can act.
Because the response is structured, the pattern works across chat, streaming and JSON mode, and it composes with any orchestration framework. On Plugsky the tool loop runs on the live OpenAI-compatible /v1/chat/completions endpoint, so existing OpenAI-style code works after a base URL change.
Designing tools the model can use
Tool quality determines agent quality. A vague description produces wrong calls no matter how strong the model is. Write descriptions as instructions: what the tool returns, when to choose it over a sibling tool, and any constraints. Prefer narrow, single-purpose tools over broad ones that require arguments the model cannot know.
- Names: verb plus object, consistent across the tool set.
- Parameters: typed, required where needed, with enums for closed sets.
- Results: small and structured; five relevant fields beat a full record.
- Errors: return them as data so the model can retry or choose another path.
The loop, retries and safety
The loop is mechanical but easy to get wrong. After the model returns tool calls, execute each one, append one tool message per call with the matching ID, then call the endpoint again. Stop when the model responds normally or you hit the turn cap. Parallel tool calls should be executed concurrently where they are independent, with per-call timeouts.
Safety lives outside the model: allowlist which tools each agent may call, validate arguments against the schema, authorize every call server-side, cap rate and spend, and log the full trajectory. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live on Plugsky; audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon. Start on the free plan with plugsky-micro and plugsky-lite, or check the live pricing page for paid tiers.
Honest comparison
| Aspect | Function calling | Prompt-only agent | Hard-coded workflow |
|---|---|---|---|
| Action choice | Model selects from declared tools | Model writes text you parse | Developer decides |
| Reliability | Structured arguments | Fragile parsing | Deterministic |
| Flexibility | High within the tool set | High but unsafe | Low |
| Validation | Schema before execution | Manual regex or none | Not needed |
| Best for | Open-ended tasks with clear tools | Prototypes only | Known, fixed processes |
Frequently asked questions
Does the model run my code?
No. The model returns a structured request naming a tool and arguments. Your application decides whether to execute it, and only your code touches real systems.
How many tools should an agent have?
Start with two to five well-described tools. Large tool sets confuse selection and inflate every prompt, so add tools only when evaluation shows a real gap.
Can the model call several tools at once?
Yes. Modern models can return multiple tool calls in one turn, which you can execute in parallel when they are independent.
What happens if a tool fails?
Return the error as a readable tool message. The model can retry, choose another tool or explain the failure. Also log it, because repeated tool errors usually indicate a schema or prompt problem.
Is function calling available on all models?
It is available on modern models in the catalogue. Check the model reference for per-model tool support before depending on it in production.
How do I test tools before shipping?
Use the function calling tester to check schemas and argument generation, then run a trajectory evaluation suite on real tasks.
Do I need a special endpoint?
No. Function calling works on the live OpenAI-compatible chat completions endpoint. Assistants-style managed endpoints are coming soon.