Key facts
| Model families | Qwen, Llama 3.x, Mistral, gpt-oss and Hermes-style function-calling models |
| Runtime support | llama.cpp, Ollama and vLLM parse tool calls; vLLM needs a parser flag per family |
| Schema format | OpenAI-style tools array with JSON Schema parameters |
| JSON mode | Constrained decoding or JSON mode reduces malformed output |
| Common failure | Wrong argument types, invented parameter names and calls outside the schema |
| Cloud option | Plugsky serves function calling and agents live with 30+ models |
| Free tier | plugsky-micro and plugsky-lite on the free plan, plus a 14-day full-access trial |
TL;DR
- Pick a model explicitly trained for function calling before tuning anything else.
- Confirm your runtime parses the model's tool-call format; formats are not interchangeable.
- Keep schemas small, with typed parameters and clear descriptions.
- Validate every tool call before execution and return structured errors to the model.
- Test parse reliability per model; one bad schema can corrupt a whole workflow.
How it works, step by step
- List each tool, its arguments, types and failure behaviour.
- Write JSON Schemas with descriptions and enums instead of free-text parameters.
- Choose two or three function-calling models that fit your memory budget.
- Serve them with a runtime that parses tool calls and enable JSON mode if available.
- Run a scripted test set that includes valid calls, missing arguments and out-of-schema requests.
- Score parse success, argument accuracy and recovery after an invalid call.
- Deploy the winner behind a validator, then add retries and a cloud fallback for hard cases.
Original data
Try it yourself
Open the function calling tester →
How local tool calling works
The model does not execute anything. It receives a list of tool definitions and returns a structured request, which your code validates and runs. The runtime's job is translating between the model's native output format and the tool-call shape your application expects.
Formats differ between model families. Some emit JSON blocks, others use special tokens or XML-like tags. llama.cpp and Ollama handle common conventions directly; vLLM requires enabling auto tool choice and selecting a parser for the family you serve. If the parser does not match the model, calls degrade into prose and the loop stalls.
Choosing and testing a local model
Model size is less important than training signal for this task. A 7B-14B model trained on function calling typically outperforms a larger general model at emitting valid calls, and it runs faster. Test with your own schemas rather than public examples, because argument complexity drives failure rates.
- Include one tool with nested object parameters.
- Include one tool with an enum and strict value set.
- Test a request that should trigger no tool call at all.
- Test recovery: send back an error and see if the model corrects the call.
Grammar-constrained decoding and JSON mode further reduce malformed output, but they do not fix bad schema design. Typed parameters with clear descriptions do more than any decoder setting.
Hardening the tool loop
Treat every model-emitted call as untrusted input. Validate against the schema, reject unknown parameters, enforce allowlists for commands and paths, and set timeouts on every tool. Return compact structured errors so the model can retry with corrected arguments, and cap retries to avoid infinite loops.
Log inputs, calls and observations so you can replay failures. Add step and token budgets, and require human confirmation for destructive operations. When a local model consistently fails a class of calls, route just that task to a hosted model. Plugsky runs function calling, agents, JSON mode, streaming and chat live behind an OpenAI-compatible endpoint with 30+ models; audio, image, moderation, batch and fine-tuning endpoints are coming soon. Plan details are on the live pricing page.
Honest comparison
| Setup | Tool-call reliability | Speed | Best for |
|---|---|---|---|
| Ollama plus a function-calling model | Good for common formats | Fast on GPU | Quick local automation |
| llama.cpp plus grammar constraints | High with constrained decoding | Hardware-dependent | Strict output control |
| vLLM with a tool-call parser | Good at scale | High throughput | Multi-user services |
| Plugsky function calling | Managed, per-model coverage | Network-bound | Production agents |
Frequently asked questions
Which local models support function calling?
Look for models trained on tool use, such as recent Qwen, Llama 3.x, Mistral and gpt-oss releases. Support varies by version, so verify the documentation for the exact weights you download.
Why does my local model return prose instead of a tool call?
Usually the runtime is not parsing the model's tool-call format, or the prompt does not present tools in the expected structure. Match the parser to the model family.
Is JSON mode enough for tool calling?
No. JSON mode guarantees well-formed JSON, not schema-correct arguments. You still need validation and retries.
How big a model do I need for tool calling?
A 7B-14B function-calling model handles most single-tool and two-tool workflows. Larger models help with multi-step planning and ambiguous requests.
Can local models call multiple tools in one turn?
Some can emit parallel calls, but support varies. If your workflow needs it, test explicitly and fall back to sequential calls when unsupported.
How do I stop the model from calling dangerous tools?
Do not let the model decide execution. Validate arguments, allowlist operations, require confirmation and run tools in a sandbox.
What is the best way to debug tool calling?
Log the raw model output alongside the parsed call. Most failures are visible in the delta between what the model emitted and what your parser extracted.
Does Plugsky support function calling?
Yes. Function calling and agents are live on Plugsky with 30+ models, so the same tools can run locally or hosted.