Local AI

What is the best local LLM for tool calling?

The best local model for tool calling is one trained on function calling and served by a runtime that parses its tool-call format. Qwen, Llama 3.x, Mistral, gpt-oss and Hermes-style models are common choices, with llama.cpp, Ollama and vLLM providing the calling mechanics. Schema design and validation matter as much as model choice, because most failures are malformed arguments, not missing capability.

Key facts

Model familiesQwen, Llama 3.x, Mistral, gpt-oss and Hermes-style function-calling models
Runtime supportllama.cpp, Ollama and vLLM parse tool calls; vLLM needs a parser flag per family
Schema formatOpenAI-style tools array with JSON Schema parameters
JSON modeConstrained decoding or JSON mode reduces malformed output
Common failureWrong argument types, invented parameter names and calls outside the schema
Cloud optionPlugsky serves function calling and agents live with 30+ models
Free tierplugsky-micro and plugsky-lite on the free plan, plus a 14-day full-access trial

TL;DR

  • Pick a model explicitly trained for function calling before tuning anything else.
  • Confirm your runtime parses the model's tool-call format; formats are not interchangeable.
  • Keep schemas small, with typed parameters and clear descriptions.
  • Validate every tool call before execution and return structured errors to the model.
  • Test parse reliability per model; one bad schema can corrupt a whole workflow.

How it works, step by step

  1. List each tool, its arguments, types and failure behaviour.
  2. Write JSON Schemas with descriptions and enums instead of free-text parameters.
  3. Choose two or three function-calling models that fit your memory budget.
  4. Serve them with a runtime that parses tool calls and enable JSON mode if available.
  5. Run a scripted test set that includes valid calls, missing arguments and out-of-schema requests.
  6. Score parse success, argument accuracy and recovery after an invalid call.
  7. Deploy the winner behind a validator, then add retries and a cloud fallback for hard cases.
1List each tool, itsarguments, typesand failure2Write JSON Schemaswith descriptionsand enums instead3Choose two or threefunction-callingmodels that fit4Serve them with aruntime that parsestool calls and5Run a scripted testset that includesvalid calls,6Score parsesuccess, argumentaccuracy and

Original data

Qwen, Llama 3.Model familiesPlugsky servesCloud optionplugsky-micro Free tierSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the function calling tester →

How local tool calling works

The model does not execute anything. It receives a list of tool definitions and returns a structured request, which your code validates and runs. The runtime's job is translating between the model's native output format and the tool-call shape your application expects.

Formats differ between model families. Some emit JSON blocks, others use special tokens or XML-like tags. llama.cpp and Ollama handle common conventions directly; vLLM requires enabling auto tool choice and selecting a parser for the family you serve. If the parser does not match the model, calls degrade into prose and the loop stalls.

Choosing and testing a local model

Model size is less important than training signal for this task. A 7B-14B model trained on function calling typically outperforms a larger general model at emitting valid calls, and it runs faster. Test with your own schemas rather than public examples, because argument complexity drives failure rates.

  • Include one tool with nested object parameters.
  • Include one tool with an enum and strict value set.
  • Test a request that should trigger no tool call at all.
  • Test recovery: send back an error and see if the model corrects the call.

Grammar-constrained decoding and JSON mode further reduce malformed output, but they do not fix bad schema design. Typed parameters with clear descriptions do more than any decoder setting.

Hardening the tool loop

Treat every model-emitted call as untrusted input. Validate against the schema, reject unknown parameters, enforce allowlists for commands and paths, and set timeouts on every tool. Return compact structured errors so the model can retry with corrected arguments, and cap retries to avoid infinite loops.

Log inputs, calls and observations so you can replay failures. Add step and token budgets, and require human confirmation for destructive operations. When a local model consistently fails a class of calls, route just that task to a hosted model. Plugsky runs function calling, agents, JSON mode, streaming and chat live behind an OpenAI-compatible endpoint with 30+ models; audio, image, moderation, batch and fine-tuning endpoints are coming soon. Plan details are on the live pricing page.

Honest comparison

SetupTool-call reliabilitySpeedBest for
Ollama plus a function-calling modelGood for common formatsFast on GPUQuick local automation
llama.cpp plus grammar constraintsHigh with constrained decodingHardware-dependentStrict output control
vLLM with a tool-call parserGood at scaleHigh throughputMulti-user services
Plugsky function callingManaged, per-model coverageNetwork-boundProduction agents

Frequently asked questions

Which local models support function calling?

Look for models trained on tool use, such as recent Qwen, Llama 3.x, Mistral and gpt-oss releases. Support varies by version, so verify the documentation for the exact weights you download.

Why does my local model return prose instead of a tool call?

Usually the runtime is not parsing the model's tool-call format, or the prompt does not present tools in the expected structure. Match the parser to the model family.

Is JSON mode enough for tool calling?

No. JSON mode guarantees well-formed JSON, not schema-correct arguments. You still need validation and retries.

How big a model do I need for tool calling?

A 7B-14B function-calling model handles most single-tool and two-tool workflows. Larger models help with multi-step planning and ambiguous requests.

Can local models call multiple tools in one turn?

Some can emit parallel calls, but support varies. If your workflow needs it, test explicitly and fall back to sequential calls when unsupported.

How do I stop the model from calling dangerous tools?

Do not let the model decide execution. Validate arguments, allowlist operations, require confirmation and run tools in a sandbox.

What is the best way to debug tool calling?

Log the raw model output alongside the parsed call. Most failures are visible in the delta between what the model emitted and what your parser extracted.

Does Plugsky support function calling?

Yes. Function calling and agents are live on Plugsky with 30+ models, so the same tools can run locally or hosted.