Local AI

How do you build a local AI agent with tools?

A local agent with tools is a loop: the model receives tool schemas, returns a structured call, your code validates and executes it, and the observation goes back into context. Local models handle this well when schemas are small and typed and the runtime parses tool calls correctly. Safety comes from execution controls, not the model, so validate arguments, sandbox tools and log every step.

Key facts

Loop shapeModel call, tool call, execution, observation, next turn
Schema formatOpenAI-style tools array with JSON Schema parameters
Runtime requirementllama.cpp, Ollama and vLLM parse tool calls with family-appropriate settings
Model requirementFunction-calling training, not just instruction tuning
Failure modesMalformed arguments, wrong tool choice, missing calls and loops
Safety controlsArgument validation, allowlists, sandboxes, confirmation gates and timeouts
Cloud optionPlugsky serves tool calling and agents live with 30+ models

TL;DR

  • Schema design is the highest-leverage part of a local tool agent.
  • Small typed schemas with enums outperform large free-form parameter lists.
  • Validate and sandbox every call; the model is not a security boundary.
  • Cap steps and retries so failures end instead of looping.
  • Log raw output and parsed calls together for fast debugging.

How it works, step by step

  1. List the tools the agent needs and the exact inputs each requires.
  2. Write JSON Schemas with types, enums and descriptions; avoid free-text where a list works.
  3. Serve a function-calling model and test one tool in isolation.
  4. Implement the loop with parsing, validation and structured error responses.
  5. Add safety controls: allowlists, sandboxes, timeouts and confirmation for destructive actions.
  6. Set step and retry budgets, then test multi-tool workflows and failure recovery.
  7. Route consistently failing tool classes to a hosted model instead of forcing the local one.
1List the tools theagent needs and theexact inputs each2Write JSON Schemaswith types, enumsand descriptions;3Serve afunction-callingmodel and test one4Implement the loopwith parsing,validation and5Add safetycontrols:allowlists,6Set step and retrybudgets, then testmulti-tool

Try it yourself

Open the function calling tester →

Writing schemas local models actually follow

Local models at 7B-14B follow compact, explicit schemas far better than sprawling ones. Give each tool one clear purpose, name parameters descriptively, and use types and enums rather than free-form strings. A date should be a string with a stated format; a category should be an enum; nested objects should stay shallow.

Descriptions do real work: they tell the model when to use the tool, not just what it accepts. Include when-not-to-use guidance for tools that overlap. If two tools can plausibly answer the same request, expect inconsistent selection until you sharpen their boundaries. And include a no-tool path, so the model can decline rather than inventing a call.

Running the loop reliably

Parsing is where most local setups break. Confirm the runtime converts the model's native output into structured calls; if it does not, calls arrive as prose and the loop stalls. Once parsing works, treat every call as untrusted: validate against the schema, reject unknown fields, coerce types carefully and return a compact structured error when something is wrong.

  • Cap steps per task so runaway loops terminate.
  • Retry a failed call once with the error message appended.
  • Deduplicate identical consecutive calls to break stuck cycles.
  • Keep observations short; large outputs crowd the context.

Multi-tool workflows need sequential calls in most local models. Parallel tool calls exist but support varies, so design for sequential execution and add parallelism only after testing.

Safety, evaluation and escalation

The model proposes; your code decides. Sandbox tools that touch the filesystem or shell, allowlist commands, mount data read-only where possible, and require human confirmation for irreversible actions. Keep credentials out of the prompt and out of logs. Retrieved documents and tool outputs can carry injection attempts, so never let untrusted text expand the tool surface.

Evaluate with a scripted set: valid calls, missing arguments, out-of-schema requests, ambiguous instructions and refusal cases. Track parse success and argument accuracy per model, because a smaller model with better schema compliance often beats a larger one in production. When a tool class consistently fails locally, send that step to a hosted model. Plugsky runs function calling, agents, JSON mode and streaming live over an OpenAI-compatible API with 30+ models; audio, image, moderation, batch and fine-tuning endpoints are coming soon. See the live pricing page for plan details.

Honest comparison

ConcernWeak approachStrong approachWhy
SchemaOne large free-form toolSeveral small typed toolsLocal models follow narrow contracts
ValidationTrust model outputValidate against schema firstModel output is untrusted input
ErrorsCrash or ignoreStructured error back to modelEnables self-correction
LoopsUnlimited stepsStep and retry capsPrevents runaway execution
ExecutionDirect system accessSandbox plus allowlistLimits blast radius

Frequently asked questions

Do local models support multiple tools?

Yes, but reliability drops as tool count and schema complexity grow. Keep the active tool set small and add tools only when testing shows they are used correctly.

Why does the model call the wrong tool?

Usually because tool descriptions overlap or the model lacks function-calling training. Sharpen boundaries and verify the model family supports tool use.

How do I handle invalid JSON from a local model?

Enable JSON mode or grammar-constrained decoding where the runtime supports it, then validate anyway and return a structured error for a single retry.

Can tools run in parallel?

Some runtimes and models support parallel calls, but support varies. Design for sequential execution first and add concurrency only after testing.

Is it safe to give an agent shell access?

Only inside a sandbox with an allowlist of commands, working-directory restrictions and confirmation for destructive operations. Never grant unrestricted shell access.

How do I test a tool agent?

Build a scripted suite covering valid calls, missing arguments, ambiguous requests and refusals, then score parse success and argument accuracy per model.

What should I log?

The raw model output, the parsed call, validation results, tool results and timing. That combination makes failures reproducible.

When should I use a hosted model instead?

When a tool class fails repeatedly locally, when planning is multi-step and ambiguous, or when concurrency exceeds local capacity. Keep the same tool schemas.