Local AI

What are local AI agents and how do you deploy them?

Local AI agents are autonomous loops that run a language model on hardware you control, calling tools and carrying out multi-step tasks. A working deployment needs four parts: a tool-capable model, a runtime that parses tool calls, an orchestrator that executes and logs steps, and memory for state. Local agents fit privacy, offline and high-volume automation; hybrid routing handles tasks beyond local capability.

Key facts

Core componentsModel, runtime, orchestrator, tool layer and memory store
Model requirementFunction-calling training plus enough context for tool observations
Runtime optionsOllama, llama.cpp server and vLLM expose OpenAI-compatible endpoints
Memory designWorking state in context; durable state in a vector or relational store
Security modelSandboxed tool execution, allowlists, confirmation gates and audit logs
Cloud escalationPlugsky runs agents live with 30+ models over an OpenAI-compatible API
Endpoint statusChat, streaming, JSON mode, function calling, embeddings and RAG are live

TL;DR

  • Treat an agent as a distributed system: model, runtime, tools, memory and audit.
  • Model choice is about reliable tool calls, not maximum size.
  • Capability comes from tools; safety comes from how you execute them.
  • Start with one agent and a handful of tools before adding orchestration layers.
  • Plan escalation from day one so hard tasks do not stall the workflow.

How it works, step by step

  1. Choose a bounded task with measurable success criteria and a real user.
  2. Select a function-calling model that fits your memory budget and context needs.
  3. Serve it with a runtime that parses tool calls and exposes an OpenAI-compatible API.
  4. Define a small tool set with typed schemas and explicit failure behaviour.
  5. Implement the orchestration loop with validation, retries, step caps and full logging.
  6. Add memory: conversation state in context, durable facts and documents in a store with retrieval.
  7. Deploy with sandboxing and access control, then monitor quality and route failures to a hosted model.
1Choose a boundedtask withmeasurable success2Select afunction-callingmodel that fits3Serve it with aruntime that parsestool calls and4Define a small toolset with typedschemas and5Implement theorchestration loopwith validation,6Add memory:conversation statein context, durable

Try it yourself

Open the AI agent builder →

The architecture of a local agent

An agent is a model wrapped in a control loop. The orchestrator assembles a prompt from the goal, the tool schemas and recent observations, calls the model, parses the response, executes any requested tool and appends the result. That cycle repeats until the task is complete, a step budget is exhausted or a human intervenes.

Four subsystems surround the loop. The tool layer defines what the agent can do and validates every request. The memory layer keeps working state in context and durable state in a store. The observability layer records inputs, calls, outputs and timings. The policy layer enforces what is allowed without human approval. Skipping any of these turns a demo into an incident.

Model, runtime and memory choices

Pick the model on structured-output reliability, then on context, then on size. A 7B-14B function-calling model that always emits valid calls outperforms a larger model that occasionally wraps a call in prose. Context must hold the system prompt, tool schemas, retrieved passages and accumulated observations, so long workflows need either summarisation or a model with grouped-query attention and a manageable cache.

  • Runtime: Ollama for simplicity, llama.cpp for quantization control, vLLM for concurrent serving.
  • Memory: summarise old turns, store facts in a vector store, and retrieve narrowly.
  • Tools: prefer many small typed tools over a few broad ones.

Security, evaluation and hybrid escalation

Local deployment protects data in transit and at rest from third parties, but an agent that can act is a new risk surface. Sandbox execution, allowlist commands and paths, require confirmation for irreversible actions, keep secrets out of prompts, and treat retrieved content as untrusted. Audit logs should capture the raw model output, parsed calls and results so incidents are reconstructable.

Evaluate on a scripted suite that includes ambiguous requests, tool failures and refusal cases. Track success rate and steps per task, then set an escalation policy: when a local agent fails repeatedly or the task exceeds its context, route that step to a hosted model. Plugsky runs agents, function calling, JSON mode, streaming, embeddings and RAG live behind an OpenAI-compatible endpoint with 30+ models, plus cloud, VPC, on-prem and air-gapped deployment options. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. See the live pricing page before choosing a plan.

Honest comparison

DeploymentModel locationBest forOperations
Fully local agentOn your hardwarePrivacy and offline automationYou operate everything
Local agent with hosted escalationLocal plus 30+ cloud modelsMixed difficulty workloadsSplit responsibilities
Hosted agent platformProvider infrastructureFast rollout and scaleManaged
Air-gapped agentIsolated networkClassified environmentsFull ownership, offline updates

Frequently asked questions

What makes an agent different from a chatbot?

A chatbot answers; an agent acts. It plans steps, calls tools, observes results and continues until the task completes or a limit is reached.

Do local agents need a GPU?

For interactive multi-step tasks, a GPU or Apple Silicon machine is strongly preferred. CPU-only agents run but are slow, which compounds across many loop iterations.

How much memory should I plan for?

Start with weights plus KV cache at your maximum context and concurrency. A 7B model at 4-bit needs roughly 4-5 GB of weights, plus cache that grows with tokens.

Can local agents be fully private?

Yes, if the model, tools and data are local. Any tool that calls an external service breaks isolation, so audit the tool set for network access.

How many tools should an agent have?

Fewer than you think. Start with five or fewer well-described tools, measure selection accuracy, and add tools only when testing justifies them.

How do I handle tasks the local model cannot do?

Route those steps to a hosted OpenAI-compatible API while keeping the same orchestration and tool schemas. Plugsky supports agents live across 30+ models.

How do I monitor agent quality?

Log every step and score success rate, steps per task and tool-error rate on a fixed evaluation suite. Track regressions whenever you change model or prompt.

Can agents run on-prem or air-gapped?

Yes. Local agents run entirely on your infrastructure, and Plugsky supports cloud, VPC, on-prem and air-gapped deployment for enterprise teams.