Local AI

How do you run AI agents locally?

Run a local model with reliable tool calling behind an OpenAI-compatible endpoint, then build the plan-act-observe loop in your own code: the model proposes a tool call, your executor runs it in a sandbox, and the result returns to context. Keep memory layered and route tasks the local model cannot handle to a hosted API.

Key facts

RuntimeOllama, LM Studio, llama.cpp server or vLLM expose OpenAI-compatible endpoints
Model requirementTool calling needs a function-calling capable model
LoopPlan, act, observe, repeat until the answer or step budget
SandboxAllowlist commands and isolate execution from the host
MemoryWorking state in context; durable facts in a vector store
Failure handlingRetries, step limits and timeouts prevent runaway loops
Hosted fallbackPlugsky agents and function calling are live over one API
Coming soonAssistants, responses and batch endpoints

TL;DR

  • Local agents run the same loop as hosted agents, on your hardware.
  • Tool-calling reliability matters more than raw model size.
  • Sandbox every tool, because local inference does not make execution safe.
  • Layer memory: context for now, vectors for later.
  • Add a hosted fallback instead of forcing one machine to do everything.

How it works, step by step

  1. Define the agent's tools, inputs and stop conditions.
  2. Choose a local model with verified function calling.
  3. Expose it on an OpenAI-compatible endpoint and test one tool.
  4. Implement the loop with retries, timeouts and a step budget.
  5. Sandbox tool execution and restrict filesystem and network access.
  6. Add memory for durable facts and log every decision.
  7. Route failed or heavy tasks to a hosted OpenAI-compatible endpoint.
1Define the agent'stools, inputs andstop conditions.2Choose a localmodel with verifiedfunction calling.3Expose it on anOpenAI-compatibleendpoint and test4Implement the loopwith retries,timeouts and a step5Sandbox toolexecution andrestrict filesystem6Add memory fordurable facts andlog every decision.

Try it yourself

Open the AI agent builder →

The local agent loop

An agent is a control loop around a model. You give it a goal and a set of tool schemas. The model returns a structured tool call, your code executes it, and the observation goes back into context for the next turn. The loop ends when the model returns a final answer or hits a step budget.

Running that loop locally changes only where inference happens. What changes is the constraint set: local models are smaller, context is more expensive in memory, and tool-call reliability varies more between models. That makes model selection and JSON validation the difference between a working agent and a loop of parse failures.

Reliability, safety and memory

Start with one tool and one scripted prompt. Only add complexity after function calling works reliably, because debugging a multi-tool agent on a model that emits malformed JSON is wasted effort.

  • Validate tool calls against a schema and fail fast with a clear error.
  • Sandbox execution: allowlist commands, mount data read-only and run tools in a container.
  • Bound the loop: step limits, timeouts and a per-task ceiling.
  • Layer memory: keep recent turns in context and write durable facts to a vector store.
  • Log everything: prompts, tool calls, arguments and observations, so failures are reproducible.

Scaling and hybrid routing

A single machine handles a personal agent or a small internal tool well. It struggles with concurrency, very long contexts and hard reasoning, because all three consume the same fixed resources.

Hybrid routing is the pragmatic answer: keep private or offline steps local, and send hard or parallel tasks to a hosted endpoint. Plugsky serves agents and function calling over an OpenAI-compatible API with 30+ models, plus chat, streaming, JSON mode, embeddings and RAG, all live; assistants, responses and batch endpoints are coming soon. Because the interface is standard, routing is configuration rather than a rewrite. See pricing for plans and start free with plugsky-micro and plugsky-lite.

Honest comparison

ConcernLocal agentHybrid local + PlugskyPure cloud agent
Data exposurePrompts stay on the machineSensitive steps stay localPrompts leave your network
Model ceilingLimited by local memoryLocal model plus 30+ hosted modelsHosted models on demand
Cost shapeHardware and powerHardware plus one cloud planSubscription or per-token
OfflineWorks with no networkDegrades to local onlyFails without connectivity
OperationsYou run and monitor the loopSplit responsibilityProvider-managed inference

Frequently asked questions

Can local models call tools reliably?

Yes, if the model supports function calling and the runtime parses its tool-call format. Verify with a single tool before building a multi-step agent.

How much memory does a local agent need?

A 7B-8B model at 4-bit occupies roughly 4-5 GB of weights plus KV cache. Multi-step agents benefit from more headroom because context grows with each observation.

Why sandbox a local agent?

Tool execution carries the same risk locally as in the cloud. A shell or file tool can damage data or leak credentials regardless of where inference runs.

How do I stop runaway loops?

Set a step budget, timeouts and a per-task cost or token ceiling, and log each iteration so you can see where agents stall.

What memory should an agent have?

Keep recent state in the prompt and store durable facts in a vector store you query deliberately. Do not treat the context window as a database.

Is a local agent private?

Inference and data can stay local, but tools may call the network. Restrict egress and review each tool's reach if privacy is the goal.

When should I use a hosted model?

For hard reasoning, large context or concurrency. The same agent code can call an OpenAI-compatible endpoint, so routing is a configuration choice.