Key facts
| Runtime | Ollama, LM Studio, llama.cpp server or vLLM expose OpenAI-compatible endpoints |
| Model requirement | Tool calling needs a function-calling capable model |
| Loop | Plan, act, observe, repeat until the answer or step budget |
| Sandbox | Allowlist commands and isolate execution from the host |
| Memory | Working state in context; durable facts in a vector store |
| Failure handling | Retries, step limits and timeouts prevent runaway loops |
| Hosted fallback | Plugsky agents and function calling are live over one API |
| Coming soon | Assistants, responses and batch endpoints |
TL;DR
- Local agents run the same loop as hosted agents, on your hardware.
- Tool-calling reliability matters more than raw model size.
- Sandbox every tool, because local inference does not make execution safe.
- Layer memory: context for now, vectors for later.
- Add a hosted fallback instead of forcing one machine to do everything.
How it works, step by step
- Define the agent's tools, inputs and stop conditions.
- Choose a local model with verified function calling.
- Expose it on an OpenAI-compatible endpoint and test one tool.
- Implement the loop with retries, timeouts and a step budget.
- Sandbox tool execution and restrict filesystem and network access.
- Add memory for durable facts and log every decision.
- Route failed or heavy tasks to a hosted OpenAI-compatible endpoint.
Try it yourself
The local agent loop
An agent is a control loop around a model. You give it a goal and a set of tool schemas. The model returns a structured tool call, your code executes it, and the observation goes back into context for the next turn. The loop ends when the model returns a final answer or hits a step budget.
Running that loop locally changes only where inference happens. What changes is the constraint set: local models are smaller, context is more expensive in memory, and tool-call reliability varies more between models. That makes model selection and JSON validation the difference between a working agent and a loop of parse failures.
Reliability, safety and memory
Start with one tool and one scripted prompt. Only add complexity after function calling works reliably, because debugging a multi-tool agent on a model that emits malformed JSON is wasted effort.
- Validate tool calls against a schema and fail fast with a clear error.
- Sandbox execution: allowlist commands, mount data read-only and run tools in a container.
- Bound the loop: step limits, timeouts and a per-task ceiling.
- Layer memory: keep recent turns in context and write durable facts to a vector store.
- Log everything: prompts, tool calls, arguments and observations, so failures are reproducible.
Scaling and hybrid routing
A single machine handles a personal agent or a small internal tool well. It struggles with concurrency, very long contexts and hard reasoning, because all three consume the same fixed resources.
Hybrid routing is the pragmatic answer: keep private or offline steps local, and send hard or parallel tasks to a hosted endpoint. Plugsky serves agents and function calling over an OpenAI-compatible API with 30+ models, plus chat, streaming, JSON mode, embeddings and RAG, all live; assistants, responses and batch endpoints are coming soon. Because the interface is standard, routing is configuration rather than a rewrite. See pricing for plans and start free with plugsky-micro and plugsky-lite.
Honest comparison
| Concern | Local agent | Hybrid local + Plugsky | Pure cloud agent |
|---|---|---|---|
| Data exposure | Prompts stay on the machine | Sensitive steps stay local | Prompts leave your network |
| Model ceiling | Limited by local memory | Local model plus 30+ hosted models | Hosted models on demand |
| Cost shape | Hardware and power | Hardware plus one cloud plan | Subscription or per-token |
| Offline | Works with no network | Degrades to local only | Fails without connectivity |
| Operations | You run and monitor the loop | Split responsibility | Provider-managed inference |
Frequently asked questions
Can local models call tools reliably?
Yes, if the model supports function calling and the runtime parses its tool-call format. Verify with a single tool before building a multi-step agent.
How much memory does a local agent need?
A 7B-8B model at 4-bit occupies roughly 4-5 GB of weights plus KV cache. Multi-step agents benefit from more headroom because context grows with each observation.
Why sandbox a local agent?
Tool execution carries the same risk locally as in the cloud. A shell or file tool can damage data or leak credentials regardless of where inference runs.
How do I stop runaway loops?
Set a step budget, timeouts and a per-task cost or token ceiling, and log each iteration so you can see where agents stall.
What memory should an agent have?
Keep recent state in the prompt and store durable facts in a vector store you query deliberately. Do not treat the context window as a database.
Is a local agent private?
Inference and data can stay local, but tools may call the network. Restrict egress and review each tool's reach if privacy is the goal.
When should I use a hosted model?
For hard reasoning, large context or concurrency. The same agent code can call an OpenAI-compatible endpoint, so routing is a configuration choice.