Key facts
| Runtime options | Ollama, LM Studio, llama.cpp server and vLLM all expose an OpenAI-compatible endpoint |
| Agent loop | Model output to tool call to observation to next step |
| Tool calling | Needs a function-calling model plus a runtime that parses tool calls |
| Memory | Working context lives in the KV cache; long-term facts go in a vector store |
| Hardware floor | A 7B-8B model at 4-bit needs roughly 4-5 GB of weights before KV cache |
| Cloud fallback | Plugsky exposes an OpenAI-compatible API with 30+ models for hybrid routing |
| Endpoint status | Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live |
TL;DR
- Local agents run the same plan-act-observe loop as cloud agents, just on your hardware.
- Pick a model with reliable tool calling; small is fine when the task graph is simple.
- Sandbox tool execution, because local inference does not make tool calls safe.
- Use an OpenAI-compatible local server so agent code can move between local and cloud.
- Route hard reasoning to a cloud API in hybrid mode instead of forcing one machine to do everything.
How it works, step by step
- Define the agent's tools and success criteria before choosing a model.
- Install a local runtime such as Ollama, LM Studio or llama.cpp and pull a tool-capable model.
- Verify function calling with one tool and a scripted prompt before adding complexity.
- Add memory: keep conversation state in the prompt and durable facts in a vector store.
- Sandbox tool execution with restricted permissions and a command allowlist.
- Log every model decision, tool call and observation for debugging and audit.
- Add a cloud fallback for tasks the local model fails or when load exceeds capacity.
Try it yourself
How a local agent loop actually works
An agent is a control loop around a language model. The model receives a goal plus tool schemas, emits a structured tool call, your code executes it, and the result is appended to the context for the next turn. That repeats until the model returns a final answer or a step budget runs out.
Locally, the difference is only where inference happens. A runtime hosts the weights and serves an HTTP API; your agent code is unchanged. The practical constraints shift to memory, context length and tool-calling reliability, which is why model choice matters more locally than it does against a frontier API.
Choosing a model and runtime for local agents
Tool calling is the first filter. Look for a model family trained on function calling, then confirm your runtime parses its tool-call format. llama.cpp and Ollama handle common formats directly, while vLLM needs a tool-call parser flag enabled for the model family you serve.
- Ollama is the fastest way to start: one command pulls a model and exposes
/v1. - llama.cpp server gives control over quantization and offload across CPU, Metal, CUDA and ROCm.
- vLLM is the choice for concurrent agent traffic because of continuous batching.
For agent work, a 7B-14B instruction model with strong structured output usually beats a larger model that emits malformed JSON, because every parse failure costs an extra loop.
Sandboxing, memory and hybrid fallback
Local inference removes network exposure, not risk. A tool that runs shell commands or writes files has the same power whether the model is local or remote, so restrict the tool surface: allowlist commands, mount read-only data, run tools in a container, and require confirmation for destructive actions. Treat retrieved documents as untrusted input, since prompt injection travels through RAG.
Memory has two layers. Short-term state is the message history in context; long-term memory is a vector store you query and write deliberately. Keep the retrieval step small so the context does not crowd out the task.
Finally, plan the fallback. When a local model stalls on a hard subtask, an OpenAI-compatible cloud API can take over without changing agent code. Plugsky serves chat, streaming, JSON mode, function calling, embeddings, RAG and agents live from one key; image, audio, moderation, batch and fine-tuning endpoints are coming soon, so keep those workloads on their current provider for now. See pricing for plan details.
Honest comparison
| Concern | Local-only agent | Hybrid local + Plugsky | Pure cloud agent |
|---|---|---|---|
| Data exposure | Nothing leaves the machine | Sensitive steps stay local | Prompts leave your network |
| Model ceiling | Limited by your memory | Local model plus 30+ cloud models | Frontier models on demand |
| Cost shape | Hardware, power and ops | Hardware plus one flat cloud plan | Per-token or flat plan |
| Offline operation | Works with no internet | Degrades to local only | Fails without connectivity |
| Operations | You patch, monitor and scale | Split responsibilities | Provider manages |
Frequently asked questions
Can local models call tools reliably?
Yes, if the model was trained for function calling and the runtime parses its tool-call format. Qwen, Llama 3.x, Mistral and gpt-oss families are common local choices.
How much VRAM does a local agent need?
A 7B-8B model at 4-bit occupies roughly 4-5 GB of weights, plus KV cache and runtime overhead. 8 GB of VRAM is a realistic floor for a simple single-tool agent.
Is a local agent private?
The model and prompt stay on your machine, but tool calls can still reach the network. Restrict egress and sandbox tools if privacy is the point.
Do I need a GPU?
No. llama.cpp runs on CPU and Apple Silicon, but multi-step agents feel much faster on a GPU or unified-memory Mac.
Can I move the same agent to the cloud later?
Yes. If your local server speaks the OpenAI-compatible /v1 surface, moving to Plugsky is a base URL and model-name change.
What is the best local agent runtime?
Ollama for quick starts, llama.cpp for quantization control, and vLLM for concurrent GPU serving. All three expose OpenAI-compatible APIs.
Does Plugsky run agents?
Yes. Agents, tool calling and RAG are live on Plugsky, so the same agent code can route to 30+ hosted models when local capacity is not enough.