Local AI

How do you run local AI agents on your own hardware?

A local AI agent runs the full plan-act-observe loop against a model on your own machine. A runtime such as Ollama, LM Studio, llama.cpp server or vLLM hosts the model, while your agent code calls tools, stores memory and executes steps. Local agents suit private, offline or high-volume workloads; hybrid setups keep sensitive steps local and send hard reasoning to a cloud API.

Key facts

Runtime optionsOllama, LM Studio, llama.cpp server and vLLM all expose an OpenAI-compatible endpoint
Agent loopModel output to tool call to observation to next step
Tool callingNeeds a function-calling model plus a runtime that parses tool calls
MemoryWorking context lives in the KV cache; long-term facts go in a vector store
Hardware floorA 7B-8B model at 4-bit needs roughly 4-5 GB of weights before KV cache
Cloud fallbackPlugsky exposes an OpenAI-compatible API with 30+ models for hybrid routing
Endpoint statusChat, streaming, JSON mode, function calling, embeddings, RAG and agents are live

TL;DR

  • Local agents run the same plan-act-observe loop as cloud agents, just on your hardware.
  • Pick a model with reliable tool calling; small is fine when the task graph is simple.
  • Sandbox tool execution, because local inference does not make tool calls safe.
  • Use an OpenAI-compatible local server so agent code can move between local and cloud.
  • Route hard reasoning to a cloud API in hybrid mode instead of forcing one machine to do everything.

How it works, step by step

  1. Define the agent's tools and success criteria before choosing a model.
  2. Install a local runtime such as Ollama, LM Studio or llama.cpp and pull a tool-capable model.
  3. Verify function calling with one tool and a scripted prompt before adding complexity.
  4. Add memory: keep conversation state in the prompt and durable facts in a vector store.
  5. Sandbox tool execution with restricted permissions and a command allowlist.
  6. Log every model decision, tool call and observation for debugging and audit.
  7. Add a cloud fallback for tasks the local model fails or when load exceeds capacity.
1Define the agent'stools and successcriteria before2Install a localruntime such asOllama, LM Studio3Verify functioncalling with onetool and a scripted4Add memory: keepconversation statein the prompt and5Sandbox toolexecution withrestricted6Log every modeldecision, tool calland observation for

Try it yourself

Open the AI agent builder →

How a local agent loop actually works

An agent is a control loop around a language model. The model receives a goal plus tool schemas, emits a structured tool call, your code executes it, and the result is appended to the context for the next turn. That repeats until the model returns a final answer or a step budget runs out.

Locally, the difference is only where inference happens. A runtime hosts the weights and serves an HTTP API; your agent code is unchanged. The practical constraints shift to memory, context length and tool-calling reliability, which is why model choice matters more locally than it does against a frontier API.

Choosing a model and runtime for local agents

Tool calling is the first filter. Look for a model family trained on function calling, then confirm your runtime parses its tool-call format. llama.cpp and Ollama handle common formats directly, while vLLM needs a tool-call parser flag enabled for the model family you serve.

  • Ollama is the fastest way to start: one command pulls a model and exposes /v1.
  • llama.cpp server gives control over quantization and offload across CPU, Metal, CUDA and ROCm.
  • vLLM is the choice for concurrent agent traffic because of continuous batching.

For agent work, a 7B-14B instruction model with strong structured output usually beats a larger model that emits malformed JSON, because every parse failure costs an extra loop.

Sandboxing, memory and hybrid fallback

Local inference removes network exposure, not risk. A tool that runs shell commands or writes files has the same power whether the model is local or remote, so restrict the tool surface: allowlist commands, mount read-only data, run tools in a container, and require confirmation for destructive actions. Treat retrieved documents as untrusted input, since prompt injection travels through RAG.

Memory has two layers. Short-term state is the message history in context; long-term memory is a vector store you query and write deliberately. Keep the retrieval step small so the context does not crowd out the task.

Finally, plan the fallback. When a local model stalls on a hard subtask, an OpenAI-compatible cloud API can take over without changing agent code. Plugsky serves chat, streaming, JSON mode, function calling, embeddings, RAG and agents live from one key; image, audio, moderation, batch and fine-tuning endpoints are coming soon, so keep those workloads on their current provider for now. See pricing for plan details.

Honest comparison

ConcernLocal-only agentHybrid local + PlugskyPure cloud agent
Data exposureNothing leaves the machineSensitive steps stay localPrompts leave your network
Model ceilingLimited by your memoryLocal model plus 30+ cloud modelsFrontier models on demand
Cost shapeHardware, power and opsHardware plus one flat cloud planPer-token or flat plan
Offline operationWorks with no internetDegrades to local onlyFails without connectivity
OperationsYou patch, monitor and scaleSplit responsibilitiesProvider manages

Frequently asked questions

Can local models call tools reliably?

Yes, if the model was trained for function calling and the runtime parses its tool-call format. Qwen, Llama 3.x, Mistral and gpt-oss families are common local choices.

How much VRAM does a local agent need?

A 7B-8B model at 4-bit occupies roughly 4-5 GB of weights, plus KV cache and runtime overhead. 8 GB of VRAM is a realistic floor for a simple single-tool agent.

Is a local agent private?

The model and prompt stay on your machine, but tool calls can still reach the network. Restrict egress and sandbox tools if privacy is the point.

Do I need a GPU?

No. llama.cpp runs on CPU and Apple Silicon, but multi-step agents feel much faster on a GPU or unified-memory Mac.

Can I move the same agent to the cloud later?

Yes. If your local server speaks the OpenAI-compatible /v1 surface, moving to Plugsky is a base URL and model-name change.

What is the best local agent runtime?

Ollama for quick starts, llama.cpp for quantization control, and vLLM for concurrent GPU serving. All three expose OpenAI-compatible APIs.

Does Plugsky run agents?

Yes. Agents, tool calling and RAG are live on Plugsky, so the same agent code can route to 30+ hosted models when local capacity is not enough.