Local AI

What are the best local AI agents you can run yourself?

The best local AI agent is a combination, not a single product: a tool-calling model, a runtime that parses tool calls, and a framework that owns the loop. Strong model families for local agents include Qwen, Llama 3.x, Mistral and gpt-oss; common runtimes are Ollama, llama.cpp and vLLM, with LangChain, LlamaIndex, CrewAI or custom loops on top. Match the stack to your hardware, not the other way around.

Key facts

Agent componentsModel, runtime, loop framework, tool layer and memory store
Tool-calling modelsQwen, Llama 3.x, Mistral and gpt-oss families are common local choices
RuntimesOllama, llama.cpp server and vLLM all expose OpenAI-compatible endpoints
FrameworksLangChain, LlamaIndex, CrewAI, AutoGen or a small custom loop
Hardware guidanceA 7B-14B model at 4-bit fits 8-12 GB of VRAM with modest context
Cloud fallbackPlugsky serves agents live with 30+ models over an OpenAI-compatible API
Free tierplugsky-micro and plugsky-lite on the free plan, plus a 14-day full-access trial

TL;DR

  • Reliable tool calling beats raw model size for agent work.
  • Choose the runtime for concurrency: vLLM for many users, Ollama or llama.cpp for single-user control.
  • Keep the model endpoint OpenAI-compatible so frameworks and cloud fallback stay portable.
  • Budget memory for context and KV cache, not just weights.
  • Test every tool schema against the local model before trusting a multi-step workflow.

How it works, step by step

  1. Write down the agent's tools, expected inputs and failure modes.
  2. Shortlist two or three tool-calling models that fit your memory budget.
  3. Serve them locally through Ollama, llama.cpp or vLLM and confirm the endpoint lists the model.
  4. Run the same scripted tool-call test against each candidate and score parse reliability.
  5. Wire the winner into your loop framework and add step and token budgets.
  6. Add a vector store for long-term memory and keep retrieval chunks small.
  7. Route failed or over-budget tasks to a hosted OpenAI-compatible API as a fallback.
1Write down theagent's tools,expected inputs and2Shortlist two orthree tool-callingmodels that fit3Serve them locallythrough Ollama,llama.cpp or vLLM4Run the samescripted tool-calltest against each5Wire the winnerinto your loopframework and add6Add a vector storefor long-termmemory and keep

Original data

Qwen, Llama 3.Tool-calling modelA 7B-14B modelHardware guidancePlugsky servesCloud fallbackplugsky-micro Free tierSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the best model for agents selector →

The four parts of a local agent stack

An agent needs a model that can emit structured tool calls, a runtime that turns those calls into a usable protocol, a loop that executes tools and feeds results back, and a memory layer. Weakness in any part caps the whole system, so diagnose before upgrading hardware.

  • Model: function-calling training, instruction following, enough context for tool output.
  • Runtime: parses the model's tool-call format and supports streaming.
  • Loop: step limits, retries, error handling, human approval for risky tools.
  • Memory: conversation state in context, durable facts in a vector store.

Most failed local agent projects are loop or schema problems, not model quality problems.

Choosing between local models

Rank candidates on structured output reliability first. A 14B model that emits valid JSON every time is worth more than a 32B model that occasionally wraps tool calls in prose. Test with your own schemas, including one tool with nested parameters and one with free-text arguments.

Context length matters because tool results accumulate. Retain only the observations you need, summarize older turns, and prefer models with grouped-query attention so the KV cache stays manageable at longer contexts. If a task requires frontier reasoning, split the workflow: local models handle routine steps, and a hosted model handles the hard planning call.

Frameworks, orchestration and when to go hybrid

Frameworks save time on plumbing but add abstraction. LangChain and LlamaIndex offer mature tool and retriever integrations; CrewAI and AutoGen model multi-agent collaboration; a custom loop of a few hundred lines is often enough for a single agent with five tools. Pick the smallest framework that covers your orchestration needs, because debugging tool calls through layers of abstraction is slow.

Hybrid routing is the pragmatic endpoint. Keep private data and routine steps local, and send hard reasoning or overflow traffic to a hosted API. Plugsky runs agents, tool calling, RAG, embeddings, streaming and JSON mode live behind an OpenAI-compatible endpoint with 30+ models, so moving a step from local to cloud is a base URL change. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. Check the live pricing page before sizing a plan.

Finally, treat tool execution as a security boundary. Sandbox shells and file access, allowlist commands, and log every call so you can replay a failed run.

Honest comparison

StackBest forStrengthsTrade-offs
Ollama plus custom loopSingle user, simple toolsFast setup, easy model pullsLimited concurrency
llama.cpp plus LangChainQuantization controlCPU and GPU flexibility, portableMore manual tuning
vLLM plus LlamaIndexMulti-user agent serviceHigh throughput batchingGPU-heavy, more ops
Hybrid local plus PlugskyMixed workloads30+ cloud models as fallbackTwo environments to monitor

Frequently asked questions

Do local agents need a big model?

No. A 7B-14B model with strong structured output handles most routine tool-calling tasks. Larger models help mainly with ambiguous planning, which can be routed to a hosted API.

Which local models support tool calling?

Families commonly used for local tool calling include Qwen, Llama 3.x, Mistral and gpt-oss. Confirm runtime support for the specific model's tool-call format before building on it.

How much VRAM should I plan for?

Weights for a 7B model at 4-bit are roughly 4-5 GB. Add KV cache, context and runtime overhead, so 8 GB works for light use and 16-24 GB is comfortable for larger models and longer contexts.

Can local agents browse the web or use APIs?

Yes, through tools you write and expose. The model only requests a call; your code decides what runs and with what permissions.

Do I need a framework?

Only if your orchestration is complex. A custom loop with explicit tool schemas is easier to debug for a single agent.

How do I keep an agent private?

Keep the model local, restrict tool egress, sandbox execution and avoid sending retrieved documents to external services. RAG documents can carry prompt injection, so treat them as untrusted.

What happens when the local agent fails at a task?

Add a fallback route to a hosted OpenAI-compatible API. Plugsky runs agents live with 30+ models, so the same agent code can retry a hard step in the cloud.

Can I run agents fully offline?

Yes, with a local model and local tools. Offline operation limits web search and any tool that depends on external services.