Key facts
| Agent components | Model, runtime, loop framework, tool layer and memory store |
| Tool-calling models | Qwen, Llama 3.x, Mistral and gpt-oss families are common local choices |
| Runtimes | Ollama, llama.cpp server and vLLM all expose OpenAI-compatible endpoints |
| Frameworks | LangChain, LlamaIndex, CrewAI, AutoGen or a small custom loop |
| Hardware guidance | A 7B-14B model at 4-bit fits 8-12 GB of VRAM with modest context |
| Cloud fallback | Plugsky serves agents live with 30+ models over an OpenAI-compatible API |
| Free tier | plugsky-micro and plugsky-lite on the free plan, plus a 14-day full-access trial |
TL;DR
- Reliable tool calling beats raw model size for agent work.
- Choose the runtime for concurrency: vLLM for many users, Ollama or llama.cpp for single-user control.
- Keep the model endpoint OpenAI-compatible so frameworks and cloud fallback stay portable.
- Budget memory for context and KV cache, not just weights.
- Test every tool schema against the local model before trusting a multi-step workflow.
How it works, step by step
- Write down the agent's tools, expected inputs and failure modes.
- Shortlist two or three tool-calling models that fit your memory budget.
- Serve them locally through Ollama, llama.cpp or vLLM and confirm the endpoint lists the model.
- Run the same scripted tool-call test against each candidate and score parse reliability.
- Wire the winner into your loop framework and add step and token budgets.
- Add a vector store for long-term memory and keep retrieval chunks small.
- Route failed or over-budget tasks to a hosted OpenAI-compatible API as a fallback.
Original data
Try it yourself
Open the best model for agents selector →
The four parts of a local agent stack
An agent needs a model that can emit structured tool calls, a runtime that turns those calls into a usable protocol, a loop that executes tools and feeds results back, and a memory layer. Weakness in any part caps the whole system, so diagnose before upgrading hardware.
- Model: function-calling training, instruction following, enough context for tool output.
- Runtime: parses the model's tool-call format and supports streaming.
- Loop: step limits, retries, error handling, human approval for risky tools.
- Memory: conversation state in context, durable facts in a vector store.
Most failed local agent projects are loop or schema problems, not model quality problems.
Choosing between local models
Rank candidates on structured output reliability first. A 14B model that emits valid JSON every time is worth more than a 32B model that occasionally wraps tool calls in prose. Test with your own schemas, including one tool with nested parameters and one with free-text arguments.
Context length matters because tool results accumulate. Retain only the observations you need, summarize older turns, and prefer models with grouped-query attention so the KV cache stays manageable at longer contexts. If a task requires frontier reasoning, split the workflow: local models handle routine steps, and a hosted model handles the hard planning call.
Frameworks, orchestration and when to go hybrid
Frameworks save time on plumbing but add abstraction. LangChain and LlamaIndex offer mature tool and retriever integrations; CrewAI and AutoGen model multi-agent collaboration; a custom loop of a few hundred lines is often enough for a single agent with five tools. Pick the smallest framework that covers your orchestration needs, because debugging tool calls through layers of abstraction is slow.
Hybrid routing is the pragmatic endpoint. Keep private data and routine steps local, and send hard reasoning or overflow traffic to a hosted API. Plugsky runs agents, tool calling, RAG, embeddings, streaming and JSON mode live behind an OpenAI-compatible endpoint with 30+ models, so moving a step from local to cloud is a base URL change. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. Check the live pricing page before sizing a plan.
Finally, treat tool execution as a security boundary. Sandbox shells and file access, allowlist commands, and log every call so you can replay a failed run.
Honest comparison
| Stack | Best for | Strengths | Trade-offs |
|---|---|---|---|
| Ollama plus custom loop | Single user, simple tools | Fast setup, easy model pulls | Limited concurrency |
| llama.cpp plus LangChain | Quantization control | CPU and GPU flexibility, portable | More manual tuning |
| vLLM plus LlamaIndex | Multi-user agent service | High throughput batching | GPU-heavy, more ops |
| Hybrid local plus Plugsky | Mixed workloads | 30+ cloud models as fallback | Two environments to monitor |
Frequently asked questions
Do local agents need a big model?
No. A 7B-14B model with strong structured output handles most routine tool-calling tasks. Larger models help mainly with ambiguous planning, which can be routed to a hosted API.
Which local models support tool calling?
Families commonly used for local tool calling include Qwen, Llama 3.x, Mistral and gpt-oss. Confirm runtime support for the specific model's tool-call format before building on it.
How much VRAM should I plan for?
Weights for a 7B model at 4-bit are roughly 4-5 GB. Add KV cache, context and runtime overhead, so 8 GB works for light use and 16-24 GB is comfortable for larger models and longer contexts.
Can local agents browse the web or use APIs?
Yes, through tools you write and expose. The model only requests a call; your code decides what runs and with what permissions.
Do I need a framework?
Only if your orchestration is complex. A custom loop with explicit tool schemas is easier to debug for a single agent.
How do I keep an agent private?
Keep the model local, restrict tool egress, sandbox execution and avoid sending retrieved documents to external services. RAG documents can carry prompt injection, so treat them as untrusted.
What happens when the local agent fails at a task?
Add a fallback route to a hosted OpenAI-compatible API. Plugsky runs agents live with 30+ models, so the same agent code can retry a hard step in the cloud.
Can I run agents fully offline?
Yes, with a local model and local tools. Offline operation limits web search and any tool that depends on external services.