Key facts
| Ollama role | Local model runner with a daemon and an OpenAI-compatible endpoint |
| Default endpoint | OpenAI-compatible routes served on localhost port 11434 |
| Model choice | Use a function-calling model such as a recent Qwen, Llama 3.x, Mistral or gpt-oss release |
| Tool loop | Model emits a call, your code validates and runs it, the result returns as an observation |
| Memory | Conversation state in context; durable facts in a local vector store |
| Cloud fallback | Plugsky is OpenAI-compatible with 30+ models and live agents |
| Free tier | plugsky-micro and plugsky-lite on the free plan, plus a 14-day full-access trial |
TL;DR
- Ollama gives you a local model plus an OpenAI-compatible endpoint in minutes.
- Pick a function-calling model; tool support is not universal across local weights.
- Your code executes tools, not the model — validate everything before running.
- Keep schemas small and typed; local models follow simple contracts best.
- Point the same agent at a hosted API when local capacity or quality runs out.
How it works, step by step
- Install Ollama and pull a function-calling model that fits your memory.
- Confirm the model answers a basic prompt through the local endpoint.
- Create an API key placeholder and point your agent framework at the local base URL.
- Define two or three tools with typed JSON Schema parameters.
- Run the loop: call the model, parse the tool call, execute it, append the observation.
- Add a step budget, a retry limit and sandboxing for any tool that touches the system.
- Log each turn, then add a hosted fallback route for failed or oversized tasks.
Original data
Try it yourself
Open the Ollama Modelfile generator →
Set up the model and endpoint
Start by checking what your machine can hold. A 7B-8B model at 4-bit needs roughly 4-5 GB for weights before cache, so 8 GB of VRAM or unified memory is a workable floor. Larger models improve planning but slow every step of the loop, and agents make many calls per task.
Ollama serves an OpenAI-compatible endpoint on localhost, so your agent code can use the same SDK it would use against a hosted provider. Verify the model list and send one chat completion before adding tools. If tool calls come back as plain text, the model or its template does not support the format you expect; switch to a function-calling model rather than trying to parse prose.
Design the tool loop
The loop is straightforward but the details decide reliability. Keep schemas small and explicit, with enums instead of free text where possible. Handle three cases: no tool call, a valid call, and a malformed call. Reject unknown parameters, return structured errors so the model can correct itself, and cap retries.
- Execute tools in your application, never inside the model.
- Validate arguments against the schema before running anything.
- Use timeouts on every tool call, including network ones.
- Require confirmation for destructive actions such as writes or deletes.
Memory has two layers: the message history that grows in context, and durable facts in a vector store. Keep retrieval chunks small so tool results and retrieved text do not crowd out the task.
Sandboxing, limits and hybrid fallback
An agent with shell or file access has real power. Run tools in a container, allowlist commands and paths, mount data read-only where possible, and keep credentials out of prompts and logs. Prompt injection can arrive through retrieved documents, so treat tool output and RAG content as untrusted input.
Add observability: log the raw model output, parsed call, arguments and result for every step, so failures are replayable. Then decide the escalation policy. Small local models handle routine tool calls well but struggle with multi-step planning and long contexts. Routing those steps to a hosted model keeps the workflow reliable without abandoning local inference. Plugsky runs agents, tool calling, JSON mode, streaming and RAG live across 30+ models behind an OpenAI-compatible endpoint, so escalation is a base URL change. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. Check the live pricing page for tiers.
Honest comparison
| Setup | Model | Loop | Best for |
|---|---|---|---|
| Ollama plus custom loop | 7B-8B tool-calling model | Your code, minimal abstraction | Learning and small automations |
| Ollama plus LangChain | Function-calling model | Framework-managed | Faster prototyping |
| Ollama plus n8n-style automation | Local tool model | Visual workflow | Business process automation |
| Ollama plus Plugsky fallback | Local first, hosted on failure | Split routing | Production reliability |
Frequently asked questions
Does Ollama support tool calling?
Ollama supports tool calling with models that were trained for it and exposes an OpenAI-compatible API. Support depends on the model, so test with your schemas before building.
Which model should I use for a local agent?
A recent function-calling model from the Qwen, Llama 3.x, Mistral or gpt-oss families at 7B-14B is a solid start. Larger models help with planning but slow each loop step.
How much memory does an Ollama agent need?
Weights for a 7B model at 4-bit are roughly 4-5 GB, plus KV cache that grows with context. 8 GB works for light use; 16 GB is comfortable.
Can Ollama agents access the internet?
Only through tools you implement. The model requests a call; your code decides whether and how to execute it.
How do I stop the agent from running dangerous commands?
Do not expose raw shell access. Allowlist commands, sandbox execution in a container, validate arguments and require confirmation for destructive operations.
How do I debug a failing local agent?
Log the raw model output next to the parsed call. Most failures are format mismatches or missing arguments, which are visible in that comparison.
Can I move the agent to Plugsky later?
Yes. Because Ollama exposes an OpenAI-compatible endpoint, switching the base URL and model name is usually the only change needed.