Local AI

How do you build a local AI agent with Ollama?

Ollama runs local models and exposes an OpenAI-compatible endpoint, which makes it a practical base for an agent. Pull a tool-capable model, point your agent loop at the local endpoint, define tool schemas, execute calls outside the model, and feed observations back. Keep execution sandboxed, log every step, and add a hosted fallback for tasks the local model cannot finish.

Key facts

Ollama roleLocal model runner with a daemon and an OpenAI-compatible endpoint
Default endpointOpenAI-compatible routes served on localhost port 11434
Model choiceUse a function-calling model such as a recent Qwen, Llama 3.x, Mistral or gpt-oss release
Tool loopModel emits a call, your code validates and runs it, the result returns as an observation
MemoryConversation state in context; durable facts in a local vector store
Cloud fallbackPlugsky is OpenAI-compatible with 30+ models and live agents
Free tierplugsky-micro and plugsky-lite on the free plan, plus a 14-day full-access trial

TL;DR

  • Ollama gives you a local model plus an OpenAI-compatible endpoint in minutes.
  • Pick a function-calling model; tool support is not universal across local weights.
  • Your code executes tools, not the model — validate everything before running.
  • Keep schemas small and typed; local models follow simple contracts best.
  • Point the same agent at a hosted API when local capacity or quality runs out.

How it works, step by step

  1. Install Ollama and pull a function-calling model that fits your memory.
  2. Confirm the model answers a basic prompt through the local endpoint.
  3. Create an API key placeholder and point your agent framework at the local base URL.
  4. Define two or three tools with typed JSON Schema parameters.
  5. Run the loop: call the model, parse the tool call, execute it, append the observation.
  6. Add a step budget, a retry limit and sandboxing for any tool that touches the system.
  7. Log each turn, then add a hosted fallback route for failed or oversized tasks.
1Install Ollama andpull afunction-calling2Confirm the modelanswers a basicprompt through the3Create an API keyplaceholder andpoint your agent4Define two or threetools with typedJSON Schema5Run the loop: callthe model, parsethe tool call,6Add a step budget,a retry limit andsandboxing for any

Original data

OpenAI-compatiDefault endpointUse a functionModel choicePlugsky is OpeCloud fallbackplugsky-micro Free tierSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the Ollama Modelfile generator →

Set up the model and endpoint

Start by checking what your machine can hold. A 7B-8B model at 4-bit needs roughly 4-5 GB for weights before cache, so 8 GB of VRAM or unified memory is a workable floor. Larger models improve planning but slow every step of the loop, and agents make many calls per task.

Ollama serves an OpenAI-compatible endpoint on localhost, so your agent code can use the same SDK it would use against a hosted provider. Verify the model list and send one chat completion before adding tools. If tool calls come back as plain text, the model or its template does not support the format you expect; switch to a function-calling model rather than trying to parse prose.

Design the tool loop

The loop is straightforward but the details decide reliability. Keep schemas small and explicit, with enums instead of free text where possible. Handle three cases: no tool call, a valid call, and a malformed call. Reject unknown parameters, return structured errors so the model can correct itself, and cap retries.

  • Execute tools in your application, never inside the model.
  • Validate arguments against the schema before running anything.
  • Use timeouts on every tool call, including network ones.
  • Require confirmation for destructive actions such as writes or deletes.

Memory has two layers: the message history that grows in context, and durable facts in a vector store. Keep retrieval chunks small so tool results and retrieved text do not crowd out the task.

Sandboxing, limits and hybrid fallback

An agent with shell or file access has real power. Run tools in a container, allowlist commands and paths, mount data read-only where possible, and keep credentials out of prompts and logs. Prompt injection can arrive through retrieved documents, so treat tool output and RAG content as untrusted input.

Add observability: log the raw model output, parsed call, arguments and result for every step, so failures are replayable. Then decide the escalation policy. Small local models handle routine tool calls well but struggle with multi-step planning and long contexts. Routing those steps to a hosted model keeps the workflow reliable without abandoning local inference. Plugsky runs agents, tool calling, JSON mode, streaming and RAG live across 30+ models behind an OpenAI-compatible endpoint, so escalation is a base URL change. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. Check the live pricing page for tiers.

Honest comparison

SetupModelLoopBest for
Ollama plus custom loop7B-8B tool-calling modelYour code, minimal abstractionLearning and small automations
Ollama plus LangChainFunction-calling modelFramework-managedFaster prototyping
Ollama plus n8n-style automationLocal tool modelVisual workflowBusiness process automation
Ollama plus Plugsky fallbackLocal first, hosted on failureSplit routingProduction reliability

Frequently asked questions

Does Ollama support tool calling?

Ollama supports tool calling with models that were trained for it and exposes an OpenAI-compatible API. Support depends on the model, so test with your schemas before building.

Which model should I use for a local agent?

A recent function-calling model from the Qwen, Llama 3.x, Mistral or gpt-oss families at 7B-14B is a solid start. Larger models help with planning but slow each loop step.

How much memory does an Ollama agent need?

Weights for a 7B model at 4-bit are roughly 4-5 GB, plus KV cache that grows with context. 8 GB works for light use; 16 GB is comfortable.

Can Ollama agents access the internet?

Only through tools you implement. The model requests a call; your code decides whether and how to execute it.

How do I stop the agent from running dangerous commands?

Do not expose raw shell access. Allowlist commands, sandbox execution in a container, validate arguments and require confirmation for destructive operations.

How do I debug a failing local agent?

Log the raw model output next to the parsed call. Most failures are format mismatches or missing arguments, which are visible in that comparison.

Can I move the agent to Plugsky later?

Yes. Because Ollama exposes an OpenAI-compatible endpoint, switching the base URL and model name is usually the only change needed.