Local AI

How do you build local RAG with Ollama?

Ollama runs both embedding and chat models behind a local API. Pull an embedding model and a chat model, embed your chunks through the embeddings endpoint, store the vectors in a local database, retrieve the closest matches for each question, and pass them to the chat model with instructions to answer only from context.

Key facts

RuntimeOllama serves local models over a REST API and an OpenAI-compatible endpoint
Embedding modelsDedicated embedding models are pulled like chat models
Chat modelsQuantized GGUF models run on CPU, Metal or CUDA
Model managementOne command pulls or updates a model by name
Vector storeBring your own: Chroma, Qdrant, pgvector or FAISS
ChunkingChunk size and overlap drive retrieval quality
Cloud fallbackPlugsky offers OpenAI-compatible chat and embeddings
Endpoint statusEmbeddings and chat are live; batch endpoints are coming soon

TL;DR

  • Ollama covers both retrieval embeddings and answer generation locally.
  • Pull a dedicated embedding model instead of reusing a chat model.
  • Bring your own vector store, because Ollama handles inference only.
  • Keep chunks small enough for precise matches and large enough for meaning.
  • The same OpenAI-compatible client can target a hosted API later.

How it works, step by step

  1. Install Ollama and pull an embedding model plus a chat model.
  2. Parse and chunk documents, attaching source metadata.
  3. Embed chunks through the local embeddings endpoint.
  4. Store vectors and metadata in Chroma, Qdrant, pgvector or FAISS.
  5. Embed each question and retrieve the top-k filtered chunks.
  6. Generate an answer with the chat model using only the retrieved context.
  7. Evaluate grounding and tune chunk size, overlap and top-k.
1Install Ollama andpull an embeddingmodel plus a chat2Parse and chunkdocuments,attaching source3Embed chunksthrough the localembeddings4Store vectors andmetadata in Chroma,Qdrant, pgvector or5Embed each questionand retrieve thetop-k filtered6Generate an answerwith the chat modelusing only the

Try it yourself

Open the RAG chunk size calculator →

Ollama as the inference layer

Ollama is a model runner, not a retrieval system. It downloads and serves models, exposes a local API, and can expose an OpenAI-compatible surface for existing clients. In a RAG stack it plays two roles: an embedding service for ingestion and queries, and a chat service for the final answer.

That separation is useful. You can benchmark embedding models and chat models independently, and you can replace either without touching the other. Keep model names and versions pinned in configuration so results stay reproducible.

From chunks to grounded answers

Ingestion is the unglamorous part that decides quality. Parse documents into clean text, split them on natural boundaries, and store metadata with each chunk so retrieval can be filtered. Embed the chunks and upsert them with stable ids.

  • Query time: embed the question with the same model used for chunks.
  • Retrieve: take a small top-k and apply metadata filters such as source or recency.
  • Generate: pass only retrieved context, and require the model to say when the answer is not present.
  • Cite: return chunk identifiers so users can open the source.

Log retrieval results. If the right chunk never appears, tuning the generator will not help.

Scaling and hybrid routing

A single machine with a 7B-8B model handles low-volume personal and small-team RAG well. Limits appear with concurrency, long documents and high-quality demands, because each of those consumes memory or time on the same hardware.

The hybrid pattern keeps ingestion and the vector store local while routing generation to a hosted endpoint. Plugsky exposes embeddings and chat over an OpenAI-compatible API, so the same client code works with a base URL change; embeddings, RAG and chat are live, and batch and file endpoints are coming soon. See pricing for plans and start free with plugsky-micro and plugsky-lite.

Honest comparison

StageOllama localPlugsky hostedCheck before deciding
DeploymentOne binary, models on diskManaged OpenAI-compatible APIOperations capacity
EmbeddingsLocal embedding modelEmbeddings endpoint, liveDimensions and quality
ChatLocal quantized models30+ models on one APIQuality and context needs
Vector storeYour choice, run locallyYour store stays as isIndex scale
OfflineWorks with no networkRequires connectivityOffline requirement

Frequently asked questions

Does Ollama support embeddings?

Yes. It serves embedding models through its API, including an OpenAI-compatible embeddings route, so the same client code works against local and hosted endpoints.

Which embedding model should I pull?

A compact multilingual embedding model is a good default. Verify its vector dimensions and keep them constant, because changing the model requires re-embedding everything.

How do I choose chunk size?

Start with a few hundred tokens and modest overlap, test retrieval on real questions, and adjust. Retriever quality usually matters more than the chat model.

Can Ollama run a vector database?

No. Ollama is an inference runtime. Pair it with Chroma, Qdrant, pgvector or FAISS for storage and similarity search.

How do I keep answers grounded?

Retrieve a small, filtered set of chunks, instruct the model to answer only from that context, and return source references so users can verify.

What hardware do I need?

A small quantized chat model runs on CPU or Apple Silicon, while a GPU makes generation much faster. Embedding is light and rarely the bottleneck.

Can I move to a hosted API without rewriting?

Yes. Ollama exposes OpenAI-compatible routes, so switching to Plugsky is a base URL and model-name change, with embeddings and chat both live.