Key facts
| Pipeline | Ingest, chunk, embed, index, retrieve, rerank, generate |
| Embedding model | A local embedding model produces vectors without network calls |
| Vector stores | Chroma, Qdrant, pgvector and FAISS are common local options |
| Chunking | Chunk size and overlap strongly affect retrieval quality |
| Reranking | A cross-encoder reranker improves top-k precision at extra latency |
| Memory budget | Embeddings are cheap; long retrieved context inflates the KV cache |
| Cloud fallback | Plugsky serves embeddings and chat over an OpenAI-compatible API |
| Endpoint status | Embeddings, RAG and chat are live; batch and file endpoints coming soon |
TL;DR
- Local RAG keeps documents, questions and answers on your hardware end to end.
- Chunking and evaluation matter more than model size for retrieval quality.
- Use a local embedding model with a vector store that supports metadata filters.
- Rerank the shortlist before sending context to the generator.
- Keep generation swappable so hybrid cloud routing stays possible.
How it works, step by step
- Collect and parse source documents into clean text with metadata.
- Choose a chunk size and overlap, then keep the strategy consistent.
- Embed chunks with a local embedding model and store vectors with metadata.
- Implement top-k retrieval with metadata filters and optional hybrid search.
- Add a reranker to reorder candidates before generation.
- Generate answers with a local chat model and require citations.
- Evaluate retrieval and answers on a fixed question set, then tune.
Try it yourself
Open the local RAG stack generator →
The local RAG pipeline
RAG separates knowledge from the model. Documents are parsed into text, split into chunks and converted into vectors by an embedding model. Those vectors live in a database you can query by similarity. At question time, the query is embedded the same way, the closest chunks are retrieved, and a chat model writes an answer grounded in them.
Every stage can run locally. Embedding models are small enough for CPU or a modest GPU; vector stores such as Chroma, Qdrant, pgvector and FAISS are open components; and a 4-bit chat model handles generation. The result is a private system where no prompt or document leaves the machine.
What breaks local RAG quality
Most RAG failures are retrieval failures, not generation failures. The model answers fluently from the wrong context. Common causes:
- Bad chunks. Splitting mid-argument or at arbitrary character counts destroys meaning. Test size and overlap on real questions.
- No metadata. Dates, sources and access labels let you filter; without them you retrieve outdated or irrelevant chunks.
- Pure vector search. Keyword or hybrid search catches exact names, codes and IDs that embeddings blur.
- No reranking. A cross-encoder reranker reorders candidates and often lifts precision noticeably for a small latency cost.
- No evaluation set. Without fixed questions and known answers, tuning is guesswork.
Build the evaluation set early. Measure whether the correct chunk appears in top-k, then whether the answer is faithful to it.
Hybrid and hosted options
Local RAG is the right default for private or offline corpora, but generation quality and throughput have hardware ceilings. A hybrid split is common: keep parsing, embeddings and the vector store local, and route only the final generation step to an OpenAI-compatible endpoint.
Plugsky exposes embeddings and chat over the same OpenAI-compatible surface, so a local stack can call it without a rewrite. Embeddings and RAG are live; batch and file endpoints are coming soon, so keep large ingestion pipelines on your current provider until then. Vector dimensions must match when you switch embedding models. See pricing for plans and start free with plugsky-micro and plugsky-lite.
Honest comparison
| Stage | Local stack | Hosted Plugsky API | Check before deciding |
|---|---|---|---|
| Ingestion | Your parsers and storage | Runs client-side, unchanged | File formats and volume |
| Embeddings | Local embedding model | Embeddings endpoint, live | Dimensions and vector parity |
| Vector store | Chroma, Qdrant, pgvector, FAISS | Your store stays as is | Index scale and filters |
| Generation | Local chat model | 30+ models on one API | Context length and quality bar |
| Operations | You run every component | Managed inference with SLA | Team capacity |
Frequently asked questions
Do I need a GPU for local RAG?
Embedding is light and often acceptable on CPU; generation is the heavy step. A small 4-bit chat model on a modest GPU or Apple Silicon is enough for low-volume use.
Which local vector database should I use?
Chroma for quick local projects, Qdrant for richer filtering and quantization, pgvector when you already run Postgres, and FAISS when you want a library rather than a service.
How do I pick a chunk size?
Start around a few hundred tokens with modest overlap, then test retrieval on real questions. Larger chunks preserve context but dilute similarity; smaller chunks sharpen matching but can split meaning.
Can I mix local and cloud in one RAG stack?
Yes. Keep embeddings and the vector store local, and route generation to an OpenAI-compatible endpoint when you need more capability. Vector dimensions must match if you change embedding models.
How do I evaluate local RAG?
Build a fixed set of questions with known answers, measure whether the right chunks appear in top-k, then score answer faithfulness and citation accuracy.
Is local RAG private?
Documents, embeddings and prompts stay on your hardware if every component is local. Check that your runtime and any user interface do not send telemetry.
What is the biggest local RAG failure mode?
Retrieval returns plausible but wrong chunks, usually because of chunking or missing metadata filters. Fix retrieval before blaming the generator.