Key facts
| Definition | Retrieval-augmented generation grounds answers in your documents |
| Pipeline | Ingest, chunk, embed, index, retrieve, rerank, generate |
| Embeddings | A local embedding model vectorises text on-device |
| Stores | Chroma, Qdrant, pgvector and FAISS are common local choices |
| Retrieval | Hybrid search plus reranking improves precision |
| Evaluation | Fixed question sets measure retrieval and answer quality |
| Hosted split | Plugsky provides OpenAI-compatible embeddings and chat, both live |
| Coming soon | Batch and file endpoints for large-scale ingestion |
TL;DR
- Local RAG grounds answers in your documents without data leaving your environment.
- Chunking and retrieval quality decide most outcomes.
- Embed locally, store locally, generate locally, then measure.
- Rerank and filter to improve precision before generation.
- Keep the generation step swappable for hybrid capacity.
How it works, step by step
- Define the corpus, its access rules and the questions users will ask.
- Parse documents cleanly and chunk them with metadata.
- Embed chunks locally and store vectors in a local database.
- Implement top-k retrieval with metadata filters and hybrid search.
- Add reranking and generate answers with citations.
- Evaluate retrieval and answers on a fixed question set.
- Tune chunking, filters and top-k before changing the model.
Try it yourself
Open the RAG architecture builder →
RAG in plain terms
A language model knows what it was trained on and nothing about your private documents. RAG closes that gap by retrieving relevant passages at question time and placing them in the prompt, so the answer is grounded in your content. The model writes the answer; the retrieval system decides what it is allowed to see.
That division explains where the failures come from. If retrieval returns the wrong passage, the model produces a confident wrong answer. Which is why retrieval quality, not model size, is the first thing to measure.
The local pipeline, stage by stage
Ingestion parses documents into clean text; OCR is needed for scanned pages. Chunking splits text into retrievable units with metadata such as source, page and date. An embedding model converts each chunk into a vector, and a local store holds them.
- Retrieve: embed the question, fetch top-k nearest chunks, apply metadata filters.
- Improve precision: combine keyword and vector search, then rerank the shortlist with a cross-encoder.
- Generate: pass only retrieved context, require citations and allow the model to say it does not know.
- Evaluate: track whether the correct chunk was retrieved before scoring the answer itself.
Every stage runs locally with open components, which is what makes the privacy claim real.
Evaluation and hybrid scaling
Build a fixed evaluation set early: real questions with known answers, plus questions whose answers are absent to test refusal behaviour. Score retrieval and generation separately, because they fail differently and need different fixes.
When local generation becomes the bottleneck, split the stack. Keep parsing, embeddings and the vector store local, and route generation to a hosted endpoint that speaks the OpenAI-compatible shape. Plugsky provides embeddings and chat as live endpoints with 30+ models; batch and file endpoints are coming soon, so keep bulk ingestion local for now. See pricing for plans and start free with plugsky-micro and plugsky-lite.
Honest comparison
| Stage | Local RAG | Plugsky hosted split | Check before deciding |
|---|---|---|---|
| Ingestion | Your parsers and pipeline | Batch and file endpoints coming soon | Volume and formats |
| Embeddings | Local embedding model | Embeddings endpoint, live | Dimensions and languages |
| Store | Chroma, Qdrant, pgvector or FAISS | Your store stays local | Corpus size and filters |
| Generation | Local chat model | 30+ models on one API | Quality and context |
| Operations | You run every component | Managed inference with SLA | Team capacity |
Frequently asked questions
What is RAG in one sentence?
RAG retrieves relevant passages from your documents and gives them to a language model so the answer is grounded in your content rather than only in training data.
Do I need a GPU for local RAG?
Embedding runs acceptably on CPU for modest corpora; generation is the heavy step. A small quantized model on a GPU or Apple Silicon covers low-volume use.
What causes wrong answers?
Usually retrieval: the correct passage was not in the top results because of chunking, missing metadata filters or weak similarity matching. Fix retrieval before changing the model.
Should I use hybrid search?
If your questions include names, codes or exact terms, yes. Combining keyword and vector search catches matches that embeddings blur.
How do I evaluate RAG?
Create fixed questions with known answers, measure whether the right chunks are retrieved, then score answer faithfulness and citation accuracy separately.
Can I keep the store local and use a hosted model?
Yes. Keep parsing, embeddings and vectors local and route generation to an OpenAI-compatible endpoint. Re-embed only if you change the embedding model.
What is the most common mistake?
Treating RAG as a prompting problem. Most failures are ingestion and retrieval problems that no prompt can fix.