Key facts
| Model priority | Instruction following and grounding matter more than raw parameter count |
| Common families | Qwen, Llama, Mistral and Gemma instruction-tuned models |
| Memory guide | 7B at 4-bit is roughly 4-5 GB of weights; 14B around 8-9 GB |
| Context budget | Chunks plus prompt plus answer must fit the context window; KV cache grows with length |
| Retrieval stack | Local embeddings plus optional reranker, or a hosted embeddings API |
| Serving | Ollama, llama.cpp and vLLM all expose OpenAI-compatible endpoints |
| Endpoint status | RAG, embeddings, chat and function calling are live on Plugsky |
TL;DR
- Pick an instruction-tuned model that cites sources and refuses to invent facts.
- Keep retrieved chunks small and relevant; context pollution hurts small models fast.
- Add a reranker before upgrading to a larger generator.
- Evaluate grounding on your own documents, not on generic benchmarks.
- Use a hosted embeddings and generation API when local hardware becomes the bottleneck.
How it works, step by step
- Define the questions your RAG system must answer and collect 20-50 with known answers.
- Prepare documents with consistent chunk sizes and overlap.
- Choose an embedding model that covers your languages, and build the index locally.
- Shortlist two instruction-tuned LLMs that fit your memory budget.
- Write a strict prompt that requires answers only from retrieved context.
- Measure grounding, citation accuracy and refusal behaviour on your evaluation set.
- Add a reranker or route hard questions to a hosted model if quality still falls short.
Try it yourself
Open the best model for RAG selector →
What RAG asks of a local model
Retrieval-augmented generation changes the model's job. Instead of relying on parametric knowledge, it must read supplied passages, extract the relevant facts and answer within them. That favours instruction following and context discipline over breadth of world knowledge.
The failure mode to watch is confident invention. A model that ignores retrieved context and answers from memory is worse than useless for factual work, because errors look authoritative. Test refusal behaviour explicitly: ask questions your corpus cannot answer and check that the model says so rather than improvising.
Memory, context and chunk strategy
RAG moves pressure from training-time knowledge to context length. Weights for a 7B model at 4-bit are roughly 4-5 GB, but KV cache scales with the number of tokens in play, and long retrieved passages inflate it. Keep prompts tight: fewer, better chunks usually beat stuffing the window.
- Chunk at 256-512 tokens with modest overlap for prose documents.
- Preserve headings and metadata so filters can narrow the search space.
- Retrieve a wide candidate set, rerank, then pass only the top passages.
- Number passages and require citations so answers are checkable.
If answers improve when you paste gold chunks manually, the generator is fine and retrieval needs work.
Local versus hosted RAG
Local RAG keeps documents, vectors and prompts on your hardware, which is often the reason to build it. The costs are operational: index rebuilds, model upgrades, server capacity and monitoring. A single GPU or Apple Silicon machine handles personal and small-team corpora well.
Hybrid is the common production shape. Embeddings and retrieval can stay local while generation bursts to a hosted API, or the whole pipeline can run on a managed API when volume grows. Plugsky serves RAG, embeddings, chat, streaming and function calling live from an OpenAI-compatible endpoint with 30+ models, and supports VPC, on-prem and air-gapped deployment for enterprise teams. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. Check the live pricing page for plan details.
Honest comparison
| Option | Generator | Retrieval | Privacy | Scaling |
|---|---|---|---|---|
| Fully local | 7B-14B instruction model | Local embeddings plus reranker | Documents stay local | Limited by one machine |
| Local retrieval, hosted generation | Hosted models | Local index | Only passages leave | Elastic |
| Fully hosted with Plugsky | 30+ hosted models | Managed embeddings and RAG | Provider-processed | Elastic |
| Local retrieval, hybrid routing | Local first, cloud fallback | Local index | Policy-dependent | Elastic for hard queries |
Frequently asked questions
Do I need a large model for RAG?
No. A well-instructed 7B-14B model often suffices when retrieval is strong. Larger models help with multi-document synthesis and ambiguous questions.
Why does my local RAG model hallucinate?
Usually because retrieved passages are missing, irrelevant or buried in noise. Improve chunking, add a reranker and require the model to cite sources.
How much context should I give the model?
Only what it needs. Two to five focused chunks usually outperform twenty loosely related ones, and smaller prompts keep KV cache memory low.
Which embedding model should I pair with it?
A multilingual model such as BGE-M3 or multilingual-e5 is a good default. Evaluate on your own documents, since retrieval quality is corpus-specific.
Can I run RAG without a GPU?
Retrieval and embeddings run acceptably on CPU. Generation is slower, so CPU-only RAG suits batch or low-frequency use rather than interactive chat.
How do I evaluate local RAG quality?
Collect 20-50 questions with known answers, measure whether answers are grounded in retrieved chunks, and track citation accuracy rather than fluency.
Can I move my local RAG pipeline to a hosted API?
Yes. If you use OpenAI-compatible endpoints for embeddings and generation, switching to Plugsky is a base URL and model-name change.
Does Plugsky support on-prem RAG?
Yes. Plugsky supports VPC, on-prem and air-gapped deployments for enterprise customers, with managed RAG and embeddings in the cloud as well.