Local AI

What is the best local LLM for RAG?

For local RAG, the best model is one that follows instructions tightly, stays grounded in the passages you supply, and fits your context and memory budget. Instruction-tuned 7B-14B models from the Qwen, Llama, Mistral and Gemma families handle many document QA workloads, provided retrieval is solid. At small sizes, retrieval quality and reranking matter more than squeezing out extra parameters.

Key facts

Model priorityInstruction following and grounding matter more than raw parameter count
Common familiesQwen, Llama, Mistral and Gemma instruction-tuned models
Memory guide7B at 4-bit is roughly 4-5 GB of weights; 14B around 8-9 GB
Context budgetChunks plus prompt plus answer must fit the context window; KV cache grows with length
Retrieval stackLocal embeddings plus optional reranker, or a hosted embeddings API
ServingOllama, llama.cpp and vLLM all expose OpenAI-compatible endpoints
Endpoint statusRAG, embeddings, chat and function calling are live on Plugsky

TL;DR

  • Pick an instruction-tuned model that cites sources and refuses to invent facts.
  • Keep retrieved chunks small and relevant; context pollution hurts small models fast.
  • Add a reranker before upgrading to a larger generator.
  • Evaluate grounding on your own documents, not on generic benchmarks.
  • Use a hosted embeddings and generation API when local hardware becomes the bottleneck.

How it works, step by step

  1. Define the questions your RAG system must answer and collect 20-50 with known answers.
  2. Prepare documents with consistent chunk sizes and overlap.
  3. Choose an embedding model that covers your languages, and build the index locally.
  4. Shortlist two instruction-tuned LLMs that fit your memory budget.
  5. Write a strict prompt that requires answers only from retrieved context.
  6. Measure grounding, citation accuracy and refusal behaviour on your evaluation set.
  7. Add a reranker or route hard questions to a hosted model if quality still falls short.
1Define thequestions your RAGsystem must answer2Prepare documentswith consistentchunk sizes and3Choose an embeddingmodel that coversyour languages, and4Shortlist twoinstruction-tunedLLMs that fit your5Write a strictprompt thatrequires answers6Measure grounding,citation accuracyand refusal

Try it yourself

Open the best model for RAG selector →

What RAG asks of a local model

Retrieval-augmented generation changes the model's job. Instead of relying on parametric knowledge, it must read supplied passages, extract the relevant facts and answer within them. That favours instruction following and context discipline over breadth of world knowledge.

The failure mode to watch is confident invention. A model that ignores retrieved context and answers from memory is worse than useless for factual work, because errors look authoritative. Test refusal behaviour explicitly: ask questions your corpus cannot answer and check that the model says so rather than improvising.

Memory, context and chunk strategy

RAG moves pressure from training-time knowledge to context length. Weights for a 7B model at 4-bit are roughly 4-5 GB, but KV cache scales with the number of tokens in play, and long retrieved passages inflate it. Keep prompts tight: fewer, better chunks usually beat stuffing the window.

  • Chunk at 256-512 tokens with modest overlap for prose documents.
  • Preserve headings and metadata so filters can narrow the search space.
  • Retrieve a wide candidate set, rerank, then pass only the top passages.
  • Number passages and require citations so answers are checkable.

If answers improve when you paste gold chunks manually, the generator is fine and retrieval needs work.

Local versus hosted RAG

Local RAG keeps documents, vectors and prompts on your hardware, which is often the reason to build it. The costs are operational: index rebuilds, model upgrades, server capacity and monitoring. A single GPU or Apple Silicon machine handles personal and small-team corpora well.

Hybrid is the common production shape. Embeddings and retrieval can stay local while generation bursts to a hosted API, or the whole pipeline can run on a managed API when volume grows. Plugsky serves RAG, embeddings, chat, streaming and function calling live from an OpenAI-compatible endpoint with 30+ models, and supports VPC, on-prem and air-gapped deployment for enterprise teams. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. Check the live pricing page for plan details.

Honest comparison

OptionGeneratorRetrievalPrivacyScaling
Fully local7B-14B instruction modelLocal embeddings plus rerankerDocuments stay localLimited by one machine
Local retrieval, hosted generationHosted modelsLocal indexOnly passages leaveElastic
Fully hosted with Plugsky30+ hosted modelsManaged embeddings and RAGProvider-processedElastic
Local retrieval, hybrid routingLocal first, cloud fallbackLocal indexPolicy-dependentElastic for hard queries

Frequently asked questions

Do I need a large model for RAG?

No. A well-instructed 7B-14B model often suffices when retrieval is strong. Larger models help with multi-document synthesis and ambiguous questions.

Why does my local RAG model hallucinate?

Usually because retrieved passages are missing, irrelevant or buried in noise. Improve chunking, add a reranker and require the model to cite sources.

How much context should I give the model?

Only what it needs. Two to five focused chunks usually outperform twenty loosely related ones, and smaller prompts keep KV cache memory low.

Which embedding model should I pair with it?

A multilingual model such as BGE-M3 or multilingual-e5 is a good default. Evaluate on your own documents, since retrieval quality is corpus-specific.

Can I run RAG without a GPU?

Retrieval and embeddings run acceptably on CPU. Generation is slower, so CPU-only RAG suits batch or low-frequency use rather than interactive chat.

How do I evaluate local RAG quality?

Collect 20-50 questions with known answers, measure whether answers are grounded in retrieved chunks, and track citation accuracy rather than fluency.

Can I move my local RAG pipeline to a hosted API?

Yes. If you use OpenAI-compatible endpoints for embeddings and generation, switching to Plugsky is a base URL and model-name change.

Does Plugsky support on-prem RAG?

Yes. Plugsky supports VPC, on-prem and air-gapped deployments for enterprise customers, with managed RAG and embeddings in the cloud as well.