Local AI

What is the best local embedding model for RAG?

There is no single best local embedding model, but strong general choices are multilingual families such as BGE-M3 and multilingual-e5, alongside Nomic Embed, GTE and Qwen embedding models. Pick by language coverage, dimension trade-offs, context length and licence. A local embedding model plus a local reranker often beats upgrading the generator, because retrieval quality sets the ceiling for every answer.

Key facts

Common local modelsBGE-M3, multilingual-e5, Nomic Embed, GTE and Qwen embedding families
Typical dimensions384, 768, 1024 and higher; larger is not automatically better
Context lengthModels commonly embed 512 to 8192 tokens per chunk, depending on family
MultilingualBGE-M3 and multilingual-e5 cover many languages in one index
Storage mathDimensionality times vector count times bytes per value sets index size
RerankingA cross-encoder reranker improves precision on top-k retrieval
Cloud optionPlugsky serves embeddings live over an OpenAI-compatible endpoint

TL;DR

  • Choose by language coverage and context length before chasing leaderboard scores.
  • Run one blind evaluation on your own documents; generic rankings rarely predict your results.
  • Higher dimensions improve recall but increase storage and query cost.
  • Add a reranker before replacing the embedding model.
  • Keep embeddings local for privacy; move them to a hosted API when volume outgrows one machine.

How it works, step by step

  1. List the languages and document types your index must cover.
  2. Shortlist two or three embedding models, including one multilingual candidate.
  3. Chunk a sample of real documents with consistent size and overlap settings.
  4. Build a small labelled query set of 20-50 questions with known answers.
  5. Measure recall at k and answer grounding for each model on that set.
  6. Add a reranker and re-measure to see whether retrieval or generation is the bottleneck.
  7. Estimate index size and query latency, then choose the model you can sustain in production.
1List the languagesand document typesyour index must2Shortlist two orthree embeddingmodels, including3Chunk a sample ofreal documents withconsistent size and4Build a smalllabelled query setof 20-50 questions5Measure recall at kand answergrounding for each6Add a reranker andre-measure to seewhether retrieval

Original data

BGE-M3, multilCommon local model384, 768, 1024Typical dimensionsModels commonlContext lengthBGE-M3 and mulMultilingualSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the embedding model comparison →

How to choose a local embedding model

Start with language coverage. If your documents mix languages, a multilingual model such as BGE-M3 or multilingual-e5 keeps one index instead of several. If everything is English and domain-specific, a smaller English-focused model may match quality at lower cost.

Then weigh dimension and context. Dimensions determine index size and query latency: 384 and 768 are cheap, 1024 and above improve recall on large corpora but consume more memory and disk. Context length decides how much text one vector represents, which shapes your chunking strategy. A 512-token limit forces smaller chunks; an 8k limit lets a vector represent a whole section, at some loss of precision.

Finally, check the licence and how the model is distributed. Local embedding weights are usually downloadable, but commercial terms differ between families and versions.

Retrieval quality comes from the pipeline

The embedding model is one variable among several. Chunk size, overlap, metadata filters and hybrid search often move metrics more than a model swap. Test combinations rather than single changes, and keep a fixed evaluation set so results compare.

  • Chunking: 256-512 tokens with 10-20% overlap is a solid default for prose.
  • Hybrid search: combine vector similarity with keyword scoring for names, codes and rare terms.
  • Reranking: retrieve a wide candidate set, then rerank to the passages you actually pass to the model.
  • Grounding: require citations and verify that answers quote real chunks.

If answers improve when you feed gold passages directly, retrieval is the problem. If they still fail, the generator or prompt is the bottleneck.

Local versus hosted embeddings

Local embeddings keep document text and vectors on your hardware, which matters for regulated content. The trade-off is operational: you own the model version, serving process and scaling. A single GPU or even a CPU can serve embeddings for moderate corpora because embedding calls are short and parallelisable, but index rebuilds on large corpora take real time.

Hosted embeddings remove that operations work and scale on demand. Plugsky serves embeddings through an OpenAI-compatible API, alongside chat, streaming, JSON mode, function calling, RAG and agents, all live with 30+ models. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. A practical pattern is local embeddings for private corpora and hosted embeddings for public or high-volume collections. Plan limits are on the live pricing page.

Honest comparison

OptionLanguage fitRuns onOperationsPrivacy
BGE-M3MultilingualCPU or GPUYou serve and version itFully local
Multilingual-e5MultilingualCPU or GPUYou serve and version itFully local
Nomic Embed / GTEMostly English, some multilingualCPU or GPUYou serve and version itFully local
Plugsky embeddingsMultilingual via hosted APIHostedManagedProvider-processed

Frequently asked questions

Do I need a GPU to run local embeddings?

No. Embedding models are small compared with chat models and run acceptably on CPU for moderate volumes. A GPU helps during large index builds and high query rates.

Which embedding model is best for Arabic or Indonesian?

Prefer a multilingual family such as BGE-M3 or multilingual-e5, then evaluate on your own documents in those languages before committing.

How many dimensions should I use?

Start at 768 or 1024 for general RAG. Higher dimensions help large corpora but raise index size and latency; test whether the recall gain justifies the cost.

Should I use the same model for indexing and querying?

Yes. Always embed queries with the exact model and version used at index time, or similarity scores become meaningless.

Do I need a reranker if my embeddings are good?

A reranker usually still helps. Retrieve a generous candidate set, then rerank to the few passages the generator sees.

Can I mix local and hosted embeddings?

You can, but keep separate indexes per model. Vectors from different models are not comparable, so never mix them in one collection.

How do I evaluate a local embedding model?

Build 20-50 questions with known answers, measure recall at k on your real documents, and compare rivals on the same chunks and queries.