RAG

What are the best embedding models for RAG?

There is no single best embedding model; the right choice depends on language coverage, dimensions, hosting and budget. For many teams the practical answer is a managed model with stable dimensions and strong multilingual support, such as plugsky-embed-v1 or plugsky-embed-large, tested against your own retrieval set before you commit.

Key facts

Managed modelsplugsky-embed-v1 (1536 dimensions) and plugsky-embed-large (3072 dimensions)
Multilingualplugsky-embed-multilingual is listed in the model catalogue
EndpointPOST /v1/embeddings on the OpenAI-compatible API
Batch limitsUp to 2,048 inputs per request, max 8,191 tokens each
Retrieval fitVectors plug into keyword, vector and hybrid search in RAG collections
Data handlingAPI data is not used to train models
Pricing modelFlat monthly self-serve plans; no per-token billing on self-serve
Product statusLive

TL;DR

  • The same model must embed both documents and queries; never mix models inside one index.
  • Dimensions drive storage and latency, not quality by themselves - measure recall on your data.
  • Multilingual corpora need a multilingual model, not translation before embedding.
  • Hosted APIs cut ops work; open weights give control at the cost of GPU and serving effort.
  • Plugsky offers managed embeddings with OpenAI-ada-compatible vectors for easy migration.

How it works, step by step

  1. Define the language mix and document types your corpus actually contains.
  2. Shortlist two or three embedding models, including one managed option.
  3. Build a golden question set with known correct chunks.
  4. Embed and index the same sample with each model and measure recall@k and MRR.
  5. Check storage, latency and re-embedding cost at your real corpus size.
  6. Pick one model, version it in your config, and plan a re-embedding path for upgrades.
  7. Monitor retrieval quality after launch so model or corpus drift is caught early.
1Define the languagemix and documenttypes your corpus2Shortlist two orthree embeddingmodels, including3Build a goldenquestion set withknown correct4Embed and index thesame sample witheach model and5Check storage,latency andre-embedding cost6Pick one model,version it in yourconfig, and plan a

Original data

plugsky-embed-Managed modelsPOST /v1/embedEndpointUp to 2,048 inBatch limitsSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the embedding model comparison →

What an embedding model actually decides

An embedding model maps text to a vector so that similar meanings land close together in vector space. In a RAG pipeline it decides what your retriever can find before any reranker or generator runs. If the embedding does not separate the relevant passage from the noise, no prompt engineering downstream can recover it.

Four properties matter in practice: training data and language coverage, output dimensions, the maximum input length the model accepts, and stability of the model version. A model that handles your domain vocabulary and languages will beat a larger model that does not. Changing the model later means re-embedding the whole corpus, so treat the embedding choice as a schema decision rather than a quick experiment.

How to compare embedding models properly

Public leaderboards are a starting point, not a decision. Build a small golden set from your own documents: 50 to 200 questions with the chunk IDs that should be retrieved. Then measure recall@k for the first-stage retrieval and mean reciprocal rank to see how high the right chunk appears. Run the same set for each candidate model and index.

Watch the practical dimensions of the comparison too. Higher-dimensional vectors cost more memory and storage and can slow searches; longer context windows help long documents but increase cost. A 768- or 1024-dimension model that fits your domain often outperforms a bigger model on real queries, and it is cheaper to run.

Hosted APIs versus open-weight models

Open-weight families such as BGE, E5, Nomic and Jina can be self-hosted, which keeps data inside your perimeter and removes per-call API dependency, but you own GPUs, serving infrastructure, upgrades and evaluation. Hosted embedding APIs remove that work and usually provide predictable latency and simple scaling.

Plugsky takes the hosted path with POST /v1/embeddings and states that API data is not used to train models. The endpoint also works standalone: if you already run a vector database such as pgvector, Qdrant or Pinecone, you can keep it and only swap the embedding call.

Choosing on Plugsky

plugsky-embed-v1 returns 1536-dimension vectors and is compatible with common OpenAI-ada-style pipelines, which keeps migrations simple. plugsky-embed-large returns 3072 dimensions when you want a richer representation and can accept the extra storage. For non-English corpora, plugsky-embed-multilingual is the model to test first.

Compare candidates with the embedding model comparison, then validate recall on your own set. Start on the free plan with plugsky-micro and plugsky-lite, or use the 14-day full-access trial; learning and evaluation prompts are cheap to run while you decide. Current plans are listed on the live pricing page.

Honest comparison

Factorplugsky-embed-v1plugsky-embed-largeSelf-hosted open weights
Dimensions15363072Varies by model
InterfacePOST /v1/embeddings, OpenAI-compatibleSame endpointYou run the model server
MultilingualDedicated multilingual model availableDedicated multilingual model availableModel-dependent
OperationsManaged, no GPU to runManaged, no GPU to runGPU, serving, upgrades
MigrationOpenAI-ada-compatible vectorsHigher-dimension vectorsFull re-embed required
Data pathAPI data not used for trainingAPI data not used for trainingYou control the host

Frequently asked questions

How many embedding dimensions do I need?

There is no fixed number. Higher dimensions store more information but cost more memory and often search slower; measure recall on your own data and pick the smallest model that clears your quality bar.

Can I mix two embedding models in one index?

No. Documents and queries must be embedded with the same model, and mixing vectors from different models in one index produces meaningless distances.

Does Plugsky have a multilingual embedding model?

Yes, plugsky-embed-multilingual is part of the catalogue, alongside plugsky-embed-v1 and plugsky-embed-large.

Can I use Plugsky embeddings with my existing vector database?

Yes. POST /v1/embeddings works standalone, so you can generate vectors with Plugsky and store them in your own vector store.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

Does Plugsky train on my data?

Plugsky states that API data is not used to train models, and collections are encrypted at rest.