Key facts
| Managed models | plugsky-embed-v1 (1536 dimensions) and plugsky-embed-large (3072 dimensions) |
| Multilingual | plugsky-embed-multilingual is listed in the model catalogue |
| Endpoint | POST /v1/embeddings on the OpenAI-compatible API |
| Batch limits | Up to 2,048 inputs per request, max 8,191 tokens each |
| Retrieval fit | Vectors plug into keyword, vector and hybrid search in RAG collections |
| Data handling | API data is not used to train models |
| Pricing model | Flat monthly self-serve plans; no per-token billing on self-serve |
| Product status | Live |
TL;DR
- The same model must embed both documents and queries; never mix models inside one index.
- Dimensions drive storage and latency, not quality by themselves - measure recall on your data.
- Multilingual corpora need a multilingual model, not translation before embedding.
- Hosted APIs cut ops work; open weights give control at the cost of GPU and serving effort.
- Plugsky offers managed embeddings with OpenAI-ada-compatible vectors for easy migration.
How it works, step by step
- Define the language mix and document types your corpus actually contains.
- Shortlist two or three embedding models, including one managed option.
- Build a golden question set with known correct chunks.
- Embed and index the same sample with each model and measure recall@k and MRR.
- Check storage, latency and re-embedding cost at your real corpus size.
- Pick one model, version it in your config, and plan a re-embedding path for upgrades.
- Monitor retrieval quality after launch so model or corpus drift is caught early.
Original data
Try it yourself
Open the embedding model comparison →
What an embedding model actually decides
An embedding model maps text to a vector so that similar meanings land close together in vector space. In a RAG pipeline it decides what your retriever can find before any reranker or generator runs. If the embedding does not separate the relevant passage from the noise, no prompt engineering downstream can recover it.
Four properties matter in practice: training data and language coverage, output dimensions, the maximum input length the model accepts, and stability of the model version. A model that handles your domain vocabulary and languages will beat a larger model that does not. Changing the model later means re-embedding the whole corpus, so treat the embedding choice as a schema decision rather than a quick experiment.
How to compare embedding models properly
Public leaderboards are a starting point, not a decision. Build a small golden set from your own documents: 50 to 200 questions with the chunk IDs that should be retrieved. Then measure recall@k for the first-stage retrieval and mean reciprocal rank to see how high the right chunk appears. Run the same set for each candidate model and index.
Watch the practical dimensions of the comparison too. Higher-dimensional vectors cost more memory and storage and can slow searches; longer context windows help long documents but increase cost. A 768- or 1024-dimension model that fits your domain often outperforms a bigger model on real queries, and it is cheaper to run.
Hosted APIs versus open-weight models
Open-weight families such as BGE, E5, Nomic and Jina can be self-hosted, which keeps data inside your perimeter and removes per-call API dependency, but you own GPUs, serving infrastructure, upgrades and evaluation. Hosted embedding APIs remove that work and usually provide predictable latency and simple scaling.
Plugsky takes the hosted path with POST /v1/embeddings and states that API data is not used to train models. The endpoint also works standalone: if you already run a vector database such as pgvector, Qdrant or Pinecone, you can keep it and only swap the embedding call.
Choosing on Plugsky
plugsky-embed-v1 returns 1536-dimension vectors and is compatible with common OpenAI-ada-style pipelines, which keeps migrations simple. plugsky-embed-large returns 3072 dimensions when you want a richer representation and can accept the extra storage. For non-English corpora, plugsky-embed-multilingual is the model to test first.
Compare candidates with the embedding model comparison, then validate recall on your own set. Start on the free plan with plugsky-micro and plugsky-lite, or use the 14-day full-access trial; learning and evaluation prompts are cheap to run while you decide. Current plans are listed on the live pricing page.
Honest comparison
| Factor | plugsky-embed-v1 | plugsky-embed-large | Self-hosted open weights |
|---|---|---|---|
| Dimensions | 1536 | 3072 | Varies by model |
| Interface | POST /v1/embeddings, OpenAI-compatible | Same endpoint | You run the model server |
| Multilingual | Dedicated multilingual model available | Dedicated multilingual model available | Model-dependent |
| Operations | Managed, no GPU to run | Managed, no GPU to run | GPU, serving, upgrades |
| Migration | OpenAI-ada-compatible vectors | Higher-dimension vectors | Full re-embed required |
| Data path | API data not used for training | API data not used for training | You control the host |
Frequently asked questions
How many embedding dimensions do I need?
There is no fixed number. Higher dimensions store more information but cost more memory and often search slower; measure recall on your own data and pick the smallest model that clears your quality bar.
Can I mix two embedding models in one index?
No. Documents and queries must be embedded with the same model, and mixing vectors from different models in one index produces meaningless distances.
Does Plugsky have a multilingual embedding model?
Yes, plugsky-embed-multilingual is part of the catalogue, alongside plugsky-embed-v1 and plugsky-embed-large.
Can I use Plugsky embeddings with my existing vector database?
Yes. POST /v1/embeddings works standalone, so you can generate vectors with Plugsky and store them in your own vector store.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
Does Plugsky train on my data?
Plugsky states that API data is not used to train models, and collections are encrypted at rest.