Key facts
| Endpoint | POST /v1/embeddings, OpenAI-compatible request shape |
| Models | plugsky-embed-v1 (1536d), plugsky-embed-large (3072d) and plugsky-embed-multilingual |
| Batch limits | Up to 2,048 inputs per request, max 8,191 tokens each |
| Compatibility | plugsky-embed-v1 returns 1536-dimension vectors for ada-style pipelines |
| RAG usage | Vectors feed keyword, vector and hybrid retrieval in RAG collections |
| Data handling | API data is not used to train models |
| Migration | Change base URL and model name; re-embed to fill a new index |
| Product status | Live |
TL;DR
- The API call is easy to swap; the re-embedding of your corpus is the real migration.
- Keep a dimension-compatible model first so application and index code change minimally.
- Build the new index in parallel and dual-query both until quality is proven.
- Never mix vectors from two embedding models in one index.
- Plugsky embeddings work standalone with your existing vector database.
How it works, step by step
- Inventory every place you call the embeddings endpoint and how dimensions are stored.
- Create a parallel index for the new model; leave the old one serving traffic.
- Re-embed a representative slice and compare recall on a golden question set.
- Re-embed the full corpus in batches, respecting the 2,048-input request limit.
- Dual-query old and new indexes and compare retrieved chunks per question.
- Cut retrieval traffic to the new index once quality is equal or better.
- Keep the old index for a rollback window, then remove it and update your docs.
Original data
Try it yourself
Open the embedding API tester →
What actually changes when you switch embeddings
If the provider is OpenAI-compatible, the HTTP call is nearly identical: same endpoint path, same input list, same response shape. Your client library keeps working after a base URL and model name change. What changes is the vector space: dimensions may differ, distances are no longer comparable with the old model, and every stored vector becomes stale for the new model.
That is why an embeddings migration is mostly a data migration. The safest pattern is a parallel index: keep the current index serving production while you re-embed the corpus into a new one, then compare retrieval quality before moving traffic.
Choosing the replacement model
Two paths work. A dimension-compatible model such as plugsky-embed-v1 at 1536 dimensions minimises code changes for pipelines built on common ada-style embeddings, because index schemas and downstream assumptions stay valid. A higher-dimension model such as plugsky-embed-large at 3072 dimensions can represent more nuance but requires a new index and more storage.
For non-English corpora, test the multilingual model before anything else; translation before embedding adds latency, cost and failure modes. Whatever you choose, validate with a golden question set rather than a public leaderboard, and check the maximum input length so long chunks are not silently truncated.
Running the re-embedding safely
Batch the work, track progress per document, and make the job idempotent so a failure can resume instead of restarting. Plugsky accepts up to 2,048 inputs per embeddings request with a maximum of 8,191 tokens per input, so chunk sizes and batch sizes should be planned together. Record the model name and version alongside the index metadata; a future migration will need to know which model produced which vectors.
While both indexes exist, dual-query them for the same user questions and compare the chunks returned and the answers produced. This catches regressions that a spot check misses, and it gives you a defensible cutover decision.
What you gain on Plugsky
Plugsky keeps embeddings on the same OpenAI-compatible API as chat and RAG, so one key and one base URL cover the stack. Vectors return from POST /v1/embeddings, collections and queries handle retrieval when you want managed RAG, and API data is not used to train models. Self-serve plans are flat monthly with no per-token billing, which removes one source of migration uncertainty.
Try requests with the embedding API tester, then start free with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Factor | plugsky-embed-v1 | plugsky-embed-large | Staying on OpenAI embeddings |
|---|---|---|---|
| Dimensions | 1536 | 3072 | Model-dependent |
| Request shape | OpenAI-compatible | OpenAI-compatible | Native |
| Migration effort | Base URL plus model name | New index for higher dimensions | None now, vendor dependency remains |
| Multilingual | Dedicated multilingual model available | Dedicated multilingual model available | Model-dependent |
| Data handling | API data not used for training | Same | Provider terms apply |
| Billing model | Flat self-serve plans | Flat self-serve plans | Usage-based |
Frequently asked questions
Can I keep my OpenAI SDK code?
Yes. The embeddings endpoint is OpenAI-compatible, so existing SDK calls keep working after you change the base URL and model name.
Do I have to re-embed my whole corpus?
Yes. Vectors are only comparable within the model that produced them, so a new embedding model requires re-embedding every document you want to search.
Can I run both embeddings indexes at once?
Yes, and you should during migration. Dual-query both indexes and compare retrieved chunks and answers before switching production traffic.
How many inputs can one embeddings request contain?
Plugsky accepts up to 2,048 inputs per request, with a maximum of 8,191 tokens per input. Batch your corpus accordingly.
Can I use Plugsky embeddings with my own vector database?
Yes. POST /v1/embeddings is standalone, so you can keep pgvector, Qdrant, Pinecone or another store and only change the embedding calls.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.