RAG

What is a practical OpenAI embeddings alternative?

A practical OpenAI embeddings alternative keeps the same request shape so your code barely changes: an OpenAI-compatible endpoint that returns vectors with documented dimensions. Plugsky exposes POST /v1/embeddings with plugsky-embed-v1 at 1536 dimensions, which matches common ada-style pipelines and makes a staged re-embedding migration realistic.

Key facts

EndpointPOST /v1/embeddings, OpenAI-compatible request shape
Modelsplugsky-embed-v1 (1536d), plugsky-embed-large (3072d) and plugsky-embed-multilingual
Batch limitsUp to 2,048 inputs per request, max 8,191 tokens each
Compatibilityplugsky-embed-v1 returns 1536-dimension vectors for ada-style pipelines
RAG usageVectors feed keyword, vector and hybrid retrieval in RAG collections
Data handlingAPI data is not used to train models
MigrationChange base URL and model name; re-embed to fill a new index
Product statusLive

TL;DR

  • The API call is easy to swap; the re-embedding of your corpus is the real migration.
  • Keep a dimension-compatible model first so application and index code change minimally.
  • Build the new index in parallel and dual-query both until quality is proven.
  • Never mix vectors from two embedding models in one index.
  • Plugsky embeddings work standalone with your existing vector database.

How it works, step by step

  1. Inventory every place you call the embeddings endpoint and how dimensions are stored.
  2. Create a parallel index for the new model; leave the old one serving traffic.
  3. Re-embed a representative slice and compare recall on a golden question set.
  4. Re-embed the full corpus in batches, respecting the 2,048-input request limit.
  5. Dual-query old and new indexes and compare retrieved chunks per question.
  6. Cut retrieval traffic to the new index once quality is equal or better.
  7. Keep the old index for a rollback window, then remove it and update your docs.
1Inventory everyplace you call theembeddings endpoint2Create a parallelindex for the newmodel; leave the3Re-embed arepresentativeslice and compare4Re-embed the fullcorpus in batches,respecting the5Dual-query old andnew indexes andcompare retrieved6Cut retrievaltraffic to the newindex once quality

Original data

POST /v1/embedEndpointplugsky-embed-ModelsUp to 2,048 inBatch limitsplugsky-embed-CompatibilitySource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the embedding API tester →

What actually changes when you switch embeddings

If the provider is OpenAI-compatible, the HTTP call is nearly identical: same endpoint path, same input list, same response shape. Your client library keeps working after a base URL and model name change. What changes is the vector space: dimensions may differ, distances are no longer comparable with the old model, and every stored vector becomes stale for the new model.

That is why an embeddings migration is mostly a data migration. The safest pattern is a parallel index: keep the current index serving production while you re-embed the corpus into a new one, then compare retrieval quality before moving traffic.

Choosing the replacement model

Two paths work. A dimension-compatible model such as plugsky-embed-v1 at 1536 dimensions minimises code changes for pipelines built on common ada-style embeddings, because index schemas and downstream assumptions stay valid. A higher-dimension model such as plugsky-embed-large at 3072 dimensions can represent more nuance but requires a new index and more storage.

For non-English corpora, test the multilingual model before anything else; translation before embedding adds latency, cost and failure modes. Whatever you choose, validate with a golden question set rather than a public leaderboard, and check the maximum input length so long chunks are not silently truncated.

Running the re-embedding safely

Batch the work, track progress per document, and make the job idempotent so a failure can resume instead of restarting. Plugsky accepts up to 2,048 inputs per embeddings request with a maximum of 8,191 tokens per input, so chunk sizes and batch sizes should be planned together. Record the model name and version alongside the index metadata; a future migration will need to know which model produced which vectors.

While both indexes exist, dual-query them for the same user questions and compare the chunks returned and the answers produced. This catches regressions that a spot check misses, and it gives you a defensible cutover decision.

What you gain on Plugsky

Plugsky keeps embeddings on the same OpenAI-compatible API as chat and RAG, so one key and one base URL cover the stack. Vectors return from POST /v1/embeddings, collections and queries handle retrieval when you want managed RAG, and API data is not used to train models. Self-serve plans are flat monthly with no per-token billing, which removes one source of migration uncertainty.

Try requests with the embedding API tester, then start free with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

Factorplugsky-embed-v1plugsky-embed-largeStaying on OpenAI embeddings
Dimensions15363072Model-dependent
Request shapeOpenAI-compatibleOpenAI-compatibleNative
Migration effortBase URL plus model nameNew index for higher dimensionsNone now, vendor dependency remains
MultilingualDedicated multilingual model availableDedicated multilingual model availableModel-dependent
Data handlingAPI data not used for trainingSameProvider terms apply
Billing modelFlat self-serve plansFlat self-serve plansUsage-based

Frequently asked questions

Can I keep my OpenAI SDK code?

Yes. The embeddings endpoint is OpenAI-compatible, so existing SDK calls keep working after you change the base URL and model name.

Do I have to re-embed my whole corpus?

Yes. Vectors are only comparable within the model that produced them, so a new embedding model requires re-embedding every document you want to search.

Can I run both embeddings indexes at once?

Yes, and you should during migration. Dual-query both indexes and compare retrieved chunks and answers before switching production traffic.

How many inputs can one embeddings request contain?

Plugsky accepts up to 2,048 inputs per request, with a maximum of 8,191 tokens per input. Batch your corpus accordingly.

Can I use Plugsky embeddings with my own vector database?

Yes. POST /v1/embeddings is standalone, so you can keep pgvector, Qdrant, Pinecone or another store and only change the embedding calls.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.