Key facts
| API shape | OpenAI-compatible POST /v1/embeddings |
| Models | plugsky-embed-v1 (1536 dimensions, ada-compatible) and plugsky-embed-large (3072 dimensions) |
| Batch input | Arrays of inputs accepted per request; chunk client-side for larger jobs |
| Use cases | Semantic search, RAG retrieval, clustering, recommendations and classification |
| Languages | Cross-lingual retrieval with strong support for Arabic and other languages |
| Vector stores | Works with pgvector, Pinecone, Qdrant or your own database |
| Pricing model | Embeddings are included in flat self-serve plans, with no per-vector meter |
| Product status | Live |
TL;DR
- Drop-in OpenAI embeddings: change the base URL and keep your code.
- 1536 dimensions for general retrieval, 3072 for high-recall multilingual work.
- Embeddings are included in flat self-serve plans with no per-vector meter.
- Pair embeddings with the RAG API for end-to-end document Q&A with citations.
- Bring your own embedding model on Enterprise private endpoints.
How it works, step by step
- Pick plugsky-embed-v1 for general retrieval or plugsky-embed-large for high recall.
- Chunk documents with enough overlap to preserve meaning across boundaries.
- Call /v1/embeddings with an array of inputs and store the resulting vectors.
- Index vectors in pgvector, Pinecone, Qdrant or your own database.
- Query with the same model used for indexing and rank by cosine similarity.
- Feed the top chunks into a chat model for grounded answers.
- Re-embed when you change models, because vectors from different models are not comparable.
Try it yourself
Open the embedding API tester →
Why embeddings
Embeddings turn text into a fixed-size vector where similar meaning lands at nearby points. That property powers semantic search, recommendation, clustering, RAG retrieval and anomaly detection, all without keyword matching.
Because the endpoint is OpenAI-compatible, any SDK that already speaks OpenAI embeddings works against Plugsky unchanged. You keep your pipeline and change the base URL and model name.
Plugsky's endpoint is part of a 30+ model platform, so embeddings, chat and RAG share one key, one SDK and one flat plan.Two models, two use cases
plugsky-embed-v1 returns 1536 dimensions and is the general-purpose, OpenAI ada-compatible option for fast retrieval. plugsky-embed-large returns 3072 dimensions and targets high-recall retrieval, multilingual content and longer text.
Choose by query type, not by size alone. If recall matters more than latency and your corpus mixes languages, the larger model usually pays for itself; for high-volume English retrieval the smaller model is often enough.
Batch, storage and use cases
The endpoint accepts an array of inputs per request, so you can embed a batch of chunks in one call and write the vectors to your store. For very large jobs, chunk client-side and submit in manageable batches.
Common patterns: RAG retrieval with the RAG API, semantic search over pgvector or Pinecone, near-duplicate detection with cosine similarity, recommendation feeds, k-means topic clustering, and zero-shot classification by comparing an embedding to labelled centroids. On Enterprise, private endpoints can host your own embedding models.
Because the endpoint accepts arrays, a corpus of a few thousand chunks can be embedded in a handful of calls rather than one request per document, which keeps large ingestion jobs practical.Honest comparison
| Task | plugsky-embed-v1 | plugsky-embed-large | Bring your own model |
|---|---|---|---|
| Dimensions | 1536 | 3072 | Varies |
| Compatibility | OpenAI ada-style | OpenAI-compatible endpoint | Varies by model |
| Best for | General retrieval, fast | High-recall multilingual and longer text | Specialised workloads |
| Vector stores | pgvector, Pinecone, Qdrant | Same | Same |
| Pricing | Included in flat plans | Included in flat plans | Enterprise contract |
| Availability | Self-serve | Self-serve | Enterprise private endpoint |
Frequently asked questions
Are embeddings cached?
No. Every request computes a fresh embedding, so cache results in your own layer if you re-embed the same content.
Can I bring my own embedding model?
Yes, on Enterprise contracts. Plugsky supports Cohere, Voyage, BGE and custom models via the private endpoint.
What languages are supported?
The platform targets cross-lingual retrieval with strong Arabic support, and plugsky-embed-large is the better choice for multilingual corpora. Check the model catalogue for current coverage.
What are the pricing implications?
Embeddings are included in flat self-serve plans with no per-vector meter. See the live pricing page for tier details.
How many inputs can one request take?
The endpoint accepts arrays of inputs, and the docs describe limits per request and per input. For larger jobs, batch client-side.
Which model should I choose?
Use plugsky-embed-v1 for general, high-volume retrieval and plugsky-embed-large when recall on multilingual or longer text matters more than latency.
Plugsky (2026). “Embeddings API: OpenAI-Compatible Vectors”. Plugsky. Available at: https://plugsky.com/articles/embeddings-api (last updated 2026-09-25).