Developer + API

What is the Plugsky embeddings API?

An embedding is a dense numeric vector that captures the semantic meaning of text. Plugsky's embeddings endpoint follows the OpenAI shape, so any tool or SDK that speaks OpenAI embeddings works after a base_url change. Plugsky ships an ada-compatible 1536-dimension model and a 3072-dimension model for high-recall multilingual retrieval, including Arabic.

Key facts

API shapeOpenAI-compatible POST /v1/embeddings
Modelsplugsky-embed-v1 (1536 dimensions, ada-compatible) and plugsky-embed-large (3072 dimensions)
Batch inputArrays of inputs accepted per request; chunk client-side for larger jobs
Use casesSemantic search, RAG retrieval, clustering, recommendations and classification
LanguagesCross-lingual retrieval with strong support for Arabic and other languages
Vector storesWorks with pgvector, Pinecone, Qdrant or your own database
Pricing modelEmbeddings are included in flat self-serve plans, with no per-vector meter
Product statusLive

TL;DR

  • Drop-in OpenAI embeddings: change the base URL and keep your code.
  • 1536 dimensions for general retrieval, 3072 for high-recall multilingual work.
  • Embeddings are included in flat self-serve plans with no per-vector meter.
  • Pair embeddings with the RAG API for end-to-end document Q&A with citations.
  • Bring your own embedding model on Enterprise private endpoints.

How it works, step by step

  1. Pick plugsky-embed-v1 for general retrieval or plugsky-embed-large for high recall.
  2. Chunk documents with enough overlap to preserve meaning across boundaries.
  3. Call /v1/embeddings with an array of inputs and store the resulting vectors.
  4. Index vectors in pgvector, Pinecone, Qdrant or your own database.
  5. Query with the same model used for indexing and rank by cosine similarity.
  6. Feed the top chunks into a chat model for grounded answers.
  7. Re-embed when you change models, because vectors from different models are not comparable.
1Pickplugsky-embed-v1for general2Chunk documentswith enough overlapto preserve meaning3Call /v1/embeddingswith an array ofinputs and store4Index vectors inpgvector, Pinecone,Qdrant or your own5Query with the samemodel used forindexing and rank6Feed the top chunksinto a chat modelfor grounded

Try it yourself

Open the embedding API tester →

Why embeddings

Embeddings turn text into a fixed-size vector where similar meaning lands at nearby points. That property powers semantic search, recommendation, clustering, RAG retrieval and anomaly detection, all without keyword matching.

Because the endpoint is OpenAI-compatible, any SDK that already speaks OpenAI embeddings works against Plugsky unchanged. You keep your pipeline and change the base URL and model name.

Plugsky's endpoint is part of a 30+ model platform, so embeddings, chat and RAG share one key, one SDK and one flat plan.

Two models, two use cases

plugsky-embed-v1 returns 1536 dimensions and is the general-purpose, OpenAI ada-compatible option for fast retrieval. plugsky-embed-large returns 3072 dimensions and targets high-recall retrieval, multilingual content and longer text.

Choose by query type, not by size alone. If recall matters more than latency and your corpus mixes languages, the larger model usually pays for itself; for high-volume English retrieval the smaller model is often enough.

Batch, storage and use cases

The endpoint accepts an array of inputs per request, so you can embed a batch of chunks in one call and write the vectors to your store. For very large jobs, chunk client-side and submit in manageable batches.

Common patterns: RAG retrieval with the RAG API, semantic search over pgvector or Pinecone, near-duplicate detection with cosine similarity, recommendation feeds, k-means topic clustering, and zero-shot classification by comparing an embedding to labelled centroids. On Enterprise, private endpoints can host your own embedding models.

Because the endpoint accepts arrays, a corpus of a few thousand chunks can be embedded in a handful of calls rather than one request per document, which keeps large ingestion jobs practical.

Honest comparison

Taskplugsky-embed-v1plugsky-embed-largeBring your own model
Dimensions15363072Varies
CompatibilityOpenAI ada-styleOpenAI-compatible endpointVaries by model
Best forGeneral retrieval, fastHigh-recall multilingual and longer textSpecialised workloads
Vector storespgvector, Pinecone, QdrantSameSame
PricingIncluded in flat plansIncluded in flat plansEnterprise contract
AvailabilitySelf-serveSelf-serveEnterprise private endpoint

Frequently asked questions

Are embeddings cached?

No. Every request computes a fresh embedding, so cache results in your own layer if you re-embed the same content.

Can I bring my own embedding model?

Yes, on Enterprise contracts. Plugsky supports Cohere, Voyage, BGE and custom models via the private endpoint.

What languages are supported?

The platform targets cross-lingual retrieval with strong Arabic support, and plugsky-embed-large is the better choice for multilingual corpora. Check the model catalogue for current coverage.

What are the pricing implications?

Embeddings are included in flat self-serve plans with no per-vector meter. See the live pricing page for tier details.

How many inputs can one request take?

The endpoint accepts arrays of inputs, and the docs describe limits per request and per input. For larger jobs, batch client-side.

Which model should I choose?

Use plugsky-embed-v1 for general, high-volume retrieval and plugsky-embed-large when recall on multilingual or longer text matters more than latency.

Cite this page

Plugsky (2026). “Embeddings API: OpenAI-Compatible Vectors”. Plugsky. Available at: https://plugsky.com/articles/embeddings-api (last updated 2026-09-25).