Key facts
| Endpoint | POST /v1/embeddings (OpenAI-compatible) |
| Models | plugsky-embed, plugsky-embed-multilingual and plugsky-embed-nim in a 30+ model catalogue |
| Input limit | 8K-class text per request; live limits published per model |
| Vector dimension | Published per model on /models — read it before creating a collection |
| Use cases | RAG, semantic search, clustering and recommendations |
| Indexing | Batch embeddings for large corpora; re-embed when the model version changes |
| Matter isolation | Separate collections per matter with RBAC and citation metadata |
| Privilege | Access controls prevent cross-matter retrieval leakage |
TL;DR
- OpenAI-compatible /v1/embeddings with the plugsky-embed family in a 30+ model catalogue.
- Dimension is published per model — read it before creating the collection.
- Isolate collections per matter so retrieval cannot bridge ethical walls.
- Test with negative queries that should return nothing.
- Start free with plugsky-micro and plugsky-lite; a 14-day full-access trial covers larger models.
How it works, step by step
- Read the model's live vector dimension from the catalogue and create the collection with it.
- Chunk by document structure, embed in batches, and store metadata for citations.
- Validate retrieval on a labelled question set before moving to production traffic.
- Index each matter into a separate collection with citation metadata.
- Run negative retrieval tests to prove cross-matter isolation.
- Apply retention and deletion rules to vectors at matter close.
Original data
Try it yourself
Open the RAG chunk size calculator →
Embeddings for legal teams: what changes
Legal retrieval must respect privilege and matter boundaries as carefully as it respects relevance. Embeddings make large document sets searchable, but a shared index can silently bridge ethical walls, so architecture comes before tuning.
Embeddings run on an OpenAI-compatible /v1/embeddings endpoint, so existing vector pipelines keep their request shape. The plugsky-embed family — plugsky-embed, plugsky-embed-multilingual and plugsky-embed-nim — sits alongside chat models in a 30+ model catalogue, each with an 8K-class input limit and a published vector dimension you can read from the catalogue before indexing.
Architecture and controls
Separate collections per matter or client, keep metadata for citations and access scoping, and enforce RBAC so retrieval cannot cross boundaries. Keep the vector store inside the selected residency plane or your own deployment.
Integration pattern and rollout
Index with structure-aware chunking — clauses, headings, defined terms — and keep source references in metadata so a passage can be produced with its context. Evaluate with questions practitioners actually ask, including negative tests for cross-matter leakage.
The pipeline is straightforward: chunk, embed in batches, store vectors with source metadata, retrieve top-k. What needs discipline is versioning — record the model and dimension with every vector, keep model choice in configuration, and treat a model swap as a data migration with dual-write and a validation phase before cutover.
Limits, evidence and cost
Similarity is not authority: retrieval can surface outdated or superseded material, so ranking and recency signals matter. Plan retention for vectors like any other client record, including deletion at matter close.
Self-serve plans are flat monthly with unlimited fair-use usage, so embedding volume does not introduce per-token billing — check the live pricing page. The free plan includes plugsky-micro and plugsky-lite with no card, and the 14-day full-access trial lets you test before committing to an index.
Honest comparison
| Concern | Plugsky embeddings | Typical API provider | Self-hosted embedder |
|---|---|---|---|
| API shape | OpenAI-compatible /v1/embeddings | Usually compatible, varies | Custom serving stack |
| Model choice | plugsky-embed family inside a 30+ model catalogue | Provider catalogue only | You package each model |
| Residency | Region-locked planes; VPC, on-prem and air-gapped | Limited region choices | Wherever you deploy |
| Dimension changes | Read live dimension from /models; plan new collections | Varies by provider | You manage every migration |
| Operational load | Managed endpoint with batching | Managed endpoint | GPU capacity, patching and autoscaling |
| Matter isolation | Separate collections per matter with RBAC | Shared indexes common | You enforce ethical walls |
Frequently asked questions
Do we need to re-embed when we change models?
Yes. Vectors depend on the model and dimensions can change. Plan a new collection, backfill, validate, then cut over — re-embedding is a data migration, not a config flip.
Are embeddings region-locked?
Yes, when the workload is pinned to a region-locked plane. Keep the vector store, backups and logs in the same jurisdiction as the source data.
Which embedding model should we start with?
plugsky-embed for English-dominant corpora, plugsky-embed-multilingual for mixed languages, and plugsky-embed-nim for NVIDIA-style profiles. Evaluate on your own content.
How do we prevent cross-matter leakage?
Use separate collections per matter with RBAC, and test retrieval with negative queries that should return nothing.
Should chunks carry citations?
Yes. Store source references and dates in metadata so every passage can be produced with its context, and recency signals can inform ranking.
How do we handle retention?
Treat vectors as client records: apply the same retention and deletion rules as the documents they came from.