Feature × Audience

How do legal teams build embeddings search on Plugsky?

For legal teams, retrieval architecture must enforce matter isolation and citations before relevance tuning becomes the priority. Plugsky exposes an OpenAI-compatible /v1/embeddings endpoint with the plugsky-embed family in a 30+ model catalogue, 8K-class inputs and a vector dimension published per model that you read before creating the collection.

Key facts

EndpointPOST /v1/embeddings (OpenAI-compatible)
Modelsplugsky-embed, plugsky-embed-multilingual and plugsky-embed-nim in a 30+ model catalogue
Input limit8K-class text per request; live limits published per model
Vector dimensionPublished per model on /models — read it before creating a collection
Use casesRAG, semantic search, clustering and recommendations
IndexingBatch embeddings for large corpora; re-embed when the model version changes
Matter isolationSeparate collections per matter with RBAC and citation metadata
PrivilegeAccess controls prevent cross-matter retrieval leakage

TL;DR

  • OpenAI-compatible /v1/embeddings with the plugsky-embed family in a 30+ model catalogue.
  • Dimension is published per model — read it before creating the collection.
  • Isolate collections per matter so retrieval cannot bridge ethical walls.
  • Test with negative queries that should return nothing.
  • Start free with plugsky-micro and plugsky-lite; a 14-day full-access trial covers larger models.

How it works, step by step

  1. Read the model's live vector dimension from the catalogue and create the collection with it.
  2. Chunk by document structure, embed in batches, and store metadata for citations.
  3. Validate retrieval on a labelled question set before moving to production traffic.
  4. Index each matter into a separate collection with citation metadata.
  5. Run negative retrieval tests to prove cross-matter isolation.
  6. Apply retention and deletion rules to vectors at matter close.
1Read the model'slive vectordimension from the2Chunk by documentstructure, embed inbatches, and store3Validate retrievalon a labelledquestion set before4Index each matterinto a separatecollection with5Run negativeretrieval tests toprove cross-matter6Apply retention anddeletion rules tovectors at matter

Original data

POST /v1/embedEndpointplugsky-embed,Models8K-class text Input limitSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the RAG chunk size calculator →

Legal retrieval must respect privilege and matter boundaries as carefully as it respects relevance. Embeddings make large document sets searchable, but a shared index can silently bridge ethical walls, so architecture comes before tuning.

Embeddings run on an OpenAI-compatible /v1/embeddings endpoint, so existing vector pipelines keep their request shape. The plugsky-embed family — plugsky-embed, plugsky-embed-multilingual and plugsky-embed-nim — sits alongside chat models in a 30+ model catalogue, each with an 8K-class input limit and a published vector dimension you can read from the catalogue before indexing.

Architecture and controls

Separate collections per matter or client, keep metadata for citations and access scoping, and enforce RBAC so retrieval cannot cross boundaries. Keep the vector store inside the selected residency plane or your own deployment.

Integration pattern and rollout

Index with structure-aware chunking — clauses, headings, defined terms — and keep source references in metadata so a passage can be produced with its context. Evaluate with questions practitioners actually ask, including negative tests for cross-matter leakage.

The pipeline is straightforward: chunk, embed in batches, store vectors with source metadata, retrieve top-k. What needs discipline is versioning — record the model and dimension with every vector, keep model choice in configuration, and treat a model swap as a data migration with dual-write and a validation phase before cutover.

Limits, evidence and cost

Similarity is not authority: retrieval can surface outdated or superseded material, so ranking and recency signals matter. Plan retention for vectors like any other client record, including deletion at matter close.

Self-serve plans are flat monthly with unlimited fair-use usage, so embedding volume does not introduce per-token billing — check the live pricing page. The free plan includes plugsky-micro and plugsky-lite with no card, and the 14-day full-access trial lets you test before committing to an index.

Honest comparison

ConcernPlugsky embeddingsTypical API providerSelf-hosted embedder
API shapeOpenAI-compatible /v1/embeddingsUsually compatible, variesCustom serving stack
Model choiceplugsky-embed family inside a 30+ model catalogueProvider catalogue onlyYou package each model
ResidencyRegion-locked planes; VPC, on-prem and air-gappedLimited region choicesWherever you deploy
Dimension changesRead live dimension from /models; plan new collectionsVaries by providerYou manage every migration
Operational loadManaged endpoint with batchingManaged endpointGPU capacity, patching and autoscaling
Matter isolationSeparate collections per matter with RBACShared indexes commonYou enforce ethical walls

Frequently asked questions

Do we need to re-embed when we change models?

Yes. Vectors depend on the model and dimensions can change. Plan a new collection, backfill, validate, then cut over — re-embedding is a data migration, not a config flip.

Are embeddings region-locked?

Yes, when the workload is pinned to a region-locked plane. Keep the vector store, backups and logs in the same jurisdiction as the source data.

Which embedding model should we start with?

plugsky-embed for English-dominant corpora, plugsky-embed-multilingual for mixed languages, and plugsky-embed-nim for NVIDIA-style profiles. Evaluate on your own content.

How do we prevent cross-matter leakage?

Use separate collections per matter with RBAC, and test retrieval with negative queries that should return nothing.

Should chunks carry citations?

Yes. Store source references and dates in metadata so every passage can be produced with its context, and recency signals can inform ranking.

How do we handle retention?

Treat vectors as client records: apply the same retention and deletion rules as the documents they came from.