Key facts
| Endpoint | POST /v1/embeddings (OpenAI-compatible) |
| Models | plugsky-embed, plugsky-embed-multilingual and plugsky-embed-nim in a 30+ model catalogue |
| Input limit | 8K-class text per request; live limits published per model |
| Vector dimension | Published per model on /models — read it before creating a collection |
| Use cases | RAG, semantic search, clustering and recommendations |
| Indexing | Batch embeddings for large corpora; re-embed when the model version changes |
| Volume | Batch indexing for large document and ticket corpora |
| Residency | Keep vectors and logs in-country planes |
TL;DR
- OpenAI-compatible /v1/embeddings with the plugsky-embed family in a 30+ model catalogue.
- Dimension is published per model — read it before creating the collection.
- Pilot embedding search on internal knowledge before customer-facing use.
- Version the pipeline; large corpora make model changes expensive.
- Start free with plugsky-micro and plugsky-lite; a 14-day full-access trial covers larger models.
How it works, step by step
- Read the model's live vector dimension from the catalogue and create the collection with it.
- Chunk by document structure, embed in batches, and store metadata for citations.
- Validate retrieval on a labelled question set before moving to production traffic.
- Pilot on runbooks and historical tickets to prove retrieval quality.
- Batch-index with model version and source metadata recorded.
- Pin vectors and logs to in-country planes and audit access.
Original data
Try it yourself
Open the embedding model comparison →
Embeddings for telcos: what changes
Telcos apply embeddings to support knowledge, ticket triage, contract search and multilingual customer service, often across archives no single team fully knows. The volume argues for batch pipelines and disciplined metadata; the data argues for in-country processing.
Embeddings run on an OpenAI-compatible /v1/embeddings endpoint, so existing vector pipelines keep their request shape. The plugsky-embed family — plugsky-embed, plugsky-embed-multilingual and plugsky-embed-nim — sits alongside chat models in a 30+ model catalogue, each with an 8K-class input limit and a published vector dimension you can read from the catalogue before indexing.
Architecture and controls
Batch-index large corpora, keep vectors and logs in the approved jurisdiction, and enforce access through RBAC aligned to operational roles. Record model version and source references so answers can be traced and audited.
Integration pattern and rollout
Pilot on internal knowledge — runbooks, procedures, past tickets — to prove retrieval quality and operational fit, then extend to customer-facing support with stricter access controls and residency pinning.
The pipeline is straightforward: chunk, embed in batches, store vectors with source metadata, retrieve top-k. What needs discipline is versioning — record the model and dimension with every vector, keep model choice in configuration, and treat a model swap as a data migration with dual-write and a validation phase before cutover.
Limits, evidence and cost
Large corpora make model changes expensive, so version the pipeline and plan re-embedding windows. Residency requirements may also limit which managed vector services you can use; confirm the full path, including backups.
Self-serve plans are flat monthly with unlimited fair-use usage, so embedding volume does not introduce per-token billing — check the live pricing page. The free plan includes plugsky-micro and plugsky-lite with no card, and the 14-day full-access trial lets you test before committing to an index.
Honest comparison
| Concern | Plugsky embeddings | Typical API provider | Self-hosted embedder |
|---|---|---|---|
| API shape | OpenAI-compatible /v1/embeddings | Usually compatible, varies | Custom serving stack |
| Model choice | plugsky-embed family inside a 30+ model catalogue | Provider catalogue only | You package each model |
| Residency | Region-locked planes; VPC, on-prem and air-gapped | Limited region choices | Wherever you deploy |
| Dimension changes | Read live dimension from /models; plan new collections | Varies by provider | You manage every migration |
| Operational load | Managed endpoint with batching | Managed endpoint | GPU capacity, patching and autoscaling |
| Corpus scale | Batch indexing with in-country storage | Varies by provider | You plan capacity |
Frequently asked questions
Do we need to re-embed when we change models?
Yes. Vectors depend on the model and dimensions can change. Plan a new collection, backfill, validate, then cut over — re-embedding is a data migration, not a config flip.
Are embeddings region-locked?
Yes, when the workload is pinned to a region-locked plane. Keep the vector store, backups and logs in the same jurisdiction as the source data.
Which embedding model should we start with?
plugsky-embed for English-dominant corpora, plugsky-embed-multilingual for mixed languages, and plugsky-embed-nim for NVIDIA-style profiles. Evaluate on your own content.
Can embeddings handle very large corpora?
Yes, with batch indexing and incremental refresh. Version the pipeline and schedule re-embedding windows when models change.
How do we keep subscriber data in-country?
Run the embedding pipeline and vector store in the approved jurisdiction, and confirm backups and logs follow the same boundary.
Do multilingual models help support workloads?
Yes — use the multilingual model where customer interactions and knowledge bases span languages, and validate with local test sets.