Key facts
| Training use | API data is not used to train models |
| Encryption | Per-collection encryption at rest; BYOK on enterprise |
| Residency | Region selection with region-locked deployment options |
| Access control | Scoped API keys, RBAC and SSO on enterprise, per-collection isolation |
| Audit logs | Per-request model, tokens, latency, user and region with SIEM export |
| Deployment | Managed, VPC, on-prem and air-gapped options |
| Retention | Configurable retention and deletion workflows |
| Product status | Live |
TL;DR
- Confirm the no-training commitment for API data in writing.
- Encryption at rest and key custody are baseline requirements, not extras.
- Residency applies to documents, embeddings, queries and logs alike.
- Scope access per collection so one key cannot reach every corpus.
- Agree deletion and retention before loading sensitive documents.
How it works, step by step
- Classify documents by sensitivity and define what each class is allowed to do.
- Verify the no-training commitment and data processing terms for your deployment.
- Choose a region or private deployment plane that satisfies residency rules.
- Create collections that match permission groups and enable encryption.
- Issue scoped keys and enable audit logging with SIEM export.
- Define retention and deletion workflows, then test them on a sample collection.
- Review the DPA and sub-processor list before onboarding regulated data.
Try it yourself
Open the AI data residency checklist →
The five questions that define RAG privacy
Is data used for training? Plugsky states that API data is not used to train models, which is the baseline commitment to verify in your agreement. Is data encrypted? Collections are encrypted at rest, with customer-managed keys available on enterprise so key custody stays with you.
Where does data live? Residency covers documents, embeddings, queries, logs and backups, not just the primary store. Who can reach it? Scoped keys, per-collection isolation, RBAC and SSO define the access surface. How long does it stay? Retention and deletion workflows decide when data actually leaves the system.
Why embeddings and queries are also sensitive
Teams often treat embeddings as anonymous. They are derived from the source text and, combined with metadata, can expose meaning and structure. Query text is at least as revealing: it shows what users are looking for, which in legal, medical or financial settings can be more sensitive than the documents themselves.
A privacy review should therefore cover the whole pipeline: ingestion, chunk storage, vector storage, retrieval logs, generation prompts and monitoring exports. Any layer that crosses a boundary without the same controls undoes the rest.
Deployment choices for stricter requirements
Managed deployment with a region lock suits many workloads and keeps operations simple. A VPC deployment keeps processing inside your cloud account with private networking. On-prem and air-gapped deployments place the workload in your data centre or an isolated network where no egress is permitted, with the same OpenAI-compatible endpoints so application code does not change.
Choose the plane per corpus. Applying the strictest option to public documentation adds cost and slows delivery, while applying a permissive one to regulated records creates risk.
Operating privacy day to day
Privacy is maintained by process as much as configuration. Rotate keys on a schedule, review collection membership when teams change, export audit logs to a system with its own retention, and rehearse deletion so a closed project can be removed completely. Keep an inventory of collections with owner, classification and retention, and treat it as a living document.
Work through the data residency checklist before launch, and prototype on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Control | Plugsky managed | VPC or on-prem | Unmanaged pipeline |
|---|---|---|---|
| Training use | API data not used to train models | Same commitment | Depends on every vendor |
| Encryption | Per-collection at rest; BYOK on enterprise | Customer keys in your environment | You implement it |
| Residency | Region selection and locking | Your infrastructure | Multiple vendors, multiple regions |
| Access | Scoped keys, RBAC and SSO | Same controls plus network policy | Custom, easy to misconfigure |
| Deletion | Documented workflows | You control the data plane | Across every component |
Frequently asked questions
Does Plugsky train on my documents?
Plugsky states that API data is not used to train models, and collections are encrypted at rest. Confirm the commitment in your agreement for your deployment.
Are embeddings considered personal data?
They can be. Embeddings derive from source content and can expose meaning when combined with metadata, so treat them with the same care as the text.
Can data stay in one region?
Yes. Plugsky supports region selection and region-locked deployment options, and private deployment planes keep processing inside your perimeter.
How do I delete a collection?
Deletion workflows remove the collection and its documents. Agree retention for logs and backups with your team so removal is complete, not partial.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.