Key facts
| Deployment | Hosted, VPC, on-prem and air-gapped options |
| RAG endpoints | POST /v1/embeddings, POST /v1/rag/collections, POST /v1/rag/query |
| Encryption | Per-collection encryption at rest; BYOK via KMS, Key Vault, Vault or HSM |
| Audit logs | Per-request model, tokens, latency, user and region; SIEM export and retention options |
| Access control | Scoped API keys, RBAC and SSO for enterprise; per-collection isolation |
| Data use | API data is not used to train models |
| Retrieval | Keyword, vector and hybrid search with optional reranking and citations |
| Product status | Live |
TL;DR
- Choose the deployment plane per corpus, not one setting for the whole company.
- Encryption at rest and customer-managed keys are table stakes for private RAG.
- Scoped keys and per-collection isolation keep access attributable.
- Audit logs turn retrieval into something you can review and evidence.
- The API stays OpenAI-compatible, so moving planes does not rewrite the app.
How it works, step by step
- Classify each corpus by sensitivity, residency requirement and retention rule.
- Map corpora to deployment planes: managed, VPC, on-prem or air-gapped.
- Create collections with encryption scopes that match the classification.
- Issue scoped API keys per team or application and enable audit logging.
- Connect logs to your SIEM and agree retention with security and compliance.
- Run a retrieval evaluation inside the target plane before loading production data.
- Document the rollback and deletion workflow for each corpus.
Try it yourself
Open the private LLM deployment estimator →
What private RAG has to protect
Private RAG protects three things: the source documents, the embeddings derived from them, and the queries users ask. Embeddings are not anonymised data; they encode content and can expose information if leaked alongside metadata. Query logs reveal what people are looking for, which is often as sensitive as the corpus itself.
That is why a private deployment decision covers the whole pipeline: ingestion, storage, retrieval, generation and observability. A pipeline that keeps documents on-prem but sends every query to a shared endpoint is only partly private, and the gap is usually the one auditors ask about.
Choosing the deployment plane
A VPC deployment keeps processing inside your cloud account with private networking, which fits teams already standardised on one provider. On-prem puts the workload in your data centre for workloads with strict internal policy. Air-gapped removes network access entirely and is used for classified or safety-critical environments where no egress is acceptable.
Match the plane to the corpus rather than applying the strictest option everywhere. A public product manual and a set of employment contracts rarely deserve the same controls, and over-classifying everything slows delivery without improving security.
The controls to configure first
Start with encryption: per-collection encryption at rest, plus customer-managed keys through a KMS, Key Vault, Vault or an on-prem HSM so key custody stays with you. Then lock down access with scoped API keys per application, RBAC and SSO for people, and per-collection isolation so one key cannot query another team's data.
Turn on per-request audit logging covering model, tokens, latency, user and region, export it to your SIEM, and agree retention with compliance. Finally, confirm in writing that API data is not used for model training and that the DPA and sub-processor list cover the deployment you chose.
Running it on Plugsky
Plugsky keeps the RAG and embeddings endpoints OpenAI-compatible across deployment planes, so code written against POST /v1/embeddings and POST /v1/rag/query keeps working when the data plane changes. Retrieval returns ranked chunks with citations, and 30+ models sit behind the same API for generation.
Estimate the shape of the deployment with the private deployment estimator, then prototype on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Control | Plugsky managed | VPC deployment | On-prem or air-gapped |
|---|---|---|---|
| Data location | Plugsky cloud with region choice | Your cloud account | Your data centre or isolated network |
| Key custody | Managed encryption | Customer-managed keys available | BYOK with KMS or on-prem HSM |
| Egress | Region-locked options | Private networking | None |
| App changes | None - OpenAI-compatible | None - same endpoints | None - same endpoints |
| Best for | Standard workloads | Residency with cloud operations | Strict policy and classified data |
Frequently asked questions
What makes RAG private?
Keeping documents, embeddings, queries and logs inside a boundary you control, with encryption, scoped access and audit logging enabled across the whole pipeline.
Can Plugsky run inside our VPC?
Yes. Plugsky supports VPC, on-prem and air-gapped deployments for enterprise customers with the same OpenAI-compatible endpoints.
Do you train on our documents?
Plugsky states that API data is not used to train models, and collections are encrypted at rest. Confirm the current terms in your agreement.
How do audit logs help?
They record per-request model, tokens, latency, user and region, which supports access reviews, incident investigation and compliance evidence when exported to your SIEM.
Is there a free plan for evaluation?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage, and enterprise deployments are quoted separately. See the live pricing page for current self-serve plans.