Key facts
| Deployment | VPC, on-prem and air-gapped options for sensitive workloads |
| Encryption | Per-collection encryption at rest; BYOK through KMS, Key Vault, Vault or HSM |
| Access control | Scoped keys and per-collection isolation; RBAC and SSO on enterprise |
| Audit logs | Per-request model, tokens, latency, user and region with SIEM export |
| Data use | API data is not used to train models |
| Retrieval | Keyword, vector and hybrid search with citations for review |
| Formats | PDF, DOCX, TXT, MD and HTML ingestion |
| Product status | Live |
TL;DR
- Classify clinical, operational and public content separately before ingestion.
- Use a private deployment plane for anything containing patient data.
- Scope keys per system and log every query with user identity.
- Citations keep answers reviewable against the source document.
- Keep humans accountable for clinical decisions; RAG is a retrieval tool.
How it works, step by step
- Classify content: clinical records, operational docs, policies and public material.
- Choose a deployment plane per class; keep patient-adjacent data inside the perimeter.
- Create collections aligned with care teams and access policies.
- Enable encryption, customer-managed keys and audit logging before loading data.
- Ingest with metadata such as source system, date and author.
- Require citations and review paths for any content used in care settings.
- Rehearse deletion and retention workflows for closed records.
Try it yourself
Open the private LLM deployment estimator →
What is different about healthcare data
Health content combines three sensitivities: patient-identifiable records, clinical guidance that shapes decisions, and operational documents that describe how care is delivered. Each carries different rules. A retrieval system that mixes them lets operational queries touch clinical content, and a query log can itself become a sensitive record of who looked for what.
That is why classification comes first. Decide what may run in a managed environment, what must stay in a VPC, and what requires on-prem or air-gapped deployment, then build collections that match the boundaries rather than one index for everything.
Controls to configure before ingestion
Encryption at rest per collection, customer-managed keys through a KMS, Key Vault, Vault or an on-prem HSM, and scoped API keys per system are the baseline. Add per-request audit logs covering model, tokens, latency, user and region, exported to your SIEM, and agree retention with your compliance team before the first document loads.
Confirm the no-training commitment for API data in your agreement, and review the DPA and sub-processor list for the deployment you chose. These are procurement blockers if discovered late, and cheap to confirm early.
Retrieval quality in a clinical context
Healthcare questions often hinge on exact terms, drug names, codes and versions of guidance. Hybrid retrieval helps because keyword matching catches identifiers that embeddings may blur, and citations keep every answer reviewable against the source passage. Keep documents versioned so an answer can be tied to the guidance in force at the time.
Refusal behaviour matters most here. When the corpus does not support an answer, the system should say so and point to the appropriate process. Build review workflows around the assistant rather than replacing professional judgment.
Deploying on Plugsky
Plugsky supports managed, VPC, on-prem and air-gapped deployment with the same OpenAI-compatible endpoints, so code does not change when the data plane does. Collections handle ingestion and retrieval with keyword, vector and hybrid modes, optional reranking and citations, and 30+ models sit behind one API for generation.
Estimate the deployment shape with the private deployment estimator, and run evaluations on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Control | Managed with region lock | VPC deployment | On-prem or air-gapped |
|---|---|---|---|
| Data location | Plugsky cloud in a chosen region | Your cloud account | Your data centre or isolated network |
| Key custody | Managed encryption; BYOK on enterprise | Customer-managed keys | KMS or on-prem HSM |
| Audit | Per-request logs with SIEM export | Same with network policy | Same, fully internal |
| App changes | None | None | None |
| Fits | Operational and policy content | Patient-adjacent workloads | Classified or offline requirements |
Frequently asked questions
Can Plugsky be used for patient data?
Plugsky offers VPC, on-prem and air-gapped deployment with encryption, scoped access and audit logs. Confirm the agreement, DPA and residency terms for your jurisdiction before loading patient data.
Does Plugsky train on the data?
Plugsky states that API data is not used to train models, and collections are encrypted at rest.
How do audit logs support compliance?
They record per-request model, tokens, latency, user and region, export to your SIEM, and can be retained according to your policy for access reviews and investigations.
Is RAG a medical device or diagnostic tool?
No. It is a retrieval and summarisation tool. Keep professional review and accountability in the workflow and be explicit about that boundary with users.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. Enterprise deployments are quoted separately.