RAG

How do you build RAG for healthcare securely?

Healthcare RAG must treat documents, embeddings and queries as sensitive data. The core controls are private deployment such as VPC, on-prem or air-gapped, per-collection encryption, scoped access, audit logging and citations for traceability. Plugsky provides these with OpenAI-compatible endpoints and no training on API data.

Key facts

DeploymentVPC, on-prem and air-gapped options for sensitive workloads
EncryptionPer-collection encryption at rest; BYOK through KMS, Key Vault, Vault or HSM
Access controlScoped keys and per-collection isolation; RBAC and SSO on enterprise
Audit logsPer-request model, tokens, latency, user and region with SIEM export
Data useAPI data is not used to train models
RetrievalKeyword, vector and hybrid search with citations for review
FormatsPDF, DOCX, TXT, MD and HTML ingestion
Product statusLive

TL;DR

  • Classify clinical, operational and public content separately before ingestion.
  • Use a private deployment plane for anything containing patient data.
  • Scope keys per system and log every query with user identity.
  • Citations keep answers reviewable against the source document.
  • Keep humans accountable for clinical decisions; RAG is a retrieval tool.

How it works, step by step

  1. Classify content: clinical records, operational docs, policies and public material.
  2. Choose a deployment plane per class; keep patient-adjacent data inside the perimeter.
  3. Create collections aligned with care teams and access policies.
  4. Enable encryption, customer-managed keys and audit logging before loading data.
  5. Ingest with metadata such as source system, date and author.
  6. Require citations and review paths for any content used in care settings.
  7. Rehearse deletion and retention workflows for closed records.
1Classify content:clinical records,operational docs,2Choose a deploymentplane per class;keep3Create collectionsaligned with careteams and access4Enable encryption,customer-managedkeys and audit5Ingest withmetadata such assource system, date6Require citationsand review pathsfor any content

Try it yourself

Open the private LLM deployment estimator →

What is different about healthcare data

Health content combines three sensitivities: patient-identifiable records, clinical guidance that shapes decisions, and operational documents that describe how care is delivered. Each carries different rules. A retrieval system that mixes them lets operational queries touch clinical content, and a query log can itself become a sensitive record of who looked for what.

That is why classification comes first. Decide what may run in a managed environment, what must stay in a VPC, and what requires on-prem or air-gapped deployment, then build collections that match the boundaries rather than one index for everything.

Controls to configure before ingestion

Encryption at rest per collection, customer-managed keys through a KMS, Key Vault, Vault or an on-prem HSM, and scoped API keys per system are the baseline. Add per-request audit logs covering model, tokens, latency, user and region, exported to your SIEM, and agree retention with your compliance team before the first document loads.

Confirm the no-training commitment for API data in your agreement, and review the DPA and sub-processor list for the deployment you chose. These are procurement blockers if discovered late, and cheap to confirm early.

Retrieval quality in a clinical context

Healthcare questions often hinge on exact terms, drug names, codes and versions of guidance. Hybrid retrieval helps because keyword matching catches identifiers that embeddings may blur, and citations keep every answer reviewable against the source passage. Keep documents versioned so an answer can be tied to the guidance in force at the time.

Refusal behaviour matters most here. When the corpus does not support an answer, the system should say so and point to the appropriate process. Build review workflows around the assistant rather than replacing professional judgment.

Deploying on Plugsky

Plugsky supports managed, VPC, on-prem and air-gapped deployment with the same OpenAI-compatible endpoints, so code does not change when the data plane does. Collections handle ingestion and retrieval with keyword, vector and hybrid modes, optional reranking and citations, and 30+ models sit behind one API for generation.

Estimate the deployment shape with the private deployment estimator, and run evaluations on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

ControlManaged with region lockVPC deploymentOn-prem or air-gapped
Data locationPlugsky cloud in a chosen regionYour cloud accountYour data centre or isolated network
Key custodyManaged encryption; BYOK on enterpriseCustomer-managed keysKMS or on-prem HSM
AuditPer-request logs with SIEM exportSame with network policySame, fully internal
App changesNoneNoneNone
FitsOperational and policy contentPatient-adjacent workloadsClassified or offline requirements

Frequently asked questions

Can Plugsky be used for patient data?

Plugsky offers VPC, on-prem and air-gapped deployment with encryption, scoped access and audit logs. Confirm the agreement, DPA and residency terms for your jurisdiction before loading patient data.

Does Plugsky train on the data?

Plugsky states that API data is not used to train models, and collections are encrypted at rest.

How do audit logs support compliance?

They record per-request model, tokens, latency, user and region, export to your SIEM, and can be retained according to your policy for access reviews and investigations.

Is RAG a medical device or diagnostic tool?

No. It is a retrieval and summarisation tool. Keep professional review and accountability in the workflow and be explicit about that boundary with users.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. Enterprise deployments are quoted separately.