Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (change the base URL) |
| Models | 30+ models from free to frontier tiers behind one API |
| Agent primitives | Function calling, JSON mode and streaming are live |
| Retrieval | Embeddings and RAG over your own corpus |
| Deployment | Plugsky cloud, VPC, on-prem or air-gapped |
| Pricing | Flat monthly self-serve plans with fair-use usage; see the live pricing page |
| Human-in-the-loop | Clinician review before clinical use is the recommended pattern |
| Scoped access | Role-based keys and per-collection retrieval controls |
TL;DR
- Keep your OpenAI SDK — change the base URL and model name.
- 30+ models behind one API, from free chat models to frontier reasoning.
- Deployment options from hosted cloud to VPC, on-prem and air-gapped.
- Start with administrative workflows before clinical ones.
- Keep PHI inside the boundary and validate your obligations yourself.
How it works, step by step
- Define the job, the permitted data sources and where a human must approve.
- Pick administrative workflows first and define the clinician review point.
- Create a Plugsky account and generate an API key (free plan, no card required).
- Point your OpenAI SDK at the Plugsky base URL and map your model names.
- Index the approved corpus with embeddings and keep retrieval role-scoped.
- Validate the deployment tier against your privacy obligations before pilots.
- Measure quality on your own samples, then scale with usage monitoring.
Try it yourself
Open the private LLM cost estimator →
Where AI agents pay off in healthcare
Healthcare teams do not lack ideas for agents; they lack a safe path from demo to production. The pattern below targets repetitive, document-heavy work where a human can check the output, which is where agents earn their place first. Treat the agent as a new team member with a narrow brief, explicit permissions and a probation period, and rollout becomes an operations exercise rather than a leap of faith.
- Documentation support — draft visit notes and summaries for clinician review
- Patient intake — collect structured history and route to the right service
- Prior authorisation — assemble payer packets from approved sources
- Knowledge Q&A — answer policy, formulary and protocol questions with citations
A reference architecture for healthcare agents
An intake agent collects structured answers and checks them against your scheduling or records system through tools; a documentation agent retrieves approved templates and drafts text for the clinician. No agent writes to a clinical record without a human accepting the draft.
- Gateway with per-service keys and strict rate limits
- Retrieval over approved protocols, policies and formularies
- Read-first tools into scheduling, billing and content systems
- Redaction and logging layer inside the data boundary
Data governance and human oversight
Treat health information as data that never leaves the approved environment. Your obligations under health-privacy rules remain yours to validate; choose the deployment tier that satisfies them, restrict retrieval by role, and log every access.
- Per-role keys and scoped retrieval
- De-identification or redaction before external calls, where applicable
- Audit logs with caller identity and timestamps
- Clinician sign-off before clinical use
From pilot to production
Begin with non-clinical administration such as intake forms and prior-authorisation checklists. Measure accuracy against clinician-reviewed samples, then extend to documentation support.
Keep the rollout reversible: run the agent in shadow mode alongside the current process, compare outputs on your own samples, and move it into the workflow only when the evidence holds. Document what you measured so expanding to the next team is a decision, not a hope.
Honest comparison
| Capability | Plugsky | Typical cloud AI API | Building in-house |
|---|---|---|---|
| API compatibility | Drop-in base URL change | Usually compatible | Full rewrite |
| Model access | 30+ models behind one API | Vendor's own catalogue | You host each model |
| Pricing | Flat monthly self-serve plans; see live pricing | Often per-token | GPU + ops cost |
| Deployment | Cloud, VPC, on-prem or air-gapped | Usually vendor cloud regions | You own the stack |
| Clinical write-back | Agents draft; clinicians commit | Varies by vendor | You build the guardrails |
| Retrieval scope | Role-scoped collections | Varies | Your responsibility |
Frequently asked questions
Do we have to rewrite our application?
No. The chat completions API is OpenAI-compatible, so you change the base URL and model name and keep your existing SDK.
Is there a free plan?
Yes — the free plan includes two free AI models, plugsky-micro and plugsky-lite, with no credit card required.
How is pricing structured?
Self-serve plans are flat monthly with fair-use usage and no per-token charges; see the live pricing page for current plans.
Which endpoints are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses endpoints are coming soon — check the docs before planning around them.
Can agents handle protected health information?
They can run inside a deployment that keeps data in your boundary — hosted VPC, on-prem or air-gapped. Compliance with health-privacy rules remains your responsibility, so verify the tier and controls against your obligations.
Do agents give clinical advice?
No. Design them for retrieval, drafting and administrative workflow, with qualified clinicians reviewing any output that influences care.
Can we restrict what each agent can read?
Yes. Scope API keys and retrieval collections per role or service so an agent only reaches the corpora and tools it needs.