Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (change the base URL) |
| Models | 30+ models from free to frontier tiers behind one API |
| Agent primitives | Function calling, JSON mode and streaming are live |
| Retrieval | Embeddings and RAG over your own corpus |
| Deployment | Plugsky cloud, VPC, on-prem or air-gapped |
| Pricing | Flat monthly self-serve plans with fair-use usage; see the live pricing page |
| Operational tools | Tool calls into scheduling, bed management and ERP systems |
| Data minimisation | Role-scoped access to the minimum data needed per task |
TL;DR
- Keep your OpenAI SDK — change the base URL and model name.
- 30+ models behind one API, from free chat models to frontier reasoning.
- Deployment options from hosted cloud to VPC, on-prem and air-gapped.
- Start with operational knowledge, not clinical workflows.
- Scope every agent to the minimum data its task needs.
How it works, step by step
- Define the job, the permitted data sources and where a human must approve.
- Choose operational workflows where mistakes are easy to catch.
- Create a Plugsky account and generate an API key (free plan, no card required).
- Point your OpenAI SDK at the Plugsky base URL and map your model names.
- Index the approved corpus with embeddings and keep retrieval role-scoped.
- Define the boundary: what the agent may read and never write.
- Measure quality on your own samples, then scale with usage monitoring.
Try it yourself
Open the private LLM cost estimator →
Where AI agents pay off in hospitals
Hospitals teams do not lack ideas for agents; they lack a safe path from demo to production. The pattern below targets repetitive, document-heavy work where a human can check the output, which is where agents earn their place first. Treat the agent as a new team member with a narrow brief, explicit permissions and a probation period, and rollout becomes an operations exercise rather than a leap of faith.
- Capacity Q&A — answer bed, theatre and equipment questions from live system data
- Discharge coordination — draft checklists and follow-up notes for staff
- Knowledge lookup — retrieve policies, protocols and equipment guidance with citations
- Patient navigation — answer administrative questions about appointments and directions
A reference architecture for hospitals agents
An operations agent queries bed management or scheduling through tools and returns a cited answer to the coordinator, while a knowledge agent answers protocol questions from an approved library. Patient-facing responses stay limited to administrative content and route through staff.
- Per-department keys with narrow scopes
- Retrieval over approved protocols and equipment manuals
- Read-only tools into bed management, scheduling and ERPs
- Full audit trail with named human sign-off points
Data governance and human oversight
Hospital data spans clinical, operational and personal categories. Keep processing inside the boundary, restrict retrieval by role, and never let an agent make or draft a clinical judgment without qualified review. Privacy obligations are yours to validate against the chosen tier.
- Role-scoped retrieval and tool permissions
- No clinical decision output without review
- Access logs with identity and purpose
- Retention rules per data class
From pilot to production
Start with operational questions and equipment knowledge, where errors are visible and low-risk. Extend to coordination workflows once accuracy and escalation behaviour are stable.
Keep the rollout reversible: run the agent in shadow mode alongside the current process, compare outputs on your own samples, and move it into the workflow only when the evidence holds. Document what you measured so expanding to the next team is a decision, not a hope.
Honest comparison
| Capability | Plugsky | Typical cloud AI API | Building in-house |
|---|---|---|---|
| API compatibility | Drop-in base URL change | Usually compatible | Full rewrite |
| Model access | 30+ models behind one API | Vendor's own catalogue | You host each model |
| Pricing | Flat monthly self-serve plans; see live pricing | Often per-token | GPU + ops cost |
| Deployment | Cloud, VPC, on-prem or air-gapped | Usually vendor cloud regions | You own the stack |
| Clinical decisions | Excluded by design; staff retain authority | Varies | You enforce it |
| On-prem fit | VPC, on-prem and air-gapped options | Often cloud-only | You own everything |
Frequently asked questions
Do we have to rewrite our application?
No. The chat completions API is OpenAI-compatible, so you change the base URL and model name and keep your existing SDK.
Is there a free plan?
Yes — the free plan includes two free AI models, plugsky-micro and plugsky-lite, with no credit card required.
How is pricing structured?
Self-serve plans are flat monthly with fair-use usage and no per-token charges; see the live pricing page for current plans.
Which endpoints are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses endpoints are coming soon — check the docs before planning around them.
Can agents access patient records?
Design them to work with the minimum necessary data: scope retrieval and tools so an agent sees only what its task requires, and keep processing inside the approved boundary.
Are agents making clinical decisions?
No. Use them for retrieval, drafting and coordination with qualified staff reviewing anything that touches care.
Can it run on-premises?
Yes — Plugsky supports VPC, on-prem and air-gapped deployment for environments that cannot use a public endpoint.