Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (change the base URL) |
| Models | 30+ models from free to frontier tiers behind one API |
| Agent primitives | Function calling, JSON mode and streaming are live |
| Retrieval | Embeddings and RAG over your own corpus |
| Deployment | Plugsky cloud, VPC, on-prem or air-gapped |
| Pricing | Flat monthly self-serve plans with fair-use usage; see the live pricing page |
| Residency | Region pinning plus sovereign deployment tiers |
| Access control | Per-department keys, RBAC and per-request audit logs |
TL;DR
- Keep your OpenAI SDK — change the base URL and model name.
- 30+ models behind one API, from free chat models to frontier reasoning.
- Deployment options from hosted cloud to VPC, on-prem and air-gapped.
- Start with internal policy Q&A before any citizen-facing workflow.
- Keep decision authority with staff — agents draft, retrieve and route.
How it works, step by step
- Define the job, the permitted data sources and where a human must approve.
- Map the case lifecycle and mark every point that needs a human decision.
- Create a Plugsky account and generate an API key (free plan, no card required).
- Point your OpenAI SDK at the Plugsky base URL and map your model names.
- Index the approved corpus with embeddings and keep retrieval role-scoped.
- Pilot one queue and sample officer-reviewed answers weekly.
- Measure quality on your own samples, then scale with usage monitoring.
Try it yourself
Open the sovereign AI readiness score →
Where AI agents pay off in government
Government teams do not lack ideas for agents; they lack a safe path from demo to production. The pattern below targets repetitive, document-heavy work where a human can check the output, which is where agents earn their place first. Treat the agent as a new team member with a narrow brief, explicit permissions and a probation period, and rollout becomes an operations exercise rather than a leap of faith.
- Citizen intake triage — classify requests, detect duplicates and route them to the right queue
- Case research — summarise long case histories and cite the source documents
- Policy Q&A — let staff query internal policy with role-based access
- Drafting — produce first-draft letters, briefs and reports for officer review
A reference architecture for government agents
A triage agent classifies each request and calls your case-management API through typed tools; a research agent retrieves policy passages, drafts an answer and attaches citations. All write actions sit behind an approval gate, so the agent prepares and the officer commits.
- API gateway with per-department keys and rate limits
- Vector store over policy, procedure and case corpora inside the approved boundary
- Tool endpoints into case management, CRM and knowledge bases
- Policy layer for redaction, allow-lists, approval gates and audit logging
Data governance and human oversight
Public-sector workloads combine resident data with procurement and records obligations. Keep the model endpoint and vector store inside the boundary your agency has approved, give each department its own key, and treat every agent output as draft material until a named officer signs it off.
- Per-team API keys with RBAC, rotation and revocation
- Request and admin audit logs exportable to a SIEM
- Retention and deletion rules per record class
- Human sign-off before any citizen-facing or financial action
From pilot to production
Pilot an internal policy assistant on a narrow, non-sensitive corpus, then measure answer quality with officer-reviewed samples before touching citizen workflows. Expand queue by queue, keeping decision authority with staff.
Keep the rollout reversible: run the agent in shadow mode alongside the current process, compare outputs on your own samples, and move it into the workflow only when the evidence holds. Document what you measured so expanding to the next team is a decision, not a hope.
Honest comparison
| Capability | Plugsky | Typical cloud AI API | Building in-house |
|---|---|---|---|
| API compatibility | Drop-in base URL change | Usually compatible | Full rewrite |
| Model access | 30+ models behind one API | Vendor's own catalogue | You host each model |
| Pricing | Flat monthly self-serve plans; see live pricing | Often per-token | GPU + ops cost |
| Deployment | Cloud, VPC, on-prem or air-gapped | Usually vendor cloud regions | You own the stack |
| System of record | Stays your case platform; agents call it through tools | Unchanged | You build the tools either way |
| Citizen-facing safeguards | Approval gates and audit logs you configure | Varies by provider | You build them |
Frequently asked questions
Do we have to rewrite our application?
No. The chat completions API is OpenAI-compatible, so you change the base URL and model name and keep your existing SDK.
Is there a free plan?
Yes — the free plan includes two free AI models, plugsky-micro and plugsky-lite, with no credit card required.
How is pricing structured?
Self-serve plans are flat monthly with fair-use usage and no per-token charges; see the live pricing page for current plans.
Which endpoints are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses endpoints are coming soon — check the docs before planning around them.
Can an agent decide eligibility or benefits?
No. Use agents for retrieval, drafting and routing; formal decisions should stay with authorised staff under your legal and policy framework.
Can this run inside a government or national cloud?
Plugsky supports VPC, on-prem and air-gapped deployment so the model endpoint and data can stay inside an environment you control.
How do we audit agent activity?
Every request carries a key identity and produces logs you can export to your SIEM; pair those with approval records from your case system.