Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (change the base URL) |
| Models | 30+ models from free to frontier tiers behind one API |
| Agent primitives | Function calling, JSON mode and streaming are live |
| Retrieval | Embeddings and RAG over your own corpus |
| Deployment | Plugsky cloud, VPC, on-prem or air-gapped |
| Pricing | Flat monthly self-serve plans with fair-use usage; see the live pricing page |
| Critical infrastructure | On-prem and air-gapped deployment with no control-path access |
| Operations control | Switching and dispatch authority stays with operators |
TL;DR
- Keep your OpenAI SDK — change the base URL and model name.
- 30+ models behind one API, from free chat models to frontier reasoning.
- Deployment options from hosted cloud to VPC, on-prem and air-gapped.
- Keep agents out of control paths entirely.
- Version safety procedures before any field use.
How it works, step by step
- Define the job, the permitted data sources and where a human must approve.
- Pilot field-procedure lookup and outage drafting.
- Create a Plugsky account and generate an API key (free plan, no card required).
- Point your OpenAI SDK at the Plugsky base URL and map your model names.
- Index the approved corpus with embeddings and keep retrieval role-scoped.
- Version controlled procedures before field use.
- Measure quality on your own samples, then scale with usage monitoring.
Try it yourself
Open the sovereign AI readiness score →
Where AI agents pay off in utilities
Utilities teams do not lack ideas for agents; they lack a safe path from demo to production. The pattern below targets repetitive, document-heavy work where a human can check the output, which is where agents earn their place first. Treat the agent as a new team member with a narrow brief, explicit permissions and a probation period, and rollout becomes an operations exercise rather than a leap of faith.
- Outage communications — draft accurate updates from incident information you supply
- Field knowledge — retrieve procedures and safety guidance for crews
- Asset history — answer maintenance and failure questions for a given asset
- Regulatory research — summarise filings and internal policy with citations
A reference architecture for utilities agents
A communications agent drafts outage updates from approved inputs, while a field agent retrieves the right procedure and asset history through tools. Grid control and switching decisions remain fully with operations.
- Site and role scoped collections
- Read-first tools into work management and asset systems
- Approval gates for customer communications
- Offline-capable deployment where required
Data governance and human oversight
Critical-infrastructure answers must be dependable and traceable. Version the procedures that feed retrieval, keep operational data in the boundary, and never let an agent sit in a control path.
- Controlled-document versioning
- Role-scoped access
- Audit logs for queries and drafts
- Operations retain switching and dispatch authority
From pilot to production
Pilot field-procedure lookup and outage drafting, then asset history. Keep control-room and switching workflows out of scope.
Keep the rollout reversible: run the agent in shadow mode alongside the current process, compare outputs on your own samples, and move it into the workflow only when the evidence holds. Document what you measured so expanding to the next team is a decision, not a hope.
Honest comparison
| Capability | Plugsky | Typical cloud AI API | Building in-house |
|---|---|---|---|
| API compatibility | Drop-in base URL change | Usually compatible | Full rewrite |
| Model access | 30+ models behind one API | Vendor's own catalogue | You host each model |
| Pricing | Flat monthly self-serve plans; see live pricing | Often per-token | GPU + ops cost |
| Deployment | Cloud, VPC, on-prem or air-gapped | Usually vendor cloud regions | You own the stack |
| Operational control | Excluded; operators retain authority | Varies | You enforce it |
| Offline operation | On-prem and air-gapped options | Often cloud-only | You own the stack |
Frequently asked questions
Do we have to rewrite our application?
No. The chat completions API is OpenAI-compatible, so you change the base URL and model name and keep your existing SDK.
Is there a free plan?
Yes — the free plan includes two free AI models, plugsky-micro and plugsky-lite, with no credit card required.
How is pricing structured?
Self-serve plans are flat monthly with fair-use usage and no per-token charges; see the live pricing page for current plans.
Which endpoints are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses endpoints are coming soon — check the docs before planning around them.
Can agents control grid equipment?
No. They retrieve, summarise and draft; operational control stays with your systems and qualified operators.
Can it run without internet access?
Yes — on-prem and air-gapped deployment supports sites that cannot depend on a public endpoint.
How do we keep safety procedures current?
Treat retrieval as a controlled-document pipeline: publish revisions on your change schedule and retire superseded versions.