Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (change the base URL) |
| Models | 30+ models from free to frontier tiers behind one API |
| Agent primitives | Function calling, JSON mode and streaming are live |
| Retrieval | Embeddings and RAG over your own corpus |
| Deployment | Plugsky cloud, VPC, on-prem or air-gapped |
| Pricing | Flat monthly self-serve plans with fair-use usage; see the live pricing page |
| Tenant isolation | Per-client keys plus separated retrieval collections |
| Per-tenant audit | Request logs you can export for client reviews |
TL;DR
- Keep your OpenAI SDK — change the base URL and model name.
- 30+ models behind one API, from free chat models to frontier reasoning.
- Deployment options from hosted cloud to VPC, on-prem and air-gapped.
- Treat tenant isolation as a hard requirement, not a feature.
- Package the capability once isolation is provable.
How it works, step by step
- Define the job, the permitted data sources and where a human must approve.
- Prove per-tenant isolation on your own desk first.
- Create a Plugsky account and generate an API key (free plan, no card required).
- Point your OpenAI SDK at the Plugsky base URL and map your model names.
- Index the approved corpus with embeddings and keep retrieval role-scoped.
- Define per-client log exports before packaging the service.
- Measure quality on your own samples, then scale with usage monitoring.
Try it yourself
Where AI agents pay off in MSPs
MSPs teams do not lack ideas for agents; they lack a safe path from demo to production. The pattern below targets repetitive, document-heavy work where a human can check the output, which is where agents earn their place first. Treat the agent as a new team member with a narrow brief, explicit permissions and a probation period, and rollout becomes an operations exercise rather than a leap of faith.
- Ticket triage — classify and route tickets per client with SLA awareness
- Alert summaries — turn monitoring noise into prioritised plain-language reports
- Client updates — draft status and incident communications for engineer review
- Onboarding — assemble runbooks and checklists from your standard playbooks
A reference architecture for MSPs agents
A triage agent uses a client-scoped key to classify tickets and enrich them with asset data, while a reporting agent summarises alerts per tenant. Engineers approve changes and client communications.
- Per-tenant keys and separated retrieval collections
- Tools into RMM, PSA, documentation and ticketing
- SLA-aware routing rules
- Tenant audit logs for client reporting
Data governance and human oversight
Isolation is the product. Keep each client's content in its own collection and key, log access per tenant, and be able to show a client exactly which requests touched their data.
- Tenant-separated keys and collections
- Per-tenant audit exports
- Approval before changes or client messages
- Retention configured per contract
From pilot to production
Pilot on your own service desk, then a single client under strict scoping. Package it as a service only after isolation and reporting are demonstrably reliable.
Keep the rollout reversible: run the agent in shadow mode alongside the current process, compare outputs on your own samples, and move it into the workflow only when the evidence holds. Document what you measured so expanding to the next team is a decision, not a hope.
Honest comparison
| Capability | Plugsky | Typical cloud AI API | Building in-house |
|---|---|---|---|
| API compatibility | Drop-in base URL change | Usually compatible | Full rewrite |
| Model access | 30+ models behind one API | Vendor's own catalogue | You host each model |
| Pricing | Flat monthly self-serve plans; see live pricing | Often per-token | GPU + ops cost |
| Deployment | Cloud, VPC, on-prem or air-gapped | Usually vendor cloud regions | You own the stack |
| Multi-tenant model | Keys plus collections per client | Varies | You build it |
| White-label fit | OpenAI-compatible, integrates behind your product | Varies | You own the UX |
Frequently asked questions
Do we have to rewrite our application?
No. The chat completions API is OpenAI-compatible, so you change the base URL and model name and keep your existing SDK.
Is there a free plan?
Yes — the free plan includes two free AI models, plugsky-micro and plugsky-lite, with no credit card required.
How is pricing structured?
Self-serve plans are flat monthly with fair-use usage and no per-token charges; see the live pricing page for current plans.
Which endpoints are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses endpoints are coming soon — check the docs before planning around them.
Can we white-label this for clients?
The API is OpenAI-compatible and integrates behind your own product, so you control the experience; specific commercial terms are defined in your agreement.
How do we prove tenant isolation?
Use dedicated keys and collections per client and export their audit logs; that gives you concrete evidence for security reviews.
Can agents run remediation scripts?
Keep execution in your RMM with approval; agents can retrieve the runbook and prepare the action.