Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (change the base URL) |
| Models | 30+ models from free to frontier tiers behind one API |
| Agent primitives | Function calling, JSON mode and streaming are live |
| Retrieval | Embeddings and RAG over your own corpus |
| Deployment | Plugsky cloud, VPC, on-prem or air-gapped |
| Pricing | Flat monthly self-serve plans with fair-use usage; see the live pricing page |
| Multi-tenant scoping | Per-client API keys and separated retrieval collections |
| Change control | Approval gate before any change or client message |
TL;DR
- Keep your OpenAI SDK — change the base URL and model name.
- 30+ models behind one API, from free chat models to frontier reasoning.
- Deployment options from hosted cloud to VPC, on-prem and air-gapped.
- Isolate each client before automating anything.
- Draft and retrieve with agents; execute with humans or automation.
How it works, step by step
- Define the job, the permitted data sources and where a human must approve.
- Start with internal desk tickets before client-facing work.
- Create a Plugsky account and generate an API key (free plan, no card required).
- Point your OpenAI SDK at the Plugsky base URL and map your model names.
- Index the approved corpus with embeddings and keep retrieval role-scoped.
- Define per-client isolation and log retention up front.
- Measure quality on your own samples, then scale with usage monitoring.
Try it yourself
Where AI agents pay off in IT services
IT Services teams do not lack ideas for agents; they lack a safe path from demo to production. The pattern below targets repetitive, document-heavy work where a human can check the output, which is where agents earn their place first. Treat the agent as a new team member with a narrow brief, explicit permissions and a probation period, and rollout becomes an operations exercise rather than a leap of faith.
- Ticket triage — categorise, prioritise and route incidents to the right queue
- Runbook assistance — retrieve the right procedure and draft the steps
- Asset Q&A — answer licence, configuration and inventory questions
- Change summaries — draft change records and post-incident notes for review
A reference architecture for IT services agents
A triage agent classifies tickets and enriches them with asset context from your CMDB through tools, while a knowledge agent retrieves runbooks and drafts next steps for the engineer. Execution stays with the engineer or your automation platform.
- Per-client keys and data scoping
- Retrieval over runbooks, KB articles and contracts
- Read-first tools into ITSM, monitoring and CMDB
- Approval gate before any change or client update
Data governance and human oversight
You carry your clients' data, so isolation matters more than features. Scope keys and retrieval per client, log access per tenant, and keep client content out of shared collections.
- Per-tenant key scoping and rotation
- Client-segregated retrieval collections
- Audit logs per tenant for your client reports
- Human approval before changes or billable actions
From pilot to production
Pilot on internal service-desk tickets first, then one client with strict scoping. Track first-contact resolution and misroute rate before offering the capability as a managed service.
Keep the rollout reversible: run the agent in shadow mode alongside the current process, compare outputs on your own samples, and move it into the workflow only when the evidence holds. Document what you measured so expanding to the next team is a decision, not a hope.
Honest comparison
| Capability | Plugsky | Typical cloud AI API | Building in-house |
|---|---|---|---|
| API compatibility | Drop-in base URL change | Usually compatible | Full rewrite |
| Model access | 30+ models behind one API | Vendor's own catalogue | You host each model |
| Pricing | Flat monthly self-serve plans; see live pricing | Often per-token | GPU + ops cost |
| Deployment | Cloud, VPC, on-prem or air-gapped | Usually vendor cloud regions | You own the stack |
| Multi-client isolation | Keys plus separated collections | Varies by provider | You build it |
| Execution | Agents draft; your automation runs | Varies | You build it |
Frequently asked questions
Do we have to rewrite our application?
No. The chat completions API is OpenAI-compatible, so you change the base URL and model name and keep your existing SDK.
Is there a free plan?
Yes — the free plan includes two free AI models, plugsky-micro and plugsky-lite, with no credit card required.
How is pricing structured?
Self-serve plans are flat monthly with fair-use usage and no per-token charges; see the live pricing page for current plans.
Which endpoints are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses endpoints are coming soon — check the docs before planning around them.
Can we offer this as part of our managed service?
Yes — the API is OpenAI-compatible and per-tenant keys let you isolate each client's data; commercial terms are yours to set with your clients.
Can agents execute remediation?
They can retrieve and draft; execution should go through your automation with approval. Keeping a human gate is the safer default for client environments.
How do we handle many clients in one account?
Use separate keys and collections per client, and log requests so you can prove isolation in reviews and audits.