Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (change the base URL) |
| Models | 30+ models from free to frontier tiers behind one API |
| Agent primitives | Function calling, JSON mode and streaming are live |
| Retrieval | Embeddings and RAG over your own corpus |
| Deployment | Plugsky cloud, VPC, on-prem or air-gapped |
| Pricing | Flat monthly self-serve plans with fair-use usage; see the live pricing page |
| Offline deployment | On-prem and air-gapped options with periodic refresh |
| Document control | Versioned collections keep retrieval on controlled revisions |
TL;DR
- Keep your OpenAI SDK — change the base URL and model name.
- 30+ models behind one API, from free chat models to frontier reasoning.
- Deployment options from hosted cloud to VPC, on-prem and air-gapped.
- Pilot maintenance troubleshooting on a single line.
- Control systems keep authority; agents assist.
How it works, step by step
- Define the job, the permitted data sources and where a human must approve.
- Start with maintenance procedures on a single line.
- Create a Plugsky account and generate an API key (free plan, no card required).
- Point your OpenAI SDK at the Plugsky base URL and map your model names.
- Index the approved corpus with embeddings and keep retrieval role-scoped.
- Version document collections so answers track controlled revisions.
- Measure quality on your own samples, then scale with usage monitoring.
Try it yourself
Open the private LLM cost estimator →
Where AI agents pay off in manufacturing
Manufacturing teams do not lack ideas for agents; they lack a safe path from demo to production. The pattern below targets repetitive, document-heavy work where a human can check the output, which is where agents earn their place first. Treat the agent as a new team member with a narrow brief, explicit permissions and a probation period, and rollout becomes an operations exercise rather than a leap of faith.
- Maintenance support — retrieve fault procedures and prior work orders
- SOP Q&A — answer line and safety questions from controlled documents
- Quality incidents — summarise non-conformance reports and draft next actions
- Supplier documents — extract key terms and certificates for review
A reference architecture for manufacturing agents
A maintenance agent takes a fault description, retrieves the right procedure and pulls historic work orders through tools, while a quality agent drafts incident summaries for engineers. Setpoints, schedules and releases stay under human or control-system authority.
- Shift-friendly interfaces on top of the API
- Retrieval over manuals, SOPs and work-order history
- Read-first tools into CMMS, MES and ERP
- Approval gates before production changes
Data governance and human oversight
Plant data is operationally sensitive and sometimes export-controlled. Keep retrieval inside the site boundary where required, scope access by role, and log queries so safety-critical answers can be reviewed.
- Site-scoped keys and collections
- On-prem or air-gapped options for sensitive plants
- Audit logs for every query and tool call
- Human approval for production-affecting actions
From pilot to production
Pilot maintenance troubleshooting on one line, measuring time-to-procedure and answer accuracy with technicians. Expand to SOP Q&A and quality drafting once trust builds.
Keep the rollout reversible: run the agent in shadow mode alongside the current process, compare outputs on your own samples, and move it into the workflow only when the evidence holds. Document what you measured so expanding to the next team is a decision, not a hope.
Honest comparison
| Capability | Plugsky | Typical cloud AI API | Building in-house |
|---|---|---|---|
| API compatibility | Drop-in base URL change | Usually compatible | Full rewrite |
| Model access | 30+ models behind one API | Vendor's own catalogue | You host each model |
| Pricing | Flat monthly self-serve plans; see live pricing | Often per-token | GPU + ops cost |
| Deployment | Cloud, VPC, on-prem or air-gapped | Usually vendor cloud regions | You own the stack |
| Plant deployment | Cloud, on-prem or air-gapped | Often cloud-only | You own the stack |
| Production changes | Human/control-system authority | Varies | You enforce it |
Frequently asked questions
Do we have to rewrite our application?
No. The chat completions API is OpenAI-compatible, so you change the base URL and model name and keep your existing SDK.
Is there a free plan?
Yes — the free plan includes two free AI models, plugsky-micro and plugsky-lite, with no credit card required.
How is pricing structured?
Self-serve plans are flat monthly with fair-use usage and no per-token charges; see the live pricing page for current plans.
Which endpoints are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses endpoints are coming soon — check the docs before planning around them.
Can this run disconnected from the internet?
Yes — Plugsky offers on-prem and air-gapped deployment for sites that cannot use a public endpoint, with periodic model refresh.
Can agents change machine settings?
No. Keep write access with your control systems; agents retrieve, summarise and draft, with engineers approving actions.
How do we keep answers current after a procedure changes?
Re-index the changed documents on your own schedule and version collections so the agent always retrieves the controlled revision.