Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (change the base URL) |
| Models | 30+ models from free to frontier tiers behind one API |
| Agent primitives | Function calling, JSON mode and streaming are live |
| Retrieval | Embeddings and RAG over your own corpus |
| Deployment | Plugsky cloud, VPC, on-prem or air-gapped |
| Pricing | Flat monthly self-serve plans with fair-use usage; see the live pricing page |
| Maintenance triage | Classification and ticket creation through function calling |
| Channel support | Web and messaging front ends you control |
TL;DR
- Keep your OpenAI SDK — change the base URL and model name.
- 30+ models behind one API, from free chat models to frontier reasoning.
- Deployment options from hosted cloud to VPC, on-prem and air-gapped.
- One property or shift is enough to validate the pattern.
- Keep compensation and refunds with staff.
How it works, step by step
- Define the job, the permitted data sources and where a human must approve.
- Pick one shift and instrument the escalation rate.
- Create a Plugsky account and generate an API key (free plan, no card required).
- Point your OpenAI SDK at the Plugsky base URL and map your model names.
- Index the approved corpus with embeddings and keep retrieval role-scoped.
- Curate local and property content before enabling recommendations.
- Measure quality on your own samples, then scale with usage monitoring.
Try it yourself
Open the LLM token calculator →
Where AI agents pay off in hotels
Hotels teams do not lack ideas for agents; they lack a safe path from demo to production. The pattern below targets repetitive, document-heavy work where a human can check the output, which is where agents earn their place first. Treat the agent as a new team member with a narrow brief, explicit permissions and a probation period, and rollout becomes an operations exercise rather than a leap of faith.
- Front desk — answer room, facility and policy questions instantly
- Concierge — recommend dining, transport and local services from curated content
- Maintenance triage — classify guest reports and open tickets in the right queue
- Post-stay — draft follow-ups and review replies for manager approval
A reference architecture for hotels agents
A guest agent answers from curated hotel content and calls the PMS for stay context; a maintenance agent turns free-text reports into structured tickets and routes them by severity. Staff keep control of compensation and exceptions.
- Web and messaging channels with guest consent
- Retrieval over hotel services, policies and local guides
- Tools into PMS, housekeeping and ticketing
- Escalation rules with per-shift review
Data governance and human oversight
Keep guest records and payment context out of agent context unless truly needed. Scope each agent to its property, log interactions, and forbid autonomous refunds or room changes.
- Per-property scoping and keys
- Read-first access to guest data
- No pricing or refund actions without approval
- Interaction logs with retention limits
From pilot to production
Pilot on one floor or one shift with common questions plus maintenance triage. Review transcripts and escalation rates before enabling concierge recommendations.
Keep the rollout reversible: run the agent in shadow mode alongside the current process, compare outputs on your own samples, and move it into the workflow only when the evidence holds. Document what you measured so expanding to the next team is a decision, not a hope.
Honest comparison
| Capability | Plugsky | Typical cloud AI API | Building in-house |
|---|---|---|---|
| API compatibility | Drop-in base URL change | Usually compatible | Full rewrite |
| Model access | 30+ models behind one API | Vendor's own catalogue | You host each model |
| Pricing | Flat monthly self-serve plans; see live pricing | Often per-token | GPU + ops cost |
| Deployment | Cloud, VPC, on-prem or air-gapped | Usually vendor cloud regions | You own the stack |
| Guest context | Read-first PMS access through tools | Varies | Your integration work |
| Duty of care | Escalation rules you configure | Varies | You build it |
Frequently asked questions
Do we have to rewrite our application?
No. The chat completions API is OpenAI-compatible, so you change the base URL and model name and keep your existing SDK.
Is there a free plan?
Yes — the free plan includes two free AI models, plugsky-micro and plugsky-lite, with no credit card required.
How is pricing structured?
Self-serve plans are flat monthly with fair-use usage and no per-token charges; see the live pricing page for current plans.
Which endpoints are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses endpoints are coming soon — check the docs before planning around them.
Will it reduce front-desk workload?
It handles repeat questions and routes maintenance reports, which frees staff for in-person service. Measure escalation rate and transcript quality in a pilot before assuming an outcome.
Can it recommend local businesses?
Yes, if you curate the content it retrieves — recommendations are only as good as the source collection you approve.
Does it integrate with our PMS?
Through tool calls you define. Plugsky exposes function calling; your team wires the PMS endpoints and keeps them authoritative.