Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (change the base URL) |
| Models | 30+ models from free to frontier tiers behind one API |
| Agent primitives | Function calling, JSON mode and streaming are live |
| Retrieval | Embeddings and RAG over your own corpus |
| Deployment | Plugsky cloud, VPC, on-prem or air-gapped |
| Pricing | Flat monthly self-serve plans with fair-use usage; see the live pricing page |
| Knowledge reuse | Retrieval over engagement libraries with scoped access |
| Review flow | Consultant approval before client delivery |
TL;DR
- Keep your OpenAI SDK — change the base URL and model name.
- 30+ models behind one API, from free chat models to frontier reasoning.
- Deployment options from hosted cloud to VPC, on-prem and air-gapped.
- Scope retrieval per engagement and run conflict checks.
- Consultants own everything that reaches clients.
How it works, step by step
- Define the job, the permitted data sources and where a human must approve.
- Curate an engagement library and pilot proposal drafting.
- Create a Plugsky account and generate an API key (free plan, no card required).
- Point your OpenAI SDK at the Plugsky base URL and map your model names.
- Index the approved corpus with embeddings and keep retrieval role-scoped.
- Run conflict checks before enabling knowledge reuse.
- Measure quality on your own samples, then scale with usage monitoring.
Try it yourself
Where AI agents pay off in professional services
Professional Services teams do not lack ideas for agents; they lack a safe path from demo to production. The pattern below targets repetitive, document-heavy work where a human can check the output, which is where agents earn their place first. Treat the agent as a new team member with a narrow brief, explicit permissions and a probation period, and rollout becomes an operations exercise rather than a leap of faith.
- Proposal drafting — assemble credentials, case studies and approach from prior work
- Engagement research — synthesise approved sources with citations
- Knowledge reuse — find prior deliverables and lessons across engagements
- Deliverable QA — check drafts against your methodology and templates
A reference architecture for professional services agents
A pursuit agent retrieves relevant credentials and drafts a proposal skeleton, while a research agent returns cited summaries. Consultants shape, verify and own everything that reaches a client.
- Engagement-scoped retrieval collections
- Tools into CRM, DMS and project systems
- Methodology and template controls
- Review before client delivery
Data governance and human oversight
Client confidentiality and independence rules shape what agents may see. Scope retrieval per engagement, check conflicts before reuse, and keep client content inside your boundary.
- Engagement-level access control
- Conflict checks before knowledge reuse
- Audit logs for generated work product
- Consultant review before delivery
From pilot to production
Pilot proposal drafting from a clean, curated knowledge base, then expand to research and QA. Track edit effort and reuse quality rather than raw output volume.
Keep the rollout reversible: run the agent in shadow mode alongside the current process, compare outputs on your own samples, and move it into the workflow only when the evidence holds. Document what you measured so expanding to the next team is a decision, not a hope.
Honest comparison
| Capability | Plugsky | Typical cloud AI API | Building in-house |
|---|---|---|---|
| API compatibility | Drop-in base URL change | Usually compatible | Full rewrite |
| Model access | 30+ models behind one API | Vendor's own catalogue | You host each model |
| Pricing | Flat monthly self-serve plans; see live pricing | Often per-token | GPU + ops cost |
| Deployment | Cloud, VPC, on-prem or air-gapped | Usually vendor cloud regions | You own the stack |
| Client confidentiality | Engagement-scoped collections | Varies | You configure |
| Independence | Conflict checks stay in your process | Varies | You enforce it |
Frequently asked questions
Do we have to rewrite our application?
No. The chat completions API is OpenAI-compatible, so you change the base URL and model name and keep your existing SDK.
Is there a free plan?
Yes — the free plan includes two free AI models, plugsky-micro and plugsky-lite, with no credit card required.
How is pricing structured?
Self-serve plans are flat monthly with fair-use usage and no per-token charges; see the live pricing page for current plans.
Which endpoints are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses endpoints are coming soon — check the docs before planning around them.
Can it reuse work from other clients?
Only within the permissions and conflict rules you enforce. Index each engagement separately and apply your independence checks before surfacing content.
Will proposals sound like us?
Ground drafting in your methodology and genuine prior work, and keep a consultant edit pass; that preserves voice and accuracy.
How do we measure value?
Compare edit time and research hours on a pilot against your baseline, then decide whether to expand.