Key facts
| API compatibility | OpenAI-compatible chat, embeddings and function calling |
| Use cases | Proposal and RFP drafting, engagement document Q&A, research synthesis, deliverable QA |
| Client isolation | Separate keys, projects or deployments per client as your engagement terms require |
| Data controls | API keys, RBAC, SSO, audit logs and region selection |
| Deployment | Plugsky cloud, your VPC, on-prem and air-gapped options |
| Models | 30+ models from free aliases to frontier reasoning |
| Pricing model | Flat monthly self-serve plans with unlimited fair-use usage |
| Free tier | plugsky-micro and plugsky-lite on the free plan, no card required |
TL;DR
- Cut proposal and memo drafting time while keeping partner review in the loop.
- Retrieve across engagement files with citations instead of trusting model memory.
- Separate keys or deployments per client when confidentiality terms demand it.
- VPC and on-prem options satisfy clients who restrict where data is processed.
- Prove value on one engagement before rolling out firm-wide.
How it works, step by step
- Choose one repeatable deliverable, such as a proposal first draft or research memo.
- Map the client confidentiality terms to a data boundary: which files may be sent, to which region, with what retention.
- Index approved engagement documents and answer with citations to the source passage.
- Build on the OpenAI-compatible endpoint so tooling can move between cloud and VPC without a rewrite.
- Add partner or manager review plus versioned prompts before anything reaches a client.
- Track draft time saved and revision count, then extend to the next deliverable type.
Try it yourself
Open the private LLM cost estimator →
Where an AI API pays off first
Firms sell expertise and documents, so document-heavy work is the natural start:
- Proposals and RFPs: assemble a first draft from past wins, bios and methodology sections.
- Engagement Q&A: ask questions across data rooms, transcripts and prior deliverables with citations.
- Research synthesis: summarise interviews, filings and market notes into a structured memo.
- Deliverable QA: check a draft against the approved checklist before it goes to a partner.
Every one of these keeps a named professional accountable for the final output.
Confidentiality, access and client isolation
Client confidentiality is the gating issue, not model quality. Practical controls include scoped API keys per engagement, RBAC and SSO for staff, audit logs that record who queried what, and region selection for processing. Where a client contract requires stronger separation, isolate at the deployment level with a dedicated VPC or on-prem instance rather than relying on prompt hygiene. Retention settings should be set before the first client document is sent, and your DPA and vendor-risk checklist should reflect whatever you configure.
Deployment without re-platforming
Because Plugsky uses the OpenAI-compatible interface, a prototype built on the free plan moves to a paid self-serve plan, a private VPC or an on-prem install with the same SDK code, prompts and tests. That portability matters in professional services, where different clients sit under different contractual constraints and the firm does not want two separate AI stacks. It also makes air-gapped delivery viable for public-sector or defence engagements that prohibit outbound traffic.
Measuring the return
Track three numbers per deliverable: time from kickoff to first reviewable draft, partner revision count, and hours spent searching for prior work. Prototype on plugsky-micro or plugsky-lite at no cost, then use the 14-day full-access trial to test frontier models on a real deliverable. Self-serve pricing is flat monthly with fair-use usage, which makes the cost per engagement predictable enough to quote internally. Expand to new deliverable types only after the first shows consistent review savings.
Honest comparison
| Capability | Plugsky | Typical per-token API | Building in-house |
|---|---|---|---|
| API compatibility | OpenAI-compatible chat, embeddings and tools | Usually compatible | Full rewrite |
| Client isolation | Separate keys, projects or dedicated deployments | Shared tenancy by default | You design tenancy |
| Deployment | Cloud, VPC, on-prem and air-gapped | Mostly cloud-only | You operate GPUs and serving |
| Pricing | Flat monthly self-serve, fair-use usage | Per-token, harder to forecast | GPU plus operations cost |
| Model choice | 30+ models behind one API | Varies by provider | You host every model |
Frequently asked questions
Can we keep our existing OpenAI SDK code?
Yes. Plugsky exposes an OpenAI-compatible API, so migration is a base URL and model-name change in code you already maintain.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with no credit card, which is enough to test document Q&A and drafting.
How do we keep one client's data away from another's?
Use separate API keys, projects or, for stricter contracts, a dedicated VPC or on-prem deployment. Set retention per engagement before sending documents.
Can we deploy where the client requires?
Yes. Plugsky supports cloud, VPC, on-prem and air-gapped deployment, so you can match the client's constraints without changing application code.
How does pricing work?
Self-serve plans are flat monthly with unlimited fair-use usage. See the live pricing page for current plans and enterprise options.
Which endpoints are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon.
How do we measure success?
Track time to first reviewable draft, partner revision count and search hours per deliverable. Those metrics map directly to billable and non-billable time.