Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (change the base URL) |
| Models | 30+ models from free to frontier tiers behind one API |
| Agent primitives | Function calling, JSON mode and streaming are live |
| Retrieval | Embeddings and RAG over your own corpus |
| Deployment | Plugsky cloud, VPC, on-prem or air-gapped |
| Pricing | Flat monthly self-serve plans with fair-use usage; see the live pricing page |
| Matter scoping | Per-matter keys and retrieval collections |
| Confidentiality | Deployment tiers from VPC to air-gapped for controlled content |
TL;DR
- Keep your OpenAI SDK — change the base URL and model name.
- 30+ models behind one API, from free chat models to frontier reasoning.
- Deployment options from hosted cloud to VPC, on-prem and air-gapped.
- Scope retrieval per matter before any live work.
- Treat output as a draft that a lawyer reviews and owns.
How it works, step by step
- Define the job, the permitted data sources and where a human must approve.
- Pilot on closed matters and internal know-how.
- Create a Plugsky account and generate an API key (free plan, no card required).
- Point your OpenAI SDK at the Plugsky base URL and map your model names.
- Index the approved corpus with embeddings and keep retrieval role-scoped.
- Define matter scoping and review rules before live work.
- Measure quality on your own samples, then scale with usage monitoring.
Try it yourself
Open the AI data residency checklist →
Where AI agents pay off in law firms
Law Firms teams do not lack ideas for agents; they lack a safe path from demo to production. The pattern below targets repetitive, document-heavy work where a human can check the output, which is where agents earn their place first. Treat the agent as a new team member with a narrow brief, explicit permissions and a probation period, and rollout becomes an operations exercise rather than a leap of faith.
- Matter intake — capture conflicts-relevant facts and open the file consistently
- Research summaries — synthesise sources you supply, with citations
- Clause extraction — pull and compare clauses across a document set
- Review support — flag issues against a checklist for lawyer review
A reference architecture for law firms agents
A research agent retrieves from a matter-scoped collection and returns cited summaries, while a review agent compares clause text against your playbook and lists deviations for a lawyer. Nothing is filed or sent without lawyer sign-off.
- Matter-scoped keys and retrieval collections
- Read-first tools into DMS and practice management
- Playbook and checklist configuration
- Full audit logs for client and regulatory inquiries
Data governance and human oversight
Client confidentiality is the controlling constraint. Scope retrieval per matter, keep content inside the boundary your firm has approved, and treat every agent output as a junior associate's draft: useful, but reviewed and owned by a qualified lawyer.
- Per-matter access control
- Deployment tier validated against confidentiality obligations
- Audit logs with user identity
- Lawyer review before client delivery
From pilot to production
Pilot on internal know-how and closed matters, then move to live matters for research and clause work. Track citation quality and review time before expanding.
Keep the rollout reversible: run the agent in shadow mode alongside the current process, compare outputs on your own samples, and move it into the workflow only when the evidence holds. Document what you measured so expanding to the next team is a decision, not a hope.
Honest comparison
| Capability | Plugsky | Typical cloud AI API | Building in-house |
|---|---|---|---|
| API compatibility | Drop-in base URL change | Usually compatible | Full rewrite |
| Model access | 30+ models behind one API | Vendor's own catalogue | You host each model |
| Pricing | Flat monthly self-serve plans; see live pricing | Often per-token | GPU + ops cost |
| Deployment | Cloud, VPC, on-prem or air-gapped | Usually vendor cloud regions | You own the stack |
| Confidentiality | Boundary-controlled deployment plus matter scoping | Varies | You build it |
| Work product | Draft only; lawyer owns delivery | Varies | You enforce it |
Frequently asked questions
Do we have to rewrite our application?
No. The chat completions API is OpenAI-compatible, so you change the base URL and model name and keep your existing SDK.
Is there a free plan?
Yes — the free plan includes two free AI models, plugsky-micro and plugsky-lite, with no credit card required.
How is pricing structured?
Self-serve plans are flat monthly with fair-use usage and no per-token charges; see the live pricing page for current plans.
Which endpoints are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses endpoints are coming soon — check the docs before planning around them.
Is client data used to train models?
Do not assume behaviour you have not verified. Choose the deployment tier that matches your confidentiality obligations and confirm data-handling terms in writing.
Can agents draft client advice?
They can draft research summaries and document work product; a qualified lawyer must review, correct and own anything delivered to a client.
Does it work with our DMS?
Through tool calls you define — your team wires the DMS endpoints and keeps permissions authoritative.