Key facts
| Router model | plugsky-fusion escalates per document or question across tiers (live) |
| JSON mode | Live for clause extraction and metadata tagging |
| Models | 30+ models; cheap tiers suit bulk review and extraction |
| Long context | plugsky-longctx for synthesis across long agreements |
| Citations | Structured claim-to-clause mapping in JSON mode |
| Pricing | Flat monthly self-serve plans with no per-token charges on self-serve |
| Free tier | plugsky-micro and plugsky-lite on the free plan, no card required |
| Deployment | Region pinning plus VPC, on-prem and air-gapped options |
TL;DR
- Bulk extraction and tagging are cheap-tier work with schemas.
- Privilege, risk and final analysis stay on strong models.
- Every claim must cite a clause so reviewers can verify it.
- Long agreements need deliberate long-context use, not default stuffing.
- Measure cost per reviewed document plus correction rate.
How it works, step by step
- Segment legal work: intake, bulk review, extraction, analysis, drafting, final review.
- Route intake, tagging, translation and clause extraction to cheap tiers with JSON schemas.
- Add deterministic checks: clause presence, defined terms, dates and cross-references.
- Escalate risk assessment, privilege questions and interpretation to strong models.
- Use plugsky-longctx explicitly for agreements that require synthesis across sections.
- Require clause citations and human sign-off for anything a client will rely on.
- Track cost per reviewed document, reviewer correction rate and turnaround time.
Try it yourself
Open the LLM context window calculator →
Long documents, high stakes
Legal documents are long, formatted densely and consequential. The temptation is to send everything to the strongest model with the longest context, which is expensive and still no substitute for a lawyer. The practical approach separates bulk processing from judgement: extraction, tagging, translation and standardised checks are mechanical and verifiable; risk assessment, privilege analysis and negotiation advice are not.
That split governs routing. Cheap tiers with strict schemas handle the bulk, and strong models handle the parts where an error has real consequences — with citations attached so a human can verify every claim.
Tiering without losing accuracy
Extraction quality is measurable, which makes it safe to optimise. Define schemas for the fields you need — parties, dates, governing law, termination clauses, liability caps — and validate presence and consistency in code before anything reaches a review queue.
- Cheap tier: metadata extraction, defined-term indexing, translation of standard clauses.
- Mid tier: first-pass summaries and comparison against playbook positions.
- Strong tier: risk analysis, privilege considerations, conflicting-clause resolution.
- Long-context model: synthesis across lengthy agreements where passage selection is unreliable.
Review, audit and measurement
Citations are the control that makes any of this usable: each assertion links to a clause, so a reviewer verifies rather than re-reads. Keep the raw model output, prompt version and model identity for every reviewed document, because legal work may need to reconstruct how a conclusion was reached.
Measure cost per reviewed document alongside reviewer correction rate and turnaround. A cheaper tier that raises corrections is not saving money — reviewer time dominates. Start on the free plan with plugsky-micro and plugsky-lite for extraction pipelines, then use the 14-day full-access trial for strong and long-context models. Enterprise deployments add region pinning, VPC and air-gapped options for privileged material; plans are on the live pricing page.
Honest comparison
| Legal task | Routed legal assistant | Strong model for all | Bulk review manually |
|---|---|---|---|
| Clause extraction | Cheap tier with schemas | Frontier price per document | Paralegal hours |
| First-pass summary | Mid tier with citations | Native strength | Slow |
| Risk analysis | Strong tier plus human review | Native strength | Expert hours |
| Long agreements | Long-context model deliberately | Oversized prompts | Manual reading |
| Audit | Clause citations and model logs | Logged | Tacit |
Frequently asked questions
Can legal work use cheap models?
Bulk extraction, tagging, defined-term indexing and translation of standard clauses can, with strict schemas and deterministic checks. Analysis, privilege questions and anything a client relies on should use strong models with human review.
How do I keep extraction accurate?
Define schemas per document type, validate clause presence and cross-references in code, and escalate failures rather than retrying blindly on the same tier.
When should I use a long-context model?
When synthesis genuinely requires many sections of an agreement and passage selection is unreliable. Use it deliberately for those documents, not as the default for every prompt.
How do citations help?
Every claim links to a clause, so reviewers verify instead of re-reading and errors are traceable. Structured output makes citation checking systematic.
What must be retained for audit?
Raw model output, prompt version, model identity, deterministic checks applied and the reviewer decision. Legal conclusions may need reconstruction later.
Can privileged material stay on-prem?
Yes. Plugsky supports region pinning and VPC, on-prem or air-gapped deployments for sensitive and privileged material.
How do I measure savings?
Cost per reviewed document plus reviewer correction rate and turnaround time. If corrections rise, reviewer cost outweighs any model saving.
Is there a free way to test?
Yes. plugsky-micro and plugsky-lite are on the free plan with no card, and the 14-day full-access trial covers strong and long-context models.