Key facts
| Citations | Every query returns ranked chunks with source attribution |
| Isolation | Per-collection encryption and scoped keys for privilege walls |
| Formats | PDF, DOCX, TXT, MD and HTML ingestion |
| Confidentiality | VPC, on-prem and air-gapped deployment options |
| Audit logs | Per-request model, tokens, latency, user and region with SIEM export |
| Data use | API data is not used to train models |
| Retrieval | Keyword, vector and hybrid search with optional reranking |
| Product status | Live |
TL;DR
- Make the collection the confidentiality wall: one per matter or client.
- Always cite the passage; a legal answer without a source cannot be checked.
- Version documents so an answer ties to the correct draft.
- Agree retention and deletion for closed matters before ingestion.
- Deploy inside your perimeter when privilege or client terms require it.
How it works, step by step
- Map matters, clients and confidentiality walls to collection boundaries.
- Define who can query each collection and issue scoped keys accordingly.
- Ingest a closed set first, such as templates or settled matters, as a pilot.
- Check chunking against contract structure: clauses, schedules and definitions.
- Build a question set with expected sources and score retrieval before drafting.
- Configure retention, deletion and audit export with risk and IT.
- Expand to live matters only after retrieval quality and access controls are proven.
Try it yourself
Why citations are non-negotiable
Legal work is verified reading. A summary that cannot be traced to a clause has little value, because the risk of acting on an inaccurate paraphrase is high. Retrieval-grounded answers fit the profession better than free-form generation: the system returns the relevant passage with its source, and the lawyer reads the evidence before relying on it.
Citations also make review efficient. Instead of checking an entire answer against an entire contract, the reviewer checks a specific passage against a specific page, which is the same workflow they already use with a colleague's research note.
Privilege walls and access design
The collection is the practical unit of confidentiality. Create one per matter or client, scope keys to the team working on it, and never mix adversarial matters in a single index. Per-collection encryption at rest and scoped API keys keep retrieval attributable, while enterprise RBAC and SSO tie human access to the identity provider.
Audit logs then record who queried what and when, with export to your SIEM. For engagements that require it, deploy in a VPC, on-prem or air-gapped plane so documents and queries never leave the firm's perimeter.
Document structure and versioning
Contracts and filings have deliberate structure: recitals, definitions, operative clauses, schedules and annexes. Chunking that respects those boundaries keeps definitions with the clauses that use them, and citations that point to a page let a reviewer jump straight to the text. Keep the document version in the metadata so an answer can be tied to the draft in force.
Hybrid retrieval helps when searching for defined terms, party names and clause references, which are exact-match queries more than semantic ones. Optional reranking then orders the candidates for precision.
Piloting and production on Plugsky
Start with a non-privileged corpus such as a template library or public precedents to tune chunking and evaluation. Plugsky collections accept PDF, DOCX, TXT, MD and HTML, handle chunking and indexing automatically, and return ranked chunks with source attribution. Any of 30+ models can compose answers behind the same OpenAI-compatible API.
Test document questions with the Chat with PDF tool, and review the DPA and deployment terms before privileged content moves. Evaluate on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial; current plans are on the live pricing page.
Honest comparison
| Requirement | Plugsky RAG | Keyword document management | General chatbot |
|---|---|---|---|
| Source linking | Passage-level citations with page references | Document-level search | None |
| Privilege walls | Collections, scoped keys and encryption | Folder permissions | No isolation model |
| Version control | Metadata on chunks and stable ids | File version history | Unknown |
| Audit | Per-request logs with SIEM export | Access logs | Chat history only |
| Deployment | Managed, VPC, on-prem and air-gapped | On-prem by default | Vendor cloud |
Frequently asked questions
Can Plugsky replace legal research databases?
No. It retrieves from the corpus you provide, such as your contracts and precedents. Public case law and licensed databases are separate sources you would ingest if permitted.
How do we protect privilege?
Keep a collection per matter or client, scope keys to the team, deploy inside your perimeter where required, and review audit logs for access.
What happens when a matter closes?
Agree deletion workflows up front: remove the collection and documents, and confirm that logs and backups follow your retention policy.
How should contracts be chunked?
Follow the document structure: clauses, definitions and schedules. Structure-aware chunks keep defined terms with the clauses that use them.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.