Key facts
| Partitioning | Collections act as retrieval and encryption boundaries |
| Keys | Scoped API keys per application or team; RBAC and SSO on enterprise |
| Audit logs | Per-request model, tokens, latency, user and region with SIEM export |
| Data handling | Per-collection encryption at rest; API data is not used to train models |
| Retrieval | Keyword, vector and hybrid search with citations for review |
| Deployment | Managed, VPC, on-prem and air-gapped options |
| Retention | Configurable retention with export up to the published maximum |
| Product status | Live |
TL;DR
- Collection boundaries should mirror the way your organisation grants access.
- Scope keys per application so a leak has a small blast radius.
- Log who queried what, and keep citations so answers stay reviewable.
- Filter at retrieval time; do not fetch everything and hide results later.
- Review permissions on a schedule, because corpora and people change.
How it works, step by step
- Map documents to permission groups, then map groups to collections.
- Create collections so that no collection mixes two different access levels.
- Issue a scoped key per application and rotate keys on a defined schedule.
- Attach user identity to every query in your application layer.
- Enable audit logging and export entries to your SIEM.
- Test negative cases: a user must fail to retrieve a document they cannot open.
- Schedule periodic access reviews and delete stale collections deliberately.
Try it yourself
Open the AI API key security checklist →
Why retrieval needs its own access model
A RAG system answers questions from documents that were written for specific audiences. If retrieval ignores permissions, a user can receive content they could not open in the source system, and the answer will look authoritative because it cites a real document. That is why access control has to be enforced at retrieval time, not by hiding chunks after the fact.
The cleanest pattern is to make the collection the permission boundary. Documents that share the same audience live together; documents with different audiences live apart. Queries then run against the collections the caller is allowed to see, which keeps filtering logic explicit and testable.
Keys, identity and least privilege
Issue a distinct API key per application or environment, never a shared key across teams. Scope each key to the collections it needs, keep it in a secret manager rather than source control, and rotate it on a schedule or after any suspected exposure. On enterprise deployments, RBAC and SSO tie human access to your identity provider rather than to another static key.
Log the end-user identity with each query at the application layer. The API key identifies the application; the user identity identifies who asked the question, and audit reviews usually need the second one.
Audit logs and evidence
Per-request audit logs should capture the model, token counts, latency, user and region, so an investigator can reconstruct what was asked and what was returned. Export logs to your SIEM and align retention with your policy; Plugsky documents retention options and export as part of the security controls.
Citations support a second kind of review: for any answer, you can trace which chunks were used and therefore which collection and document were accessed. That makes the RAG pipeline auditable in a way that a plain chat log is not.
Putting it together on Plugsky
Plugsky supports per-collection isolation and encryption at rest, scoped keys, and enterprise controls including RBAC, SSO and audit logging with SIEM export. The same OpenAI-compatible endpoints - POST /v1/embeddings, POST /v1/rag/collections and POST /v1/rag/query - work across managed, VPC, on-prem and air-gapped deployments.
Walk through the API key security checklist before launch, and prototype the permission model on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Control | Plugsky RAG | Single shared index with filters | No access control |
|---|---|---|---|
| Boundary | Per-collection isolation and encryption | One index with metadata filters | One corpus for everyone |
| Keys | Scoped keys, RBAC and SSO on enterprise | Application-level filtering | Shared credential |
| Auditability | Per-request logs plus citations | Depends on custom logging | None |
| Negative tests | Can be validated per collection | Filter bugs leak results | Everything is exposed |
| Fit | Multi-team and regulated workloads | Small trusted groups | Public data only |
Frequently asked questions
Can RAG enforce document-level permissions?
Yes, by mapping documents to collections that match your permission groups and querying only the collections a caller may access. Mixing access levels in one collection makes enforcement harder.
How many API keys should we create?
One per application or environment, scoped to the collections it needs. Rotate keys on a schedule and store them in a secret manager.
What do audit logs contain?
Plugsky records per-request model, tokens, latency, user and region, supports SIEM export, and documents retention options for compliance review.
Does a reranker see documents the user cannot access?
It should not. Enforce permissions before reranking and generation, so only authorised chunks enter the candidate set.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.