RAG

How do you implement access control in RAG?

RAG access control means partitioning documents into collections that match your permission model, issuing scoped API keys per team or application, and logging every query with the user and collection involved. Plugsky supports per-collection encryption, scoped keys, RBAC and SSO on enterprise, and per-request audit logs that export to your SIEM.

Key facts

PartitioningCollections act as retrieval and encryption boundaries
KeysScoped API keys per application or team; RBAC and SSO on enterprise
Audit logsPer-request model, tokens, latency, user and region with SIEM export
Data handlingPer-collection encryption at rest; API data is not used to train models
RetrievalKeyword, vector and hybrid search with citations for review
DeploymentManaged, VPC, on-prem and air-gapped options
RetentionConfigurable retention with export up to the published maximum
Product statusLive

TL;DR

  • Collection boundaries should mirror the way your organisation grants access.
  • Scope keys per application so a leak has a small blast radius.
  • Log who queried what, and keep citations so answers stay reviewable.
  • Filter at retrieval time; do not fetch everything and hide results later.
  • Review permissions on a schedule, because corpora and people change.

How it works, step by step

  1. Map documents to permission groups, then map groups to collections.
  2. Create collections so that no collection mixes two different access levels.
  3. Issue a scoped key per application and rotate keys on a defined schedule.
  4. Attach user identity to every query in your application layer.
  5. Enable audit logging and export entries to your SIEM.
  6. Test negative cases: a user must fail to retrieve a document they cannot open.
  7. Schedule periodic access reviews and delete stale collections deliberately.
1Map documents topermission groups,then map groups to2Create collectionsso that nocollection mixes3Issue a scoped keyper application androtate keys on a4Attach useridentity to everyquery in your5Enable auditlogging and exportentries to your6Test negativecases: a user mustfail to retrieve a

Try it yourself

Open the AI API key security checklist →

Why retrieval needs its own access model

A RAG system answers questions from documents that were written for specific audiences. If retrieval ignores permissions, a user can receive content they could not open in the source system, and the answer will look authoritative because it cites a real document. That is why access control has to be enforced at retrieval time, not by hiding chunks after the fact.

The cleanest pattern is to make the collection the permission boundary. Documents that share the same audience live together; documents with different audiences live apart. Queries then run against the collections the caller is allowed to see, which keeps filtering logic explicit and testable.

Keys, identity and least privilege

Issue a distinct API key per application or environment, never a shared key across teams. Scope each key to the collections it needs, keep it in a secret manager rather than source control, and rotate it on a schedule or after any suspected exposure. On enterprise deployments, RBAC and SSO tie human access to your identity provider rather than to another static key.

Log the end-user identity with each query at the application layer. The API key identifies the application; the user identity identifies who asked the question, and audit reviews usually need the second one.

Audit logs and evidence

Per-request audit logs should capture the model, token counts, latency, user and region, so an investigator can reconstruct what was asked and what was returned. Export logs to your SIEM and align retention with your policy; Plugsky documents retention options and export as part of the security controls.

Citations support a second kind of review: for any answer, you can trace which chunks were used and therefore which collection and document were accessed. That makes the RAG pipeline auditable in a way that a plain chat log is not.

Putting it together on Plugsky

Plugsky supports per-collection isolation and encryption at rest, scoped keys, and enterprise controls including RBAC, SSO and audit logging with SIEM export. The same OpenAI-compatible endpoints - POST /v1/embeddings, POST /v1/rag/collections and POST /v1/rag/query - work across managed, VPC, on-prem and air-gapped deployments.

Walk through the API key security checklist before launch, and prototype the permission model on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

ControlPlugsky RAGSingle shared index with filtersNo access control
BoundaryPer-collection isolation and encryptionOne index with metadata filtersOne corpus for everyone
KeysScoped keys, RBAC and SSO on enterpriseApplication-level filteringShared credential
AuditabilityPer-request logs plus citationsDepends on custom loggingNone
Negative testsCan be validated per collectionFilter bugs leak resultsEverything is exposed
FitMulti-team and regulated workloadsSmall trusted groupsPublic data only

Frequently asked questions

Can RAG enforce document-level permissions?

Yes, by mapping documents to collections that match your permission groups and querying only the collections a caller may access. Mixing access levels in one collection makes enforcement harder.

How many API keys should we create?

One per application or environment, scoped to the collections it needs. Rotate keys on a schedule and store them in a secret manager.

What do audit logs contain?

Plugsky records per-request model, tokens, latency, user and region, supports SIEM export, and documents retention options for compliance review.

Does a reranker see documents the user cannot access?

It should not. Enforce permissions before reranking and generation, so only authorised chunks enter the candidate set.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.