RAG

How do you build RAG over Google Drive?

Google Drive RAG needs a connector, not a one-time export. Pull files through the Drive API, convert Google Docs formats to text, mirror file permissions into collection boundaries, and sync incrementally with changes. Then query with citations. Plugsky collections handle chunking, embedding and retrieval for the synced content.

Key facts

ApproachSync files via the Drive API into Plugsky collections
FormatsPDF, DOCX, TXT, MD and HTML ingestion; convert Google Docs exports to supported text
PermissionsScoped keys and per-collection isolation; mirror Drive sharing rules
SyncIncremental updates via the Drive changes feed; remove deleted files from the index
RetrievalKeyword, vector and hybrid search with optional reranking and citations
Data handlingPer-collection encryption at rest; API data is not used to train models
DeploymentManaged, VPC, on-prem and air-gapped options
Product statusLive

TL;DR

  • Treat Drive as a live source: sync changes, do not snapshot once.
  • Convert Google Docs, Sheets and Slides to text before ingestion.
  • Mirror Drive sharing rules into collection boundaries.
  • Remove deleted or unshared files from the index promptly.
  • Keep metadata such as owner, folder and modified date on every chunk.

How it works, step by step

  1. Set up a service account with read access scoped to the drives you need.
  2. Export supported formats and convert native Google files to text.
  3. Map sharing rules to collections or metadata filters for permissions.
  4. Upload documents to collections with owner, folder and modified date metadata.
  5. Store the Drive file id and revision so updates can be matched.
  6. Subscribe to the Drive changes feed and apply incremental updates.
  7. Test that a user who loses access also loses retrieval access.
1Set up a serviceaccount with readaccess scoped to2Export supportedformats and convertnative Google files3Map sharing rulesto collections ormetadata filters4Upload documents tocollections withowner, folder and5Store the Drivefile id andrevision so updates6Subscribe to theDrive changes feedand apply

Try it yourself

Open the RAG sandbox →

Why Drive needs a connector, not an export

A one-time export produces an index that is wrong within days. Files are edited, renamed, moved and shared constantly, and each of those events changes what should be retrievable. A connector keeps a stable mapping between the Drive file id and the document stored in your collection, so updates replace existing chunks instead of creating duplicates.

The Drive changes feed provides the incremental path: record a start page token, poll for changes on a schedule, and apply additions, updates and deletions. Batch small updates together, and handle rate limits with backoff rather than retrying immediately.

Converting Google formats to text

Google Docs, Sheets and Slides are not plain files; they need an export step. Export Docs to a text or HTML representation, Sheets to a structured format that preserves headers where possible, and Slides to notes and text frames. Files already stored as PDF or DOCX can be uploaded directly because Plugsky accepts PDF, DOCX, TXT, MD and HTML.

Preserve structure from headings where the export provides it, since structure-aware chunking keeps sections coherent. Attach metadata such as owner, folder path, modified date and source URL so citations resolve back to the original file.

Mirroring permissions

Drive permissions are the hardest part of the project. A retrieval index that ignores sharing rules will answer a question using a file the asker cannot open. The workable pattern is to mirror permissions into your retrieval layer: either split content into collections per audience, or store a permission label on each chunk and apply it as a filter on every query.

Whatever you choose, verify it continuously. Permission changes in Drive should propagate to the index, and negative tests should confirm that revoked access also revokes retrieval. This is the control auditors ask about first.

Querying the synced corpus

Once synced, queries behave like any other RAG workload: keyword, vector or hybrid retrieval with optional reranking, ranked chunks with citations, and answers composed by any of 30+ models on the OpenAI-compatible API. Users can ask in natural language and get a passage with a link back to the Drive file.

Start with a small scope, such as one shared drive, and validate permissions and freshness before expanding. Try the flow with the RAG sandbox, then run sync jobs against collections created on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

ConcernOne-time exportScheduled full resyncIncremental connector
FreshnessStale within daysCurrent at sync timeCurrent within the poll interval
CostLowest onceHigh repeated workProportional to changes
PermissionsBaked in at exportCan drift between runsPropagates on change
DeletesOften missedHandled if comparedHandled explicitly
OperationsManualScheduled jobsConnector to maintain

Frequently asked questions

Can Plugsky connect to Google Drive directly?

No. You run a small connector using the Drive API to sync files into Plugsky collections, where chunking, embedding and retrieval are managed.

Which Drive formats are supported?

Plugsky ingests PDF, DOCX, TXT, MD and HTML. Export Google Docs to text or HTML, Sheets to a structured format, and Slides to their text content before upload.

How do I handle Drive permissions?

Mirror sharing rules into collections or attach a permission label to each chunk and apply it as a query filter. Verify that revoked access also revokes retrieval.

How often should the sync run?

As often as your freshness needs. The Drive changes feed supports incremental updates, so frequent polls only process what changed.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.