Developer + API

How do you build retrieval-augmented generation with Plugsky?

Create a RAG collection, upload your documents, and query it: Plugsky chunks, embeds and indexes the files automatically, then returns ranked chunks with citations you can pass into a chat completion. Supported formats include PDF, DOCX, TXT, MD and HTML. Your documents stay encrypted and are never used to train models.

Key facts

CollectionsCreate with POST /v1/rag/collections
UploadMultipart document upload with optional metadata
FormatsPDF, DOCX, TXT, MD, HTML
Chunking500-token chunks with 50-token overlap by default
Querytop_k plus optional rerank returns ranked chunks
CitationsChunks include source references such as file and page
Data useDocuments never used to train models
Product statusLive

TL;DR

  • Three steps: create a collection, upload documents, query.
  • Chunking, embedding and indexing are automatic.
  • Every result carries its source so answers can cite evidence.
  • Compose retrieved chunks into a chat completion for grounded answers.
  • Bring your own embeddings or vector store when you need control.

How it works, step by step

  1. Create a collection for each knowledge base you want to isolate.
  2. Upload documents as PDF, DOCX, TXT, MD or HTML with metadata.
  3. Confirm chunking and embedding defaults fit your content.
  4. Query with a natural-language question, top_k and rerank enabled.
  5. Read chunk text, scores and source references from the response.
  6. Inject the top chunks into a chat completion and require citations.
1Create a collectionfor each knowledgebase you want to2Upload documents asPDF, DOCX, TXT, MDor HTML with3Confirm chunkingand embeddingdefaults fit your4Query with anatural-languagequestion, top_k and5Read chunk text,scores and sourcereferences from the6Inject the topchunks into a chatcompletion and

Try it yourself

Open the RAG sandbox →

The three-step RAG workflow

Plugsky's RAG API reduces retrieval-augmented generation to three calls. Create a collection to hold one knowledge base, upload the documents that belong to it, then query the collection with natural-language questions. There is no separate embedding or indexing step to orchestrate — the platform parses files, splits them into chunks, embeds each chunk and indexes the vectors for search.

Keep collections narrow. One collection per domain, product or department is easier to evaluate, secure and refresh than a single pool of unrelated documents.

Chunking, embeddings and storage

Documents are chunked at 500 tokens with 50 tokens of overlap by default, then embedded with the plugsky-embed model. Overlap keeps context that would otherwise be lost at boundaries. For structured material, test alternate chunk sizes against a small labelled question set before loading thousands of files.

  • pgvector is the default vector store.
  • Pinecone, Qdrant and Weaviate are supported on Enterprise.
  • Bring your own embeddings by setting embedding_model to custom at collection creation.

Querying and composing cited answers

A query returns chunks ranked by similarity, each with text, score, source and metadata; enable rerank: true to add a second-stage ordering. To compose an answer, join the top chunks into a system or context message and instruct the chat model to answer only from that context and cite sources. Because retrieval and generation are separate calls, you can inspect exactly what evidence the model received — which is the difference between a demo and a system an auditor will accept.

Keep top_k small and relevant: every extra chunk consumes context and can distract the model.

Data handling and deployment choices

Documents are encrypted at rest, used only for retrieval and never used to train models. The RAG API runs in region-locked data planes and is available in VPC, on-prem and air-gapped deployments for teams that cannot use multi-tenant infrastructure. Self-serve collections scale to large document counts, and Enterprise removes the hard cap.

Honest boundary: the RAG API is a retrieval layer, not a document management system. If you need versioning, complex per-user permissions or strict retention rules, keep the source of truth in your own system and sync only what should be searchable.

Honest comparison

ApproachPlugsky RAG APIEmbeddings API plus your own storeOff-the-shelf RAG SaaS
Setup effortThree callsChunk, embed, store and query codeVaries
Chunking and embeddingManagedYoursVendor defaults
CitationsSource and score per chunkYou implementOften available
Vector storepgvector default; BYO on EnterpriseYour choiceVendor's store
DeploymentCloud, VPC, on-prem, air-gappedWherever you hostUsually SaaS only
Data controlsNo training on your data; regional planesYou controlCheck vendor terms

Frequently asked questions

Which file formats can I upload?

PDF, DOCX, TXT, MD and HTML are supported by default.

How does chunking work?

Documents are split into 500-token chunks with 50-token overlap by default; you can test different sizes for structured content.

Do I need a separate vector database?

No. pgvector is the default store. Enterprise can use Pinecone, Qdrant or Weaviate, and Private Endpoint deployments can bring their own.

Can I use my own embeddings?

Yes. Set embedding_model to custom when creating the collection and supply vectors with the document upload.

Does RAG train on my documents?

No. Documents are stored encrypted and used only for retrieval; they are never used to train any model.

How do I get citations in answers?

Query results include the source of each chunk; pass those sources into the chat prompt and instruct the model to cite them.

Where can RAG run for regulated workloads?

In region-locked data planes or in VPC, on-prem and air-gapped deployments on Enterprise.

Cite this page

Plugsky (2026). “Plugsky RAG Docs — Chat With Your Documents”. Plugsky. Available at: https://plugsky.com/articles/rag-docs (last updated 2026-09-25).