Key facts
| Collections | Create with POST /v1/rag/collections |
| Upload | Multipart document upload with optional metadata |
| Formats | PDF, DOCX, TXT, MD, HTML |
| Chunking | 500-token chunks with 50-token overlap by default |
| Query | top_k plus optional rerank returns ranked chunks |
| Citations | Chunks include source references such as file and page |
| Data use | Documents never used to train models |
| Product status | Live |
TL;DR
- Three steps: create a collection, upload documents, query.
- Chunking, embedding and indexing are automatic.
- Every result carries its source so answers can cite evidence.
- Compose retrieved chunks into a chat completion for grounded answers.
- Bring your own embeddings or vector store when you need control.
How it works, step by step
- Create a collection for each knowledge base you want to isolate.
- Upload documents as PDF, DOCX, TXT, MD or HTML with metadata.
- Confirm chunking and embedding defaults fit your content.
- Query with a natural-language question, top_k and rerank enabled.
- Read chunk text, scores and source references from the response.
- Inject the top chunks into a chat completion and require citations.
Try it yourself
The three-step RAG workflow
Plugsky's RAG API reduces retrieval-augmented generation to three calls. Create a collection to hold one knowledge base, upload the documents that belong to it, then query the collection with natural-language questions. There is no separate embedding or indexing step to orchestrate — the platform parses files, splits them into chunks, embeds each chunk and indexes the vectors for search.
Keep collections narrow. One collection per domain, product or department is easier to evaluate, secure and refresh than a single pool of unrelated documents.
Chunking, embeddings and storage
Documents are chunked at 500 tokens with 50 tokens of overlap by default, then embedded with the plugsky-embed model. Overlap keeps context that would otherwise be lost at boundaries. For structured material, test alternate chunk sizes against a small labelled question set before loading thousands of files.
- pgvector is the default vector store.
- Pinecone, Qdrant and Weaviate are supported on Enterprise.
- Bring your own embeddings by setting embedding_model to custom at collection creation.
Querying and composing cited answers
A query returns chunks ranked by similarity, each with text, score, source and metadata; enable rerank: true to add a second-stage ordering. To compose an answer, join the top chunks into a system or context message and instruct the chat model to answer only from that context and cite sources. Because retrieval and generation are separate calls, you can inspect exactly what evidence the model received — which is the difference between a demo and a system an auditor will accept.
Keep top_k small and relevant: every extra chunk consumes context and can distract the model.
Data handling and deployment choices
Documents are encrypted at rest, used only for retrieval and never used to train models. The RAG API runs in region-locked data planes and is available in VPC, on-prem and air-gapped deployments for teams that cannot use multi-tenant infrastructure. Self-serve collections scale to large document counts, and Enterprise removes the hard cap.
Honest boundary: the RAG API is a retrieval layer, not a document management system. If you need versioning, complex per-user permissions or strict retention rules, keep the source of truth in your own system and sync only what should be searchable.
Honest comparison
| Approach | Plugsky RAG API | Embeddings API plus your own store | Off-the-shelf RAG SaaS |
|---|---|---|---|
| Setup effort | Three calls | Chunk, embed, store and query code | Varies |
| Chunking and embedding | Managed | Yours | Vendor defaults |
| Citations | Source and score per chunk | You implement | Often available |
| Vector store | pgvector default; BYO on Enterprise | Your choice | Vendor's store |
| Deployment | Cloud, VPC, on-prem, air-gapped | Wherever you host | Usually SaaS only |
| Data controls | No training on your data; regional planes | You control | Check vendor terms |
Frequently asked questions
Which file formats can I upload?
PDF, DOCX, TXT, MD and HTML are supported by default.
How does chunking work?
Documents are split into 500-token chunks with 50-token overlap by default; you can test different sizes for structured content.
Do I need a separate vector database?
No. pgvector is the default store. Enterprise can use Pinecone, Qdrant or Weaviate, and Private Endpoint deployments can bring their own.
Can I use my own embeddings?
Yes. Set embedding_model to custom when creating the collection and supply vectors with the document upload.
Does RAG train on my documents?
No. Documents are stored encrypted and used only for retrieval; they are never used to train any model.
How do I get citations in answers?
Query results include the source of each chunk; pass those sources into the chat prompt and instruct the model to cite them.
Where can RAG run for regulated workloads?
In region-locked data planes or in VPC, on-prem and air-gapped deployments on Enterprise.
Plugsky (2026). “Plugsky RAG Docs — Chat With Your Documents”. Plugsky. Available at: https://plugsky.com/articles/rag-docs (last updated 2026-09-25).