Key facts
| Endpoints | POST /v1/embeddings, POST /v1/rag/collections and POST /v1/rag/query |
| Ingestion | Upload PDF, DOCX, TXT, MD or HTML; chunking and embedding are automatic |
| Chunking | 500-token chunks with 50-token overlap by default |
| Retrieval | Keyword, vector and hybrid search with optional cross-encoder reranking |
| Citations | Every query returns ranked chunks with source attribution |
| Privacy | Per-collection encryption at rest; API data is not used for training |
| Deployment | Managed on self-serve; VPC, on-prem and air-gapped on enterprise |
| Product status | Live |
TL;DR
- Ingestion quality decides answer quality more than prompt wording does.
- Store metadata with every document so filters and permissions work later.
- Start with vector search, add hybrid and reranking when queries need it.
- Always pass citations to the generator and validate answers against them.
- Evaluate retrieval and generation separately so failures are diagnosable.
How it works, step by step
- Create a collection and record its identifier.
- Upload documents with metadata such as department, source and access level.
- Let the platform chunk, embed and index the content automatically.
- Query with top_k and inspect ranked chunks, scores and source references.
- Compose only the retrieved chunks into the chat prompt with citation instructions.
- Add hybrid search or reranking when keyword signals or precision matter.
- Build a golden question set and re-run it after every corpus or model change.
Try it yourself
Stage one: ingestion and indexing
Everything downstream depends on what ingestion produces. Create a collection as the unit of isolation and configuration, then upload files with metadata. Plugsky accepts PDF, DOCX, TXT, MD and HTML, chunks documents into 500-token segments with 50-token overlap by default, and embeds and indexes them automatically. Metadata attached at upload time becomes the basis for filtering and access control later.
Keep collections aligned with meaning: one per bounded domain, product line or permission group. Mixing unrelated documents into one collection forces every query to compete against irrelevant chunks, which lowers precision and makes evaluation harder to interpret.
Stage two: retrieval that matches the question
Retrieval runs against a collection with a configurable top_k, and supports keyword, vector and hybrid modes plus optional cross-encoder reranking. Vector search handles paraphrase and meaning; keyword search catches identifiers, error codes and exact names; hybrid combines both. Reranking reorders a wider candidate set for precision. Every query returns ranked chunks with source attribution.
Choose the mode from the question mix, not from preference. Support queries with model numbers or policy codes usually need hybrid retrieval, while conceptual questions are often well served by vector search alone. Measure with a fixed question set before and after each change.
Stage three: grounded generation
Send the retrieved chunks to a chat completion with an instruction to answer only from the provided context and to cite sources. Keep the prompt explicit about refusal: if the chunks do not contain the answer, the model should say so rather than improvise. Because chunks carry file and page references, citations can be rendered inline or as a source list.
Keep the generator swappable. Any of the 30+ models on the same API can serve as the answer model, so a fast model can handle routine questions while a stronger model takes complex synthesis. Retrieval code does not change when the model does.
Stage four: evaluation and operations
Evaluate retrieval and generation separately. For retrieval, measure whether the correct chunk appears in the top k. For generation, check groundedness and citation accuracy against the retrieved chunks. Re-run the same golden set after every corpus update, chunking change or model switch so regressions are caught before users find them.
Operationally, monitor query volume, latency and empty-result rates, and use per-request audit logs for traceability. Try the flow with the RAG sandbox, then start free with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Stage | What Plugsky provides | Assembling components | Doing nothing |
|---|---|---|---|
| Ingestion | Automatic chunking, embedding and indexing | You build the pipeline | No retrieval |
| Retrieval | Keyword, vector and hybrid with optional reranking | Integrate store and reranker | Generic model answers |
| Citations | Ranked chunks with source attribution | You implement attribution | No sources |
| Evaluation | Fixed endpoints make A/B testing practical | Each component needs its own harness | Quality is unknown |
| Operations | Managed, with private deployment options | You run every service | No system to run |
Frequently asked questions
What are the RAG API endpoints?
Plugsky exposes POST /v1/embeddings for vectors, POST /v1/rag/collections for ingestion and management, and POST /v1/rag/query for ranked chunks with citations.
What file formats can I ingest?
PDF, DOCX, TXT, MD and HTML. Documents are chunked, embedded and indexed automatically when uploaded to a collection.
How large are the chunks?
The default is 500-token chunks with 50-token overlap. Chunking choices affect retrieval quality, so test against your own questions.
Do I need my own vector database?
No. Collections store and retrieve chunks for you. You can still use POST /v1/embeddings standalone if you prefer to keep a vector store you already run.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.