RAG

How does the RAG API work end to end?

The RAG API lifecycle has four stages: ingest documents into a collection, let the service chunk and embed them, query with keyword, vector or hybrid retrieval, and pass the ranked chunks with citations to a chat model. Plugsky exposes this on an OpenAI-compatible API with three endpoints and no vector database to operate.

Key facts

EndpointsPOST /v1/embeddings, POST /v1/rag/collections and POST /v1/rag/query
IngestionUpload PDF, DOCX, TXT, MD or HTML; chunking and embedding are automatic
Chunking500-token chunks with 50-token overlap by default
RetrievalKeyword, vector and hybrid search with optional cross-encoder reranking
CitationsEvery query returns ranked chunks with source attribution
PrivacyPer-collection encryption at rest; API data is not used for training
DeploymentManaged on self-serve; VPC, on-prem and air-gapped on enterprise
Product statusLive

TL;DR

  • Ingestion quality decides answer quality more than prompt wording does.
  • Store metadata with every document so filters and permissions work later.
  • Start with vector search, add hybrid and reranking when queries need it.
  • Always pass citations to the generator and validate answers against them.
  • Evaluate retrieval and generation separately so failures are diagnosable.

How it works, step by step

  1. Create a collection and record its identifier.
  2. Upload documents with metadata such as department, source and access level.
  3. Let the platform chunk, embed and index the content automatically.
  4. Query with top_k and inspect ranked chunks, scores and source references.
  5. Compose only the retrieved chunks into the chat prompt with citation instructions.
  6. Add hybrid search or reranking when keyword signals or precision matter.
  7. Build a golden question set and re-run it after every corpus or model change.
1Create a collectionand record itsidentifier.2Upload documentswith metadata suchas department,3Let the platformchunk, embed andindex the content4Query with top_kand inspect rankedchunks, scores and5Compose only theretrieved chunksinto the chat6Add hybrid searchor reranking whenkeyword signals or

Try it yourself

Open the RAG sandbox →

Stage one: ingestion and indexing

Everything downstream depends on what ingestion produces. Create a collection as the unit of isolation and configuration, then upload files with metadata. Plugsky accepts PDF, DOCX, TXT, MD and HTML, chunks documents into 500-token segments with 50-token overlap by default, and embeds and indexes them automatically. Metadata attached at upload time becomes the basis for filtering and access control later.

Keep collections aligned with meaning: one per bounded domain, product line or permission group. Mixing unrelated documents into one collection forces every query to compete against irrelevant chunks, which lowers precision and makes evaluation harder to interpret.

Stage two: retrieval that matches the question

Retrieval runs against a collection with a configurable top_k, and supports keyword, vector and hybrid modes plus optional cross-encoder reranking. Vector search handles paraphrase and meaning; keyword search catches identifiers, error codes and exact names; hybrid combines both. Reranking reorders a wider candidate set for precision. Every query returns ranked chunks with source attribution.

Choose the mode from the question mix, not from preference. Support queries with model numbers or policy codes usually need hybrid retrieval, while conceptual questions are often well served by vector search alone. Measure with a fixed question set before and after each change.

Stage three: grounded generation

Send the retrieved chunks to a chat completion with an instruction to answer only from the provided context and to cite sources. Keep the prompt explicit about refusal: if the chunks do not contain the answer, the model should say so rather than improvise. Because chunks carry file and page references, citations can be rendered inline or as a source list.

Keep the generator swappable. Any of the 30+ models on the same API can serve as the answer model, so a fast model can handle routine questions while a stronger model takes complex synthesis. Retrieval code does not change when the model does.

Stage four: evaluation and operations

Evaluate retrieval and generation separately. For retrieval, measure whether the correct chunk appears in the top k. For generation, check groundedness and citation accuracy against the retrieved chunks. Re-run the same golden set after every corpus update, chunking change or model switch so regressions are caught before users find them.

Operationally, monitor query volume, latency and empty-result rates, and use per-request audit logs for traceability. Try the flow with the RAG sandbox, then start free with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

StageWhat Plugsky providesAssembling componentsDoing nothing
IngestionAutomatic chunking, embedding and indexingYou build the pipelineNo retrieval
RetrievalKeyword, vector and hybrid with optional rerankingIntegrate store and rerankerGeneric model answers
CitationsRanked chunks with source attributionYou implement attributionNo sources
EvaluationFixed endpoints make A/B testing practicalEach component needs its own harnessQuality is unknown
OperationsManaged, with private deployment optionsYou run every serviceNo system to run

Frequently asked questions

What are the RAG API endpoints?

Plugsky exposes POST /v1/embeddings for vectors, POST /v1/rag/collections for ingestion and management, and POST /v1/rag/query for ranked chunks with citations.

What file formats can I ingest?

PDF, DOCX, TXT, MD and HTML. Documents are chunked, embedded and indexed automatically when uploaded to a collection.

How large are the chunks?

The default is 500-token chunks with 50-token overlap. Chunking choices affect retrieval quality, so test against your own questions.

Do I need my own vector database?

No. Collections store and retrieve chunks for you. You can still use POST /v1/embeddings standalone if you prefer to keep a vector store you already run.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.