RAG

How do you build RAG in Python with the Plugsky API?

Use the OpenAI Python SDK with the Plugsky base URL to build RAG in a few calls: upload documents to a collection, query it for ranked chunks with citations, then pass those chunks to chat completions. Because the API is OpenAI-compatible, your existing Python client code changes only in the base URL and model name.

Key facts

ClientOpenAI Python SDK with base_url set to the Plugsky API
RAG endpointsPOST /v1/rag/collections, document upload and POST /v1/rag/query
EmbeddingsPOST /v1/embeddings with plugsky-embed-v1 (1536d) or plugsky-embed-large (3072d)
RetrievalKeyword, vector and hybrid search with optional reranking and citations
FormatsPDF, DOCX, TXT, MD and HTML
Free tierplugsky-micro and plugsky-lite, no card required
Trial14-day full-access trial available
Product statusLive

TL;DR

  • One SDK covers chat, embeddings and RAG when the base URL points to Plugsky.
  • Create a collection first, then upload documents with metadata.
  • Query for chunks, then build the prompt from chunk text only.
  • Ask the model to cite sources and to refuse when context is missing.
  • Keep an evaluation script in the repo so changes can be compared.

How it works, step by step

  1. Install the OpenAI Python SDK and set your API key as an environment variable.
  2. Point the client at the Plugsky base URL instead of the default endpoint.
  3. Create a collection and store its id in your configuration.
  4. Upload documents with metadata that reflects department or access level.
  5. Query the collection with top_k and inspect chunks, scores and sources.
  6. Compose a chat request using only the retrieved chunks and require citations.
  7. Write the answer plus sources to your logs so evaluation has raw material.
1Install the OpenAIPython SDK and setyour API key as an2Point the client atthe Plugsky baseURL instead of the3Create a collectionand store its id inyour configuration.4Upload documentswith metadata thatreflects department5Query thecollection withtop_k and inspect6Compose a chatrequest using onlythe retrieved

Original data

POST /v1/rag/cRAG endpointsPOST /v1/embedEmbeddings14-day full-acTrialSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the OpenAI-compatible API tester →

Setting up the client

Install the OpenAI SDK as usual and create a client with the Plugsky base URL and your API key from the environment. Nothing else about the client changes: the same chat.completions.create and embeddings.create calls work. Keep the key out of source control and load it from a secret manager or environment variable in every environment.

Because chat, embeddings and RAG share one OpenAI-compatible surface, a single client instance can drive the whole pipeline. That keeps dependencies small and makes it easy to swap models later by changing a model name string.

Ingesting documents into a collection

Create a collection with a POST request and store the returned identifier. Upload documents with metadata that your retrieval filters will need later, such as source, department, version or access level. Plugsky accepts PDF, DOCX, TXT, MD and HTML, and handles chunking, embedding and indexing automatically using 500-token chunks with 50-token overlap by default.

Keep ingestion scripts idempotent: use a stable document identifier so re-uploads update rather than duplicate content. Metadata fields set here become the basis for filtered queries and, in multi-team systems, for permissions.

Querying and generating with citations

Send the user question to the query endpoint with a top_k value and optional reranking, then read the ranked chunks, scores and source references from the response. Build the chat prompt from those chunks only, and instruct the model to answer strictly from the provided context, cite the sources it used, and say it does not know when the context is insufficient.

Log the query, the retrieved chunk identifiers and the final answer together. That single log record supports debugging, evaluation and audit review, and it is much cheaper to add now than to reconstruct later.

Evaluating and iterating

Write a small evaluation script that runs a fixed set of questions against the collection and reports whether the correct chunks were retrieved and whether the answer stayed grounded. Run it whenever documents change, chunking changes or the answer model changes. Treat retrieval and generation scores separately so a failure points at one stage.

Try requests interactively with the OpenAI-compatible API tester, then start free with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

TaskPlugsky callYour codeThird-party stack
Chat generationChat completions with any of 30+ modelsPrompt and citation parsingSeparate provider SDK
EmbeddingsPOST /v1/embeddingsBatch corpus textEmbedding vendor plus store
IngestionCollection upload with automatic chunkingMetadata and idempotencyChunking and indexing pipeline
RetrievalQuery with top_k, modes and optional rerankingResult handling and prompt assemblyStore client plus reranker
OperationsManaged endpointsApplication monitoringYou run every component

Frequently asked questions

Which Python package should I use?

The standard OpenAI Python SDK works because the API is OpenAI-compatible. Set the base URL to Plugsky and keep the rest of your client code.

Do I need LangChain for RAG?

No. The pipeline is a few HTTP calls. LangChain or LlamaIndex can help with orchestration, but they are optional.

How do I handle large PDFs?

Upload them to a collection and let the platform chunk and index them. Long documents become multiple chunks, each carrying source references.

Can I use my existing embeddings code?

Yes. POST /v1/embeddings matches the OpenAI request shape, so existing Python embedding calls keep working after a base URL change.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.