RAG

How do you build RAG in TypeScript with the Plugsky API?

In TypeScript, build RAG with the OpenAI Node SDK pointed at the Plugsky base URL: upload documents to a collection, query it for ranked chunks, then pass those chunks to a chat completion that must cite sources. Types come from the SDK, and the API stays OpenAI-compatible so the code is portable.

Key facts

ClientOpenAI Node SDK with baseURL set to the Plugsky API
RAG endpointsPOST /v1/rag/collections, document upload and POST /v1/rag/query
EmbeddingsPOST /v1/embeddings with plugsky-embed-v1 (1536d) or plugsky-embed-large (3072d)
RetrievalKeyword, vector and hybrid search with optional reranking and citations
StreamingChat streaming is live for token-by-token responses in Node
Free tierplugsky-micro and plugsky-lite, no card required
DeploymentManaged, VPC, on-prem and air-gapped options
Product statusLive

TL;DR

  • One SDK for chat, embeddings and RAG keeps the TypeScript surface small.
  • Create the collection, upload with metadata, then query for chunks.
  • Keep prompt assembly typed so citation parsing cannot drift.
  • Streaming improves perceived latency when the answer is long.
  • Store the query plus retrieved chunk ids for evaluation and audit.

How it works, step by step

  1. Install the OpenAI package and set the API key in the environment.
  2. Create a client with the Plugsky base URL and reuse it across the app.
  3. Create a collection and keep its id in typed configuration.
  4. Upload documents with metadata fields used later for filters.
  5. Query the collection with top_k and read chunks, scores and sources.
  6. Build a typed prompt from retrieved chunks and require inline citations.
  7. Log the question, retrieved chunk ids and answer for evaluation.
1Install the OpenAIpackage and set theAPI key in the2Create a clientwith the Plugskybase URL and reuse3Create a collectionand keep its id intyped4Upload documentswith metadatafields used later5Query thecollection withtop_k and read6Build a typedprompt fromretrieved chunks

Try it yourself

Open the OpenAI-compatible API tester →

Project setup and client configuration

Install the OpenAI Node SDK, then construct the client once with the Plugsky base URL and an API key read from the environment. Keep the client in a module that the rest of the application imports, so there is exactly one place where the base URL and default model live. That single point makes environment changes and future migrations trivial.

TypeScript helps most around the response shapes: define a small type for a retrieved chunk with text, score and source fields, and map the API response into it immediately. Everything downstream then works with a stable internal contract instead of raw JSON.

Creating a collection and uploading documents

Create a collection with a POST call and persist its identifier. Upload documents with metadata such as source, department or version; the platform accepts PDF, DOCX, TXT, MD and HTML, and chunking, embedding and indexing happen automatically with 500-token chunks and 50-token overlap by default.

Make the upload path idempotent by using a deterministic document id, so a retry or a re-run updates existing content instead of duplicating it. In serverless deployments, move uploads into background jobs: ingestion can take longer than a request timeout.

Querying and composing the answer

Send the user question to the query endpoint with top_k and optional reranking, then map the ranked chunks into your internal type. Compose the chat prompt from those chunks only, and instruct the model to answer strictly from context, cite the sources used, and decline when the answer is not present.

For user-facing chat, enable streaming so tokens appear as they are generated. Stream the grounded answer, then attach the citation list once the response completes. Keep the prompt template in one file and treat changes to it like code changes that require evaluation.

Wrapping it as a service

Expose a small internal endpoint that accepts a question, calls the query endpoint, assembles the prompt and returns the answer with citations. Add request logging with the question, retrieved chunk ids and model used, then run a fixed evaluation set through the same path on every change. That closes the loop between development and production behaviour.

Use the OpenAI-compatible API tester to check requests, then start free with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

TaskPlugsky callYour TypeScript codeThird-party stack
Chat generationChat completions with streamingPrompt template and citation parsingSeparate provider SDK
EmbeddingsPOST /v1/embeddingsBatch and map responsesEmbedding vendor plus store
IngestionCollection upload with automatic chunkingBackground job and idempotencyCustom ingestion pipeline
RetrievalQuery with top_k, modes and rerankingTyped chunk mappingStore client plus reranker
OperationsManaged endpointsLogging and evaluation harnessYou run every component

Frequently asked questions

Which TypeScript package should I use?

The OpenAI Node SDK works because the API is OpenAI-compatible. Set the base URL to Plugsky and keep your existing typed client code.

Does streaming work with RAG answers?

Yes. Streaming is live on the chat endpoint, so you can stream the grounded answer and append citations when the stream finishes.

Can I run ingestion in a serverless function?

You can start it there, but long uploads may exceed request timeouts. Move ingestion into a background job or queue for reliability.

Can I use my existing embeddings code?

Yes. POST /v1/embeddings matches the OpenAI request shape, so existing TypeScript embedding calls keep working after a base URL change.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.