RAG

How do you reduce the cost of a RAG pipeline?

RAG cost comes from four places: embedding and indexing, retrieval, prompt tokens sent to the model, and reranking or extra agent steps. The biggest levers are tighter retrieval so prompts stay small, caching repeated questions, routing easy requests to cheaper models, and re-embedding only when necessary. Flat self-serve plans remove per-token surprises.

Key facts

Pricing modelFlat monthly self-serve plans with unlimited fair-use usage
Per-token billingNo per-token charges or overage fees on self-serve plans
Cost driversEmbedding batches, top_k size, prompt tokens, reranking and agent steps
IngestionChunking, embedding and indexing are automatic per collection
RerankingOptional on queries; tune candidate count to control extra work
Batch limitsUp to 2,048 embedding inputs per request, max 8,191 tokens each
Models30+ models across tiers; route tasks by difficulty
Product statusLive

TL;DR

  • The prompt is usually the largest recurring cost: retrieve less, more precisely.
  • Cache repeated questions and stable answers before optimising anything else.
  • Route classification and extraction to fast models, synthesis to strong ones.
  • Re-embed only when the model or documents change, not on every deploy.
  • Measure quality and cost together; cheap answers that users re-ask are not cheap.

How it works, step by step

  1. Instrument the pipeline: retrieval calls, prompt tokens, model used and latency per request.
  2. Find the top question patterns and measure how many chunks each really needs.
  3. Reduce top_k and enable reranking so fewer, better chunks reach the model.
  4. Add a cache for repeated or near-identical questions with a sensible expiry.
  5. Route easy tasks to fast models and reserve frontier models for hard synthesis.
  6. Audit ingestion for duplicate documents and stale chunks to shrink the corpus.
  7. Re-check quality after every change so savings do not hide regressions.
1Instrument thepipeline: retrievalcalls, prompt2Find the topquestion patternsand measure how3Reduce top_k andenable reranking sofewer, better4Add a cache forrepeated ornear-identical5Route easy tasks tofast models andreserve frontier6Audit ingestion forduplicate documentsand stale chunks to

Try it yourself

Open the RAG cost calculator →

Where RAG cost actually accumulates

Ingestion is a batch cost: every document is chunked, embedded and indexed, and every chunk is stored. Retrieval is a per-query cost that scales with the candidate set you ask for. Generation is usually the largest recurring line, because prompt tokens include all retrieved chunks plus conversation history, and answers add completion tokens on top. Reranking and agent steps multiply the number of model calls per user request.

That ordering matters. Teams often start by shrinking the embedding model, when the real money is in sending ten loosely relevant chunks into every generation call. Instrument first, then optimise the biggest line.

The levers that usually pay off

Retrieve less, better. Lower top_k and add reranking so the prompt carries a few highly relevant chunks instead of many marginal ones. Cache. Repeated questions, shared documents and stable summaries can be served from a cache with a defined expiry. Route models. Classification, extraction and short answers rarely need a frontier model; reserve the strongest model for multi-step synthesis.

Keep the corpus clean. Duplicate documents, outdated versions and abandoned drafts all consume storage and pollute retrieval. Re-embed deliberately. Changing embedding models invalidates vectors, so batch those migrations instead of re-embedding on every deploy.

Cost models differ, so choose deliberately

Usage-based pricing makes cost proportional to traffic, which is comfortable at low volume and expensive at scale. Flat monthly pricing makes cost predictable and independent of how many tokens a prompt contains, which changes the optimisation question from how little can we send to how well can we answer. Plugsky self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees, so retrieval settings can be tuned for quality rather than token price.

Whatever the model, keep quality and cost on the same dashboard. A cheaper configuration that forces users to ask twice is a false economy, and it is easy to miss without measurement.

A practical optimisation loop

Start from production logs: which questions dominate, how many chunks are retrieved, which model answers, and how often users rephrase or retry. Then make one change at a time and re-run the evaluation set. Tightening top_k, adding reranking, caching frequent answers and routing easy requests to fast models are the changes most likely to show immediate savings without visible quality loss.

Model the numbers with the RAG cost calculator, then validate on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

LeverEffectRiskWhen to use
Lower top_kSmaller prompts, faster answersCan miss contextWhen chunks are precise
RerankingFewer, better chunks in the promptAdds a scoring passWhen retrieval over-fetches
CachingAvoids repeated generationStale answersFor stable, frequent questions
Model routingCheaper tasks on fast modelsQuality drop if misroutedWhen tasks differ in difficulty
Corpus hygieneLess storage and noiseCleanup effortAlways, on a schedule
Flat pricingPredictable monthly costFair-use policy appliesWhen token forecasting is hard

Frequently asked questions

What is the biggest cost in a RAG pipeline?

Usually generation: prompt tokens include every retrieved chunk plus history, and completion tokens add to that. Retrieval settings that reduce irrelevant context often save the most.

Does Plugsky charge per token?

No. Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.

How often should I re-embed documents?

Only when the embedding model changes or the documents change. Re-embedding on every deploy wastes work because vectors stay valid otherwise.

Is caching safe for RAG answers?

For stable, frequently asked questions, yes, with a sensible expiry. Questions involving live data or user-specific permissions should not be served from a shared cache.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How do I compare configurations fairly?

Fix a golden question set, change one variable at a time, and record retrieval quality, answer quality and latency alongside cost.