Key facts
| Pricing model | Flat monthly self-serve plans with unlimited fair-use usage |
| Per-token billing | No per-token charges or overage fees on self-serve plans |
| Cost drivers | Embedding batches, top_k size, prompt tokens, reranking and agent steps |
| Ingestion | Chunking, embedding and indexing are automatic per collection |
| Reranking | Optional on queries; tune candidate count to control extra work |
| Batch limits | Up to 2,048 embedding inputs per request, max 8,191 tokens each |
| Models | 30+ models across tiers; route tasks by difficulty |
| Product status | Live |
TL;DR
- The prompt is usually the largest recurring cost: retrieve less, more precisely.
- Cache repeated questions and stable answers before optimising anything else.
- Route classification and extraction to fast models, synthesis to strong ones.
- Re-embed only when the model or documents change, not on every deploy.
- Measure quality and cost together; cheap answers that users re-ask are not cheap.
How it works, step by step
- Instrument the pipeline: retrieval calls, prompt tokens, model used and latency per request.
- Find the top question patterns and measure how many chunks each really needs.
- Reduce top_k and enable reranking so fewer, better chunks reach the model.
- Add a cache for repeated or near-identical questions with a sensible expiry.
- Route easy tasks to fast models and reserve frontier models for hard synthesis.
- Audit ingestion for duplicate documents and stale chunks to shrink the corpus.
- Re-check quality after every change so savings do not hide regressions.
Try it yourself
Open the RAG cost calculator →
Where RAG cost actually accumulates
Ingestion is a batch cost: every document is chunked, embedded and indexed, and every chunk is stored. Retrieval is a per-query cost that scales with the candidate set you ask for. Generation is usually the largest recurring line, because prompt tokens include all retrieved chunks plus conversation history, and answers add completion tokens on top. Reranking and agent steps multiply the number of model calls per user request.
That ordering matters. Teams often start by shrinking the embedding model, when the real money is in sending ten loosely relevant chunks into every generation call. Instrument first, then optimise the biggest line.
The levers that usually pay off
Retrieve less, better. Lower top_k and add reranking so the prompt carries a few highly relevant chunks instead of many marginal ones. Cache. Repeated questions, shared documents and stable summaries can be served from a cache with a defined expiry. Route models. Classification, extraction and short answers rarely need a frontier model; reserve the strongest model for multi-step synthesis.
Keep the corpus clean. Duplicate documents, outdated versions and abandoned drafts all consume storage and pollute retrieval. Re-embed deliberately. Changing embedding models invalidates vectors, so batch those migrations instead of re-embedding on every deploy.
Cost models differ, so choose deliberately
Usage-based pricing makes cost proportional to traffic, which is comfortable at low volume and expensive at scale. Flat monthly pricing makes cost predictable and independent of how many tokens a prompt contains, which changes the optimisation question from how little can we send to how well can we answer. Plugsky self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees, so retrieval settings can be tuned for quality rather than token price.
Whatever the model, keep quality and cost on the same dashboard. A cheaper configuration that forces users to ask twice is a false economy, and it is easy to miss without measurement.
A practical optimisation loop
Start from production logs: which questions dominate, how many chunks are retrieved, which model answers, and how often users rephrase or retry. Then make one change at a time and re-run the evaluation set. Tightening top_k, adding reranking, caching frequent answers and routing easy requests to fast models are the changes most likely to show immediate savings without visible quality loss.
Model the numbers with the RAG cost calculator, then validate on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Lever | Effect | Risk | When to use |
|---|---|---|---|
| Lower top_k | Smaller prompts, faster answers | Can miss context | When chunks are precise |
| Reranking | Fewer, better chunks in the prompt | Adds a scoring pass | When retrieval over-fetches |
| Caching | Avoids repeated generation | Stale answers | For stable, frequent questions |
| Model routing | Cheaper tasks on fast models | Quality drop if misrouted | When tasks differ in difficulty |
| Corpus hygiene | Less storage and noise | Cleanup effort | Always, on a schedule |
| Flat pricing | Predictable monthly cost | Fair-use policy applies | When token forecasting is hard |
Frequently asked questions
What is the biggest cost in a RAG pipeline?
Usually generation: prompt tokens include every retrieved chunk plus history, and completion tokens add to that. Retrieval settings that reduce irrelevant context often save the most.
Does Plugsky charge per token?
No. Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.
How often should I re-embed documents?
Only when the embedding model changes or the documents change. Re-embedding on every deploy wastes work because vectors stay valid otherwise.
Is caching safe for RAG answers?
For stable, frequently asked questions, yes, with a sensible expiry. Questions involving live data or user-specific permissions should not be served from a shared cache.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How do I compare configurations fairly?
Fix a golden question set, change one variable at a time, and record retrieval quality, answer quality and latency alongside cost.