RAG

How do you make a RAG pipeline faster?

RAG latency comes from retrieval stages, prompt size, model speed and network distance. The biggest wins are usually smaller candidate sets, streaming so users see tokens early, routing simple questions to fast models, and running in a region close to users. Measure time to first token and total latency separately.

Key facts

StreamingChat streaming is live for token-by-token delivery
RetrievalKeyword, vector and hybrid search with optional reranking
Latency leverstop_k size, reranking, prompt tokens, model tier and region
Models30+ models across speed and capability tiers
RoutingSend easy questions to fast models and hard ones to stronger models
DeploymentManaged with region choice; VPC, on-prem and air-gapped options
Pricing modelFlat self-serve plans; no per-token billing to constrain prompt size
Product statusLive

TL;DR

  • Measure time to first token and total latency separately; they have different fixes.
  • Stream the answer so perceived latency drops without changing the pipeline.
  • Trim top_k and rerank so the prompt carries fewer, better chunks.
  • Route easy questions to fast models and escalate only when needed.
  • Reduce network distance by choosing the closest region for your users.

How it works, step by step

  1. Instrument the pipeline: retrieval time, reranking time, time to first token and total time.
  2. Establish a baseline at realistic concurrency rather than a single request.
  3. Reduce top_k and tune reranking to cut prompt size and scoring work.
  4. Enable streaming in the client so users see progress immediately.
  5. Add model routing for question types that do not need a frontier model.
  6. Cache repeated questions and stable answers with a defined expiry.
  7. Choose a region near your users and re-measure end to end.
1Instrument thepipeline: retrievaltime, reranking2Establish abaseline atrealistic3Reduce top_k andtune reranking tocut prompt size and4Enable streaming inthe client so userssee progress5Add model routingfor question typesthat do not need a6Cache repeatedquestions andstable answers with

Try it yourself

Open the API latency tester →

Where the time actually goes

A RAG request passes through several stages: query rewriting if you use it, retrieval, optional reranking, prompt assembly, model inference and possibly citation validation. Each has a different fix. Retrieval is usually milliseconds to tens of milliseconds; reranking adds a scoring pass; generation dominates total time and grows with prompt length.

Measure time to first token separately from total completion time. A long total with a fast first token feels responsive if you stream. A slow first token feels broken no matter how quickly the rest arrives.

The levers that matter most

Smaller prompts. Fewer, more relevant chunks reduce both retrieval work and generation time. Lower top_k and rerank to keep quality. Streaming. Delivering tokens as they generate transforms perceived latency without changing backend work. Model routing. Classification, extraction and short factual answers rarely need a frontier model; reserve the strongest model for complex synthesis.

Region and network. Distance adds fixed milliseconds to every call. Choose a deployment region close to your users and, for strict requirements, a private plane with predictable routing.

Caching and concurrency

Repeated questions are common in support and internal knowledge tools. A cache keyed on the normalised question, with a sensible expiry, removes both retrieval and generation for those hits. Cache only stable content; anything permission-sensitive or time-sensitive should bypass it.

Concurrency reveals different problems. Test with realistic parallel requests, because queueing, cold starts and rate limits affect tail latency far more than the average. Track percentiles rather than means, since users remember the slow requests.

Measuring and iterating on Plugsky

Run tests on the same endpoints used in production: POST /v1/rag/query for retrieval and chat completions with streaming for answers. Compare model tiers under identical retrieval settings so the model variable is isolated, and re-check after any corpus or chunking change since prompt size follows content.

Baseline your API latency with the API latency tester, then iterate on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

LeverEffect on first tokenEffect on totalTrade-off
StreamingLarge improvement in perceptionNoneClient must handle partial output
Smaller top_kSome improvementModerateMay miss context if set too low
Model routingDepends on modelLarge for easy tasksQuality risk if misrouted
CachingLarge for cache hitsLarge for cache hitsStale answers if expiry is wrong
Closer regionSmall fixed gainSmall fixed gainResidency constraints may decide

Frequently asked questions

What latency should I expect from RAG?

It depends on retrieval settings, model size, prompt length and network distance. Benchmark with your own pipeline and track percentiles rather than expecting a universal number.

Does reranking slow things down?

It adds a scoring pass over the candidate set. The cost is usually worthwhile when it lets you send fewer, better chunks to the model, but measure both effects.

How does streaming help if total time is unchanged?

Users see the answer begin immediately, which makes the interaction feel fast. Perceived latency often matters more than total completion time.

Should I cache RAG answers?

For stable, frequently asked questions, yes. Avoid caching permission-sensitive or time-sensitive answers, and set a short expiry for content that changes.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.