Key facts
| Streaming | Chat streaming is live for token-by-token delivery |
| Retrieval | Keyword, vector and hybrid search with optional reranking |
| Latency levers | top_k size, reranking, prompt tokens, model tier and region |
| Models | 30+ models across speed and capability tiers |
| Routing | Send easy questions to fast models and hard ones to stronger models |
| Deployment | Managed with region choice; VPC, on-prem and air-gapped options |
| Pricing model | Flat self-serve plans; no per-token billing to constrain prompt size |
| Product status | Live |
TL;DR
- Measure time to first token and total latency separately; they have different fixes.
- Stream the answer so perceived latency drops without changing the pipeline.
- Trim top_k and rerank so the prompt carries fewer, better chunks.
- Route easy questions to fast models and escalate only when needed.
- Reduce network distance by choosing the closest region for your users.
How it works, step by step
- Instrument the pipeline: retrieval time, reranking time, time to first token and total time.
- Establish a baseline at realistic concurrency rather than a single request.
- Reduce top_k and tune reranking to cut prompt size and scoring work.
- Enable streaming in the client so users see progress immediately.
- Add model routing for question types that do not need a frontier model.
- Cache repeated questions and stable answers with a defined expiry.
- Choose a region near your users and re-measure end to end.
Try it yourself
Where the time actually goes
A RAG request passes through several stages: query rewriting if you use it, retrieval, optional reranking, prompt assembly, model inference and possibly citation validation. Each has a different fix. Retrieval is usually milliseconds to tens of milliseconds; reranking adds a scoring pass; generation dominates total time and grows with prompt length.
Measure time to first token separately from total completion time. A long total with a fast first token feels responsive if you stream. A slow first token feels broken no matter how quickly the rest arrives.
The levers that matter most
Smaller prompts. Fewer, more relevant chunks reduce both retrieval work and generation time. Lower top_k and rerank to keep quality. Streaming. Delivering tokens as they generate transforms perceived latency without changing backend work. Model routing. Classification, extraction and short factual answers rarely need a frontier model; reserve the strongest model for complex synthesis.
Region and network. Distance adds fixed milliseconds to every call. Choose a deployment region close to your users and, for strict requirements, a private plane with predictable routing.
Caching and concurrency
Repeated questions are common in support and internal knowledge tools. A cache keyed on the normalised question, with a sensible expiry, removes both retrieval and generation for those hits. Cache only stable content; anything permission-sensitive or time-sensitive should bypass it.
Concurrency reveals different problems. Test with realistic parallel requests, because queueing, cold starts and rate limits affect tail latency far more than the average. Track percentiles rather than means, since users remember the slow requests.
Measuring and iterating on Plugsky
Run tests on the same endpoints used in production: POST /v1/rag/query for retrieval and chat completions with streaming for answers. Compare model tiers under identical retrieval settings so the model variable is isolated, and re-check after any corpus or chunking change since prompt size follows content.
Baseline your API latency with the API latency tester, then iterate on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Lever | Effect on first token | Effect on total | Trade-off |
|---|---|---|---|
| Streaming | Large improvement in perception | None | Client must handle partial output |
| Smaller top_k | Some improvement | Moderate | May miss context if set too low |
| Model routing | Depends on model | Large for easy tasks | Quality risk if misrouted |
| Caching | Large for cache hits | Large for cache hits | Stale answers if expiry is wrong |
| Closer region | Small fixed gain | Small fixed gain | Residency constraints may decide |
Frequently asked questions
What latency should I expect from RAG?
It depends on retrieval settings, model size, prompt length and network distance. Benchmark with your own pipeline and track percentiles rather than expecting a universal number.
Does reranking slow things down?
It adds a scoring pass over the candidate set. The cost is usually worthwhile when it lets you send fewer, better chunks to the model, but measure both effects.
How does streaming help if total time is unchanged?
Users see the answer begin immediately, which makes the interaction feel fast. Perceived latency often matters more than total completion time.
Should I cache RAG answers?
For stable, frequently asked questions, yes. Avoid caching permission-sensitive or time-sensitive answers, and set a short expiry for content that changes.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.