Key facts
| RAG retrieval | Keyword, vector and hybrid search with optional reranking |
| Citations | Ranked chunks return with source attribution |
| Ingestion | Documents chunked, embedded and indexed automatically per collection |
| Models | 30+ models with different context capabilities behind one API |
| Cost profile | Flat self-serve plans; no per-token billing pressure on prompt size |
| Latency | Smaller prompts generally produce faster first tokens |
| Deployment | Managed, VPC, on-prem and air-gapped options |
| Product status | Live |
TL;DR
- Long context is simple but scales poorly as the corpus grows.
- RAG narrows the input and keeps answers traceable to sources.
- Attention can thin out across very long inputs, so more text is not always better.
- Whole-document tasks favour long context; large corpora favour retrieval.
- Combine them: retrieve candidates, then reason over a richer context.
How it works, step by step
- Estimate the corpus size and how often it changes.
- Identify whether questions need one passage or reasoning across a whole document.
- Test answering with the full document in context on a sample of real questions.
- Test the same questions with retrieval and compare quality and latency.
- Check citation requirements: retrieval gives references, long context does not by default.
- Model the cost of sending full documents versus retrieved passages at your volume.
- Choose a hybrid design where retrieval selects and long context reasons.
Try it yourself
Open the context window comparison →
What long context is good at
Long context is the simplest architecture: put the documents in the prompt and ask the question. It works well when the corpus is small, when the task genuinely needs the whole document such as summarising a report or comparing two contracts, and when latency is less important than depth of reasoning.
It also removes retrieval risk. There is no chance of missing a relevant passage because the retriever failed, provided the entire document fits within the window along with the instructions and the answer.
Where retrieval still wins
RAG scales to corpora that will never fit in a context window and keeps prompt size proportional to the question rather than the library. That matters for latency and for cost on providers that bill by token. It also provides citations by construction: the answer references the chunks that were retrieved, so users can verify the evidence.
Practical experience also shows that models do not always use very long inputs evenly; information buried in the middle of a large context can be under-weighted. Retrieval puts the relevant passage in a small, focused context where attention is concentrated.
The hybrid pattern
The strongest designs borrow from both. Use retrieval to select a candidate set, then expand each hit to its surrounding section so the model sees complete arguments rather than fragments. For whole-document questions, retrieve the document and pass it in full, using retrieval to decide which documents deserve the space.
This keeps the benefits of traceability while giving the model richer material for reasoning. It also gives you a dial: widen the context for hard questions, narrow it for high-volume ones.
Deciding with measurements
Run the same question set through both approaches and compare answer quality, citation accuracy, latency and cost. Include questions whose answer spans two documents, because that is where long context often shines, and questions over a large corpus, where retrieval usually wins.
Compare context options with the context window comparison, then test retrieval on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Dimension | Long context only | RAG only | Retrieval plus long context |
|---|---|---|---|
| Corpus scale | Limited by window size | Large corpora | Large corpora with richer context |
| Prompt size | Large and fixed per document | Small and query-dependent | Tunable per question type |
| Citations | Not automatic | Chunk-level references | References plus expanded passages |
| Latency | Grows with input length | Generally lower | Middle ground |
| Best for | Whole-document reasoning | High-volume grounded Q&A | Mixed workloads |
Frequently asked questions
Is long context replacing RAG?
No. Long context solves small-corpus and whole-document tasks well, while retrieval scales to large corpora and provides citations. Most production systems use both.
Why not always send everything?
Prompt size affects latency and cost, and models do not always weight very long inputs evenly. Retrieval keeps context focused on what the question needs.
Does RAG work with long-context models?
Yes, and they pair well: retrieve candidates for relevance, then pass expanded or full passages to a long-context model for reasoning.
How do I choose for my corpus?
Measure. If documents fit comfortably in a window and questions require whole-document reasoning, long context is simple. If the corpus is large or growing, retrieval is the foundation.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.