RAG

RAG or long context: which should you use?

Long context lets a model read many documents in one request, which suits whole-document reasoning and small corpora. RAG selects the relevant passages, which is cheaper at scale and provides citations. The two combine well: retrieve candidates with RAG, then use a longer context to reason across the passages that matter.

Key facts

RAG retrievalKeyword, vector and hybrid search with optional reranking
CitationsRanked chunks return with source attribution
IngestionDocuments chunked, embedded and indexed automatically per collection
Models30+ models with different context capabilities behind one API
Cost profileFlat self-serve plans; no per-token billing pressure on prompt size
LatencySmaller prompts generally produce faster first tokens
DeploymentManaged, VPC, on-prem and air-gapped options
Product statusLive

TL;DR

  • Long context is simple but scales poorly as the corpus grows.
  • RAG narrows the input and keeps answers traceable to sources.
  • Attention can thin out across very long inputs, so more text is not always better.
  • Whole-document tasks favour long context; large corpora favour retrieval.
  • Combine them: retrieve candidates, then reason over a richer context.

How it works, step by step

  1. Estimate the corpus size and how often it changes.
  2. Identify whether questions need one passage or reasoning across a whole document.
  3. Test answering with the full document in context on a sample of real questions.
  4. Test the same questions with retrieval and compare quality and latency.
  5. Check citation requirements: retrieval gives references, long context does not by default.
  6. Model the cost of sending full documents versus retrieved passages at your volume.
  7. Choose a hybrid design where retrieval selects and long context reasons.
1Estimate the corpussize and how oftenit changes.2Identify whetherquestions need onepassage or3Test answering withthe full documentin context on a4Test the samequestions withretrieval and5Check citationrequirements:retrieval gives6Model the cost ofsending fulldocuments versus

Try it yourself

Open the context window comparison →

What long context is good at

Long context is the simplest architecture: put the documents in the prompt and ask the question. It works well when the corpus is small, when the task genuinely needs the whole document such as summarising a report or comparing two contracts, and when latency is less important than depth of reasoning.

It also removes retrieval risk. There is no chance of missing a relevant passage because the retriever failed, provided the entire document fits within the window along with the instructions and the answer.

Where retrieval still wins

RAG scales to corpora that will never fit in a context window and keeps prompt size proportional to the question rather than the library. That matters for latency and for cost on providers that bill by token. It also provides citations by construction: the answer references the chunks that were retrieved, so users can verify the evidence.

Practical experience also shows that models do not always use very long inputs evenly; information buried in the middle of a large context can be under-weighted. Retrieval puts the relevant passage in a small, focused context where attention is concentrated.

The hybrid pattern

The strongest designs borrow from both. Use retrieval to select a candidate set, then expand each hit to its surrounding section so the model sees complete arguments rather than fragments. For whole-document questions, retrieve the document and pass it in full, using retrieval to decide which documents deserve the space.

This keeps the benefits of traceability while giving the model richer material for reasoning. It also gives you a dial: widen the context for hard questions, narrow it for high-volume ones.

Deciding with measurements

Run the same question set through both approaches and compare answer quality, citation accuracy, latency and cost. Include questions whose answer spans two documents, because that is where long context often shines, and questions over a large corpus, where retrieval usually wins.

Compare context options with the context window comparison, then test retrieval on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

DimensionLong context onlyRAG onlyRetrieval plus long context
Corpus scaleLimited by window sizeLarge corporaLarge corpora with richer context
Prompt sizeLarge and fixed per documentSmall and query-dependentTunable per question type
CitationsNot automaticChunk-level referencesReferences plus expanded passages
LatencyGrows with input lengthGenerally lowerMiddle ground
Best forWhole-document reasoningHigh-volume grounded Q&AMixed workloads

Frequently asked questions

Is long context replacing RAG?

No. Long context solves small-corpus and whole-document tasks well, while retrieval scales to large corpora and provides citations. Most production systems use both.

Why not always send everything?

Prompt size affects latency and cost, and models do not always weight very long inputs evenly. Retrieval keeps context focused on what the question needs.

Does RAG work with long-context models?

Yes, and they pair well: retrieve candidates for relevance, then pass expanded or full passages to a long-context model for reasoning.

How do I choose for my corpus?

Measure. If documents fit comfortably in a window and questions require whole-document reasoning, long context is simple. If the corpus is large or growing, retrieval is the foundation.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.