RAG

What are the best reranking models for RAG?

A reranker rescoring the top candidates from first-stage retrieval is usually the cheapest quality gain in a RAG pipeline. The best choice depends on language, latency budget and whether you can host a cross-encoder or prefer a managed option. Plugsky includes optional cross-encoder reranking on RAG queries, so you can measure the gain first.

Key facts

RetrievalKeyword, vector and hybrid search with optional cross-encoder reranking
Query endpointPOST /v1/rag/query on a collection with top_k and reranking enabled
CitationsRanked chunks return with source attribution after reranking
First stageVector and hybrid candidate generation via POST /v1/embeddings
DeploymentManaged, VPC, on-prem and air-gapped options
Models30+ answer models; change the generator without touching the retriever
Pricing modelRAG queries share the flat self-serve pool; see the live pricing page
Product statusLive

TL;DR

  • Reranking is a second stage: retrieve many candidates, then score them more carefully.
  • Cross-encoders read query and document together, which usually beats vector similarity alone.
  • Measure nDCG or MRR before and after on your own questions before enabling it everywhere.
  • Reranking cannot fix candidates that the first-stage retriever never returned.
  • Plugsky offers optional reranking directly on RAG collection queries.

How it works, step by step

  1. Log retrieval quality today: recall@k and the rank of the correct chunk.
  2. Increase first-stage top_k to gather a wider candidate set.
  3. Enable reranking on the query and compare the top 5 results for the same questions.
  4. Measure the added latency per query at realistic concurrency.
  5. Tune the candidate count so quality gain stays worth the latency.
  6. Keep the reranker and the embedder in the same deployment boundary when data cannot leave your perimeter.
1Log retrievalquality today:recall@k and the2Increasefirst-stage top_kto gather a wider3Enable reranking onthe query andcompare the top 54Measure the addedlatency per queryat realistic5Tune the candidatecount so qualitygain stays worth6Keep the rerankerand the embedder inthe same deployment

Original data

POST /v1/rag/qQuery endpointVector and hybFirst stage30+ answer modModelsSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the reranker comparison →

Why reranking works

Vector search uses a bi-encoder: documents and queries are embedded separately, then compared with a distance function. That design is fast and scales to millions of chunks, but it is approximate. The embedding has to compress all possible meanings of a passage into one vector, so near-misses are common.

A reranker is a cross-encoder that reads the query and each candidate document together and produces a relevance score. It sees word order, negation and context that the bi-encoder flattened away. The result is a better ordering of a small candidate set, which directly improves what the generator gets to read.

What to compare between rerankers

Compare quality on your own query set first: rank-aware metrics such as nDCG and MRR show whether the right passage moved to the top. Then compare the operational properties: latency added per query, maximum input length, language coverage, and whether the model can run inside your deployment boundary.

Do not choose on a public leaderboard alone. A reranker that is marginally stronger on generic passages but adds 200ms and a second vendor to every request can be worse than a simpler option that keeps data and latency under your control.

Two-stage retrieval in practice

The standard pattern is retrieve 50 to 100 candidates from hybrid search, rerank them to a top 5 or 10, then pass only those chunks to the answer model. More candidates give the reranker more to work with, but each one costs scoring time. Tune the candidate count against your latency budget instead of assuming a fixed value.

Check the first stage separately. If the correct chunk never appears in the candidate list, reranking has nothing to promote; that is a signal to revisit chunking or the embedding model, not the reranker.

Reranking on Plugsky

Plugsky exposes reranking as an option on RAG collection queries, alongside keyword, vector and hybrid retrieval, and returns ranked chunks with source attribution. The same query can be tested with and without reranking, which makes the cost and quality trade-off measurable on your own data.

Use the reranker comparison to frame the evaluation, then validate on your corpus. Free-plan models plugsky-micro and plugsky-lite cover evaluation work, and the 14-day full-access trial covers larger tests. Current plans are on the live pricing page.

Honest comparison

ApproachPlugsky RAG queryManaged reranker APISelf-hosted cross-encoder
StagesOptional reranking on collection queriesSeparate call after retrievalYou host and call the model
OperationsOne API, no extra vendorSecond API, key and billGPU, serving stack and upgrades
LatencyAdds one scoring pass over candidatesAdds a network round tripHardware-dependent
LanguagesDepends on the configured reranker and embedderModel-dependentModel-dependent
Data pathStays inside your chosen deploymentLeaves your perimeterStays in your infrastructure

Frequently asked questions

What is the difference between an embedding model and a reranker?

The embedding model generates vectors for first-stage retrieval and is fast but approximate. The reranker scores query and document pairs directly and is slower but more precise, so it runs on a small candidate set.

How many candidates should I rerank?

Start with 50 to 100 candidates and tune down if latency matters more than the quality gain. The right number depends on corpus size, query difficulty and your response-time target.

Can reranking fix poor retrieval?

No. If the correct passage is never in the candidate set, reranking cannot surface it. Fix chunking or the embedding model first, then add reranking.

Does Plugsky have a standalone reranking endpoint?

Reranking is available as an option on RAG collection queries rather than a separate endpoint. Check the docs for the current capability matrix.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.