Key facts
| Retrieval | Keyword, vector and hybrid search with optional cross-encoder reranking |
| Query endpoint | POST /v1/rag/query on a collection with top_k and reranking enabled |
| Citations | Ranked chunks return with source attribution after reranking |
| First stage | Vector and hybrid candidate generation via POST /v1/embeddings |
| Deployment | Managed, VPC, on-prem and air-gapped options |
| Models | 30+ answer models; change the generator without touching the retriever |
| Pricing model | RAG queries share the flat self-serve pool; see the live pricing page |
| Product status | Live |
TL;DR
- Reranking is a second stage: retrieve many candidates, then score them more carefully.
- Cross-encoders read query and document together, which usually beats vector similarity alone.
- Measure nDCG or MRR before and after on your own questions before enabling it everywhere.
- Reranking cannot fix candidates that the first-stage retriever never returned.
- Plugsky offers optional reranking directly on RAG collection queries.
How it works, step by step
- Log retrieval quality today: recall@k and the rank of the correct chunk.
- Increase first-stage top_k to gather a wider candidate set.
- Enable reranking on the query and compare the top 5 results for the same questions.
- Measure the added latency per query at realistic concurrency.
- Tune the candidate count so quality gain stays worth the latency.
- Keep the reranker and the embedder in the same deployment boundary when data cannot leave your perimeter.
Original data
Try it yourself
Open the reranker comparison →
Why reranking works
Vector search uses a bi-encoder: documents and queries are embedded separately, then compared with a distance function. That design is fast and scales to millions of chunks, but it is approximate. The embedding has to compress all possible meanings of a passage into one vector, so near-misses are common.
A reranker is a cross-encoder that reads the query and each candidate document together and produces a relevance score. It sees word order, negation and context that the bi-encoder flattened away. The result is a better ordering of a small candidate set, which directly improves what the generator gets to read.
What to compare between rerankers
Compare quality on your own query set first: rank-aware metrics such as nDCG and MRR show whether the right passage moved to the top. Then compare the operational properties: latency added per query, maximum input length, language coverage, and whether the model can run inside your deployment boundary.
Do not choose on a public leaderboard alone. A reranker that is marginally stronger on generic passages but adds 200ms and a second vendor to every request can be worse than a simpler option that keeps data and latency under your control.
Two-stage retrieval in practice
The standard pattern is retrieve 50 to 100 candidates from hybrid search, rerank them to a top 5 or 10, then pass only those chunks to the answer model. More candidates give the reranker more to work with, but each one costs scoring time. Tune the candidate count against your latency budget instead of assuming a fixed value.
Check the first stage separately. If the correct chunk never appears in the candidate list, reranking has nothing to promote; that is a signal to revisit chunking or the embedding model, not the reranker.
Reranking on Plugsky
Plugsky exposes reranking as an option on RAG collection queries, alongside keyword, vector and hybrid retrieval, and returns ranked chunks with source attribution. The same query can be tested with and without reranking, which makes the cost and quality trade-off measurable on your own data.
Use the reranker comparison to frame the evaluation, then validate on your corpus. Free-plan models plugsky-micro and plugsky-lite cover evaluation work, and the 14-day full-access trial covers larger tests. Current plans are on the live pricing page.
Honest comparison
| Approach | Plugsky RAG query | Managed reranker API | Self-hosted cross-encoder |
|---|---|---|---|
| Stages | Optional reranking on collection queries | Separate call after retrieval | You host and call the model |
| Operations | One API, no extra vendor | Second API, key and bill | GPU, serving stack and upgrades |
| Latency | Adds one scoring pass over candidates | Adds a network round trip | Hardware-dependent |
| Languages | Depends on the configured reranker and embedder | Model-dependent | Model-dependent |
| Data path | Stays inside your chosen deployment | Leaves your perimeter | Stays in your infrastructure |
Frequently asked questions
What is the difference between an embedding model and a reranker?
The embedding model generates vectors for first-stage retrieval and is fast but approximate. The reranker scores query and document pairs directly and is slower but more precise, so it runs on a small candidate set.
How many candidates should I rerank?
Start with 50 to 100 candidates and tune down if latency matters more than the quality gain. The right number depends on corpus size, query difficulty and your response-time target.
Can reranking fix poor retrieval?
No. If the correct passage is never in the candidate set, reranking cannot surface it. Fix chunking or the embedding model first, then add reranking.
Does Plugsky have a standalone reranking endpoint?
Reranking is available as an option on RAG collection queries rather than a separate endpoint. Check the docs for the current capability matrix.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.