RAG

How do you evaluate a RAG pipeline?

Evaluate RAG in two layers. For retrieval, measure whether the correct chunks appear in the top k using recall, MRR or nDCG on a golden question set. For generation, measure groundedness, answer relevance and citation accuracy against those chunks. Plugsky's fixed endpoints and cited chunks make both layers straightforward to test.

Key facts

Retrieval signalsRanked chunks with scores and source attribution are returned per query
Retrieval modesKeyword, vector and hybrid search with optional reranking
GenerationChat completions across 30+ models for answer composition
CitationsSource references let answers be checked against retrieved text
Audit logsPer-request model, tokens, latency, user and region
Free tierplugsky-micro and plugsky-lite for evaluation runs, no card
Trial14-day full-access trial available
Product statusLive

TL;DR

  • Score retrieval and generation separately so failures point to one stage.
  • Recall@k answers whether the right chunk was reachable at all.
  • MRR and nDCG show whether the right chunk ranked near the top.
  • Groundedness and citation accuracy catch answers that overshoot the sources.
  • Re-run the same golden set after every corpus, chunking or model change.

How it works, step by step

  1. Collect 50 to 200 real questions with the documents and chunks that answer them.
  2. Run retrieval and record whether the correct chunk appears in the top k.
  3. Track rank-aware metrics to see where the right chunk lands, not just if it appears.
  4. Generate answers from the retrieved chunks and score groundedness and citation accuracy.
  5. Split results by query type to find systematic weaknesses.
  6. Automate the run so it executes after every ingestion or configuration change.
  7. Keep the set current: retire questions that no longer reflect real traffic.
1Collect 50 to 200real questions withthe documents and2Run retrieval andrecord whether thecorrect chunk3Track rank-awaremetrics to seewhere the right4Generate answersfrom the retrievedchunks and score5Split results byquery type to findsystematic6Automate the run soit executes afterevery ingestion or

Try it yourself

Open the prompt diff and evaluator →

Start with a golden question set

Evaluation begins with a fixed set of questions and known correct answers, written from real usage rather than invented examples. Each item should record the question, the expected document or chunk, and optionally an expected answer. Fifty to two hundred items is usually enough to detect regressions if the set spans the query types you actually receive.

Keep the set in version control alongside the pipeline configuration. When it changes, the change should be deliberate and reviewed, because the set becomes the definition of quality for the system.

Measuring retrieval quality

Retrieval metrics answer a narrow, powerful question: did the right content reach the model? Recall@k measures whether the correct chunk appears in the top k. MRR measures how high the first correct result ranks. nDCG weighs the whole ordering, which matters when several chunks are relevant.

Measure at the k you actually pass to the model, and remember the first-stage candidate count separately if reranking is enabled. If recall is low, the problem is chunking, embedding or filters. If recall is high but answers are wrong, the problem is reranking or generation.

Measuring answer quality

Answer metrics check whether the generated text stays within the evidence. Groundedness asks whether every claim is supported by the retrieved chunks. Answer relevance asks whether the response addresses the question. Citation accuracy asks whether the cited chunks actually support the claims they are attached to. Refusal correctness is the fourth: when the corpus lacks the answer, saying so is the right outcome.

Automated scoring with a capable model can approximate these judgments at scale, but calibrate it against human review on a sample before trusting the numbers.

Running evaluation as a habit

Wire the golden set into a script that runs after ingestion jobs, configuration changes and model switches, and store the results so trends are visible. Track retrieval and generation metrics side by side with latency, because improvements in one dimension can hide regressions in another.

Use the same endpoints in evaluation and production: POST /v1/rag/query for retrieval and chat completions for answers, with citations preserved. Structure prompt changes with the prompt diff and evaluator, then run evaluation on plugsky-micro and plugsky-lite for free or the 14-day full-access trial for larger sets. Current plans are on the live pricing page.

Honest comparison

LayerMetricWhat it catchesWhat it misses
RetrievalRecall@kCorrect chunk never retrievedRanking quality
RetrievalMRR and nDCGCorrect chunk ranked lowAnswer quality
GenerationGroundednessClaims unsupported by contextMissing information in the corpus
GenerationCitation accuracyWrong or invented sourcesRetrieval coverage
SystemRefusal correctnessConfident answers without evidenceLatency and cost

Frequently asked questions

How many test questions do I need?

Fifty to two hundred well-chosen questions covering your real query types usually reveal regressions. More coverage matters more than a large number of similar items.

Should I use an LLM to judge answers?

It is practical for scale, but calibrate the judge against human review on a sample and keep the rubric explicit, especially for groundedness and citation accuracy.

What is a good recall@k?

Any target depends on your domain and risk tolerance. What matters is having a baseline you measure consistently and improving it without hurting answer quality.

How often should evaluation run?

After every ingestion batch, chunking change, model switch or prompt revision. Automated runs make this cheap and keep quality visible over time.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.