Key facts
| Retrieval signals | Ranked chunks with scores and source attribution are returned per query |
| Retrieval modes | Keyword, vector and hybrid search with optional reranking |
| Generation | Chat completions across 30+ models for answer composition |
| Citations | Source references let answers be checked against retrieved text |
| Audit logs | Per-request model, tokens, latency, user and region |
| Free tier | plugsky-micro and plugsky-lite for evaluation runs, no card |
| Trial | 14-day full-access trial available |
| Product status | Live |
TL;DR
- Score retrieval and generation separately so failures point to one stage.
- Recall@k answers whether the right chunk was reachable at all.
- MRR and nDCG show whether the right chunk ranked near the top.
- Groundedness and citation accuracy catch answers that overshoot the sources.
- Re-run the same golden set after every corpus, chunking or model change.
How it works, step by step
- Collect 50 to 200 real questions with the documents and chunks that answer them.
- Run retrieval and record whether the correct chunk appears in the top k.
- Track rank-aware metrics to see where the right chunk lands, not just if it appears.
- Generate answers from the retrieved chunks and score groundedness and citation accuracy.
- Split results by query type to find systematic weaknesses.
- Automate the run so it executes after every ingestion or configuration change.
- Keep the set current: retire questions that no longer reflect real traffic.
Try it yourself
Open the prompt diff and evaluator →
Start with a golden question set
Evaluation begins with a fixed set of questions and known correct answers, written from real usage rather than invented examples. Each item should record the question, the expected document or chunk, and optionally an expected answer. Fifty to two hundred items is usually enough to detect regressions if the set spans the query types you actually receive.
Keep the set in version control alongside the pipeline configuration. When it changes, the change should be deliberate and reviewed, because the set becomes the definition of quality for the system.
Measuring retrieval quality
Retrieval metrics answer a narrow, powerful question: did the right content reach the model? Recall@k measures whether the correct chunk appears in the top k. MRR measures how high the first correct result ranks. nDCG weighs the whole ordering, which matters when several chunks are relevant.
Measure at the k you actually pass to the model, and remember the first-stage candidate count separately if reranking is enabled. If recall is low, the problem is chunking, embedding or filters. If recall is high but answers are wrong, the problem is reranking or generation.
Measuring answer quality
Answer metrics check whether the generated text stays within the evidence. Groundedness asks whether every claim is supported by the retrieved chunks. Answer relevance asks whether the response addresses the question. Citation accuracy asks whether the cited chunks actually support the claims they are attached to. Refusal correctness is the fourth: when the corpus lacks the answer, saying so is the right outcome.
Automated scoring with a capable model can approximate these judgments at scale, but calibrate it against human review on a sample before trusting the numbers.
Running evaluation as a habit
Wire the golden set into a script that runs after ingestion jobs, configuration changes and model switches, and store the results so trends are visible. Track retrieval and generation metrics side by side with latency, because improvements in one dimension can hide regressions in another.
Use the same endpoints in evaluation and production: POST /v1/rag/query for retrieval and chat completions for answers, with citations preserved. Structure prompt changes with the prompt diff and evaluator, then run evaluation on plugsky-micro and plugsky-lite for free or the 14-day full-access trial for larger sets. Current plans are on the live pricing page.
Honest comparison
| Layer | Metric | What it catches | What it misses |
|---|---|---|---|
| Retrieval | Recall@k | Correct chunk never retrieved | Ranking quality |
| Retrieval | MRR and nDCG | Correct chunk ranked low | Answer quality |
| Generation | Groundedness | Claims unsupported by context | Missing information in the corpus |
| Generation | Citation accuracy | Wrong or invented sources | Retrieval coverage |
| System | Refusal correctness | Confident answers without evidence | Latency and cost |
Frequently asked questions
How many test questions do I need?
Fifty to two hundred well-chosen questions covering your real query types usually reveal regressions. More coverage matters more than a large number of similar items.
Should I use an LLM to judge answers?
It is practical for scale, but calibrate the judge against human review on a sample and keep the rubric explicit, especially for groundedness and citation accuracy.
What is a good recall@k?
Any target depends on your domain and risk tolerance. What matters is having a baseline you measure consistently and improving it without hurting answer quality.
How often should evaluation run?
After every ingestion batch, chunking change, model switch or prompt revision. Automated runs make this cheap and keep quality visible over time.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.