Key facts
| Retrieval | Keyword, vector and hybrid search with optional cross-encoder reranking |
| Citations | Every query returns ranked chunks with source attribution |
| Refusal | Prompts can require declining when the corpus lacks an answer |
| Chunking | 500-token chunks with 50-token overlap by default |
| Models | 30+ models; route grounding-critical answers to stronger models |
| Audit logs | Per-request model, tokens, latency, user and region |
| Evaluation | Fixed endpoints make groundedness testing repeatable |
| Product status | Live |
TL;DR
- Most hallucinations are retrieval failures; fix recall before tuning prompts.
- Pass only retrieved chunks and forbid outside knowledge in the prompt.
- Require per-claim citations and validate them against the retrieved set.
- Give the model an explicit way to say the answer is not in the corpus.
- Measure groundedness and refusal correctness as first-class metrics.
How it works, step by step
- Build a question set that includes deliberately unanswerable questions.
- Measure recall first: the correct chunk must reach the model.
- Tighten chunking and add reranking so context is precise, not just present.
- Prompt for context-only answers with per-claim citations.
- Validate every citation against the retrieved chunk ids and fail closed.
- Review answers that cite nothing or refuse incorrectly to find gaps.
- Re-run the evaluation after every corpus or prompt change.
Try it yourself
Open the AI citation checker →
Most hallucinations start in retrieval
When a model invents a policy or a number, the first question is whether the correct passage was in the context at all. If retrieval missed it, the model either guessed or blended unrelated chunks. Fixing recall through better chunking, hybrid search and reranking removes a large share of hallucinated answers before any prompt change.
Include unanswerable questions in your test set. They reveal whether the system can recognise absence, which is the behaviour that protects users from confident fabrication.
Constrain the answer, then verify it
Prompt design should make the boundary explicit: answer only from the provided chunks, cite the chunk for each claim, and state that the answer is unavailable when the chunks do not contain it. Vague instructions such as be accurate do little; a concrete refusal rule changes behaviour measurably.
Then verify programmatically. Parse citations, confirm each references a chunk that was actually retrieved, and treat unresolvable citations as failures. Optionally add a second pass that checks whether the cited chunk supports its claim. Validation is what turns a policy into a guarantee.
Reduce the opportunity to guess
Smaller, more precise context helps. A prompt stuffed with ten marginal chunks gives the model more material to blend into a plausible but unsupported statement. Retrieved chunks should be relevant, current and deduplicated; stale versions of the same document are a common source of contradictory answers.
For high-stakes answers, route generation to a stronger model and keep the retrieval settings identical, so the comparison isolates the model variable. Temperature settings and deterministic options also reduce variation, though they do not substitute for grounding.
Make it measurable on Plugsky
Plugsky queries return ranked chunks with source references, so retrieval quality and citation validity can be scored on the same run. Keep the corpus clean, use hybrid retrieval where exact terms matter, and enable reranking when candidates are close. Audit logs then capture model, tokens, latency, user and region for each request.
Test answers with the AI citation checker, then run the evaluation loop on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Failure | Typical cause | Fix | How to verify |
|---|---|---|---|
| Invented policy | Correct chunk never retrieved | Improve chunking and hybrid search | Recall on a golden set |
| Wrong number | Stale or duplicate chunks | Refresh corpus and deduplicate | Answer review against sources |
| Blended facts | Too many marginal chunks in context | Lower top_k and rerank | Citation support checks |
| Confident answer without evidence | No refusal path in the prompt | Explicit refusal rule | Unanswerable test questions |
| Broken citation | No validation step | Check citations against retrieved ids | Automated citation validation |
Frequently asked questions
Can hallucinations be eliminated completely?
No. Models can still err, which is why validation and human review matter for high-stakes answers. Good retrieval, strict grounding and citation checks reduce the frequency substantially.
Does a lower temperature stop hallucinations?
It reduces variation but does not ground the model. Retrieval quality and prompt constraints matter far more.
How do I test for hallucinations?
Use a golden set that includes unanswerable questions, then score groundedness, citation accuracy and refusal correctness on every pipeline change.
Should I use the biggest model?
Use a capable model for grounding-critical answers, but remember that retrieval failures dominate. A strong model cannot cite a passage that was never retrieved.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.