RAG

How do you implement citations in a RAG pipeline?

RAG citations work when three layers agree: retrieval returns chunks with source references, the prompt instructs the model to cite only those chunks, and the application validates and renders the references. Plugsky queries return ranked chunks with source attribution such as file and page, which gives you the identifiers needed for every step.

Key facts

CitationsEvery RAG query returns ranked chunks with source attribution
Reference fieldsChunks carry references such as file and page
RetrievalKeyword, vector and hybrid search with optional reranking
GenerationAny of 30+ models can compose the cited answer on the chat endpoint
RefusalGrounding instructions can require the model to decline without context
Audit logsPer-request model, tokens, latency, user and region are recorded
DeploymentManaged, VPC, on-prem and air-gapped options
Product statusLive

TL;DR

  • Citations begin at retrieval: a chunk without a source reference cannot be cited.
  • Number chunks before prompting and require the model to use those numbers.
  • Validate every citation against the retrieved set before showing the answer.
  • Render sources close to the claim so users can verify quickly.
  • Log citations with the answer for evaluation and audit.

How it works, step by step

  1. Ensure every retrieved chunk carries a stable id plus file and page references.
  2. Number the chunks in the prompt and pass the text with those numbers.
  3. Instruct the model to cite chunk numbers inline and never invent sources.
  4. Parse the answer, extract citations and check each one against the retrieved set.
  5. Render citations as links or tooltips anchored to the supporting sentence.
  6. Require refusal when no retrieved chunk supports the answer.
  7. Track citation coverage and accuracy over time as part of evaluation.
1Ensure everyretrieved chunkcarries a stable id2Number the chunksin the prompt andpass the text with3Instruct the modelto cite chunknumbers inline and4Parse the answer,extract citationsand check each one5Render citations aslinks or tooltipsanchored to the6Require refusalwhen no retrievedchunk supports the

Try it yourself

Open the AI citation checker →

Citations start at retrieval

An answer can only be cited if the retrieved chunk carries an identifier and a human-readable source. Plugsky RAG queries return ranked chunks with source attribution such as file and page, so the pipeline has the raw material from the start. Preserve those references through your own data structures rather than collapsing chunks into a single context string.

Number the chunks when building the prompt, and keep a map from number to source in your application. The model then references numbers, and your code resolves them back to documents. That indirection keeps the prompt compact and makes validation possible.

Prompting for grounded citations

State the rules explicitly: answer only from the provided chunks, cite the chunk number after each claim, and say the answer is not available when the chunks do not contain it. Providing an explicit refusal path matters, because a model told to always answer will happily fill gaps from its own knowledge.

Ask for citations per claim rather than one list at the end. A per-claim citation makes it obvious which sentence is supported by which chunk, and it exposes answers that are only partly grounded. Keep the instruction short; long prompt rules dilute each other.

Validating before display

Parse citation markers from the generated answer, then verify each referenced chunk number exists in the retrieved set. Any citation that does not resolve should be treated as a failed answer, not silently dropped. For higher assurance, add a second pass that checks whether each cited chunk actually supports its claim, or route low-confidence answers to review.

Track two metrics as part of evaluation: citation coverage, meaning how many claims carry a citation, and citation accuracy, meaning how many citations genuinely support the claim. Both catch regressions that generic answer-quality reviews miss.

Rendering and operations

In the interface, attach sources to the claim they support: inline markers that open the passage, or a source list grouped by document with page numbers. Users trust what they can check, and a citation they cannot inspect is only decoration. Keep the original text available so the cited passage can be shown verbatim.

Log the answer together with the retrieved chunk ids and the model used, so citation issues can be reproduced later. Test the implementation with the AI citation checker, then start free with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

LayerWith citationsWithout citationsUnvalidated citations
Retrieval outputChunks with ids and source referencesPlain text blocksChunks with references
PromptNumbered chunks and citation rulesAnswer from contextAsks for sources loosely
ValidationEvery citation checked against the setNot possibleNone
User trustClaims can be verifiedAnswer must be taken on faithFalse confidence when markers are wrong
AuditSources logged with the answerAnswer onlyUnreliable trail

Frequently asked questions

Do I need a special model for citations?

No. Any capable chat model can follow citation instructions when the prompt numbers the chunks and states the rules clearly.

What if the model invents a citation?

Validate every citation against the retrieved chunk set and treat unresolvable citations as failures. This catches fabrications before they reach users.

Should citations be inline or in a list?

Inline markers tied to specific claims are more useful for verification. A document list is a helpful complement, not a replacement.

What does Plugsky return for citations?

RAG queries return ranked chunks with source attribution such as file and page references, which your application can map to UI elements.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.