Verifying RAG Citations: Beyond Source Existence to Claim Support
A citation identifier only proves a source was retrieved, not that it supports the claim. This article shows how to decompose answers, locate supporting passages, and separate retrieval checks from groundedness checks using a hypothetical policy example.
On this page
The short answer
To verify RAG citations, first separate the mechanical check that a source identifier exists from the semantic check that the cited passage actually supports the claim. Break the answer into atomic claims, locate the exact supporting text in each cited source, and judge entailment. Flag missing evidence and conflicting sources. Use a structured checklist and be aware that automated groundedness metrics can miss subtle mismatches.
Separate a link from evidence
A citation label is a reference to check, not proof that the model retrieved or read the correct passage. First resolve the label against the actual retrieved context and confirm that it points to an available source chunk. Then ask whether that chunk supports the specific claim. In our hypothetical example, Source A allows returns of unused items within 14 days and Source B covers manufacturing defects for 12 months. A claim that all used items can be refunded is not supported by Source A, even if its identifier resolves correctly.
Break the answer into claims
To systematically verify support, decompose the answer into individual factual statements. The Microsoft prompt engineering article recommends instructing the model to cite the source for each claim, which makes this decomposition easier. For the example answer 'You can refund all used items within 30 days,' the atomic claims are 'refund all used items' and 'within 30 days.' Each claim may be supported by different sources, or none at all. Breaking the answer into claims prevents a situation where a partially supported answer is accepted as a whole. After decomposition, you can check each claim against its cited source independently.
Locate the supporting passage
For each claim, find the exact text in the cited source that is supposed to support it. In the example, the claim 'refund all used items' cites Source A. The text of Source A is 'Unused items can be returned within 14 days.' There is no mention of used items, so no supporting passage exists. The claim 'unused items can be returned within 14 days' citing Source A, however, directly matches the source text. This step requires reading the source content, not just checking that the source identifier appears. The evaluation article describes groundedness calculations that use natural language inference to determine if claims are based on context, which automates this matching to some extent.
Trace a supported and unsupported claim
In the hypothetical example, 'unused items can be returned within 14 days' is supported by Source A. 'All used items can be refunded' is not supported by the same passage. A correct citation target therefore does not establish entailment. Assess each claim separately, including conditions, quantities and time limits, and distinguish an unsupported extension from information that the context explicitly contradicts.
Handle absent evidence
When a claim has no citation or the cited source lacks relevant information, it is missing evidence. The prompt engineering article advises instructing the model to state what information is missing when the context does not fully answer the question. In the example, if the answer includes 'refund all used items' without any source supporting it, that claim is unsupported due to absent evidence. The reviewer should flag it and check whether the model should have responded with 'I don't know' or a clarification. The evaluation metric completeness measures whether all parts of the query are answered; a claim with absent evidence may make completeness appear higher than it should be while groundedness remains low.
Handle source disagreement
Retrieved chunks can contain contradictory information. The prompt engineering article explicitly instructs: 'If the provided sources contain conflicting information, present both perspectives and cite the respective sources.' For example, if Source A says returns within 14 days and Source B says 30 days, the answer should acknowledge both. If the answer silently picks one, it is a groundedness issue because it ignores part of the context. The evaluation article's correctness metric can be affected if the answer chooses the wrong source. When reviewing, check whether the answer transparently presents the conflict or resolves it without justification.
Record a practical review checklist
For each claim, record its citation, the exact supporting passage and your support verdict. Check missing or irrelevant sources separately from a passage that exists but does not support the claim. If sources disagree, record the conflict and whether the answer explains it. The checklist at the end of this guide summarizes these steps.
Understand automated-verifier limits
Automated groundedness checks can help identify unsupported claims, but a score is not a guarantee. A verifier may miss a changed condition, a negation or an incorrect time limit. Groundedness and factual correctness also answer different questions: an answer can faithfully repeat a false or outdated source and still be grounded in that source. Conversely, a correct fact from outside the supplied context may lack support in that context. Review important claim-passage pairs directly and check the source's reliability, date and applicability separately. The policy snippets in this guide are constructed examples, not results of an executed retrieval experiment.
Things to check
- Verify that each citation identifier matches a retrieved source chunk.
- For each claim, locate the exact passage in the cited source that is supposed to support it.
- Assess whether the claim is logically entailed by the passage, not just mentioned.
- Flag claims that lack any supporting passage as unsupported.
- When sources conflict, check if the answer acknowledges the conflict or silently picks one.
- Separate retrieval quality (source existence) from groundedness (semantic support).
- Use a structured checklist to review each claim-source pair.
- Be aware that automated groundedness metrics may miss subtle mismatches.
Where this applies
Automated groundedness metrics may not catch all semantic mismatches; human review is needed for high-stakes decisions. The example uses hypothetical policy snippets, not real retrieval results.