Blog/RAG & retrieval

Reducing LLM hallucinations in production: grounding, citations and knowing when to say no

Cut hallucinations in production RAG: citation APIs, abstention, claim-level checks, faithfulness metrics, source UI, and the failures that still slip through.

··13 min read

  • Hallucinations
  • RAG
  • Citations
  • Grounding
  • Faithfulness
Diagram: a retrieval step feeds an evidence gate, a cited answer and a claim verifier, ending in an answer with sources, with abstain and flag paths branching off.

Key takeaways

  • Retrieval does not remove hallucinations. A Stanford study found leading RAG legal tools still hallucinated between 17% and 33% of the time, often by citing a real source that does not support the claim.
  • Treat grounding as a pipeline, not a prompt: an evidence gate before generation, native citations during generation, claim-level verification after, and a UI that shows the evidence.
  • Abstention is a product state, not an error message. Design the "I cannot answer this from the documents" path, measure it, and give it a next step.
  • Citation APIs make pointers valid, not conclusions true. Anthropic guarantees valid document pointers; whether the passage supports the claim still has to be checked, because up to 57% of citations in one study were post-rationalised.
  • Measure faithfulness (claims supported by retrieved context) separately from answer correctness, on a test set that includes questions the documents cannot answer.

Every team that ships a retrieval-augmented assistant goes through the same stages. First comes the demo, where it answers beautifully. Then comes the first real user, who finds a confident, fluent, wrong answer on day two. Then someone says "we already use RAG, why does it still make things up?"

The honest answer is that retrieval changes the odds, not the nature of the system. A language model still produces plausible text; retrieval only gives it better material and a chance to show its work. This article is how I would build the part around the model so that wrong answers become rarer, visible and recoverable. It builds on my posts about chunking, hybrid search and reranking and evals for product features; here the focus is what happens after the passages are retrieved.

What counts as a hallucination in a RAG system

In a closed-book chatbot, a hallucination is a false statement. In a RAG system there are two distinct failures, and they need different fixes:

  • Wrong content: the answer states something false, either because retrieval returned the wrong or outdated passage, or because the model ignored the context and answered from memory.
  • Misgrounded content: the answer cites a source, but the source does not say that. The Stanford and Yale team that evaluated commercial legal research tools defines it precisely: a response is hallucinated if it is incorrect or misgrounded, meaning the answer "falsely asserts that a source supports a statement".

The second type is the dangerous one. The same study tested tools from LexisNexis and Thomson Reuters that were marketed with claims like "hallucination-free" citations, and found they hallucinated between 17% and 33% of the time. The authors note that checking such errors means clicking through, reading the source and comparing it to the claim, which is exactly the work users expect the tool to have done for them.

Why do models guess at all? The paper "Why Language Models Hallucinate" argues that training and evaluation procedures reward guessing over acknowledging uncertainty, because a model that guesses scores better on benchmarks than one that abstains. That matters for you in a practical way: the default behaviour of the model is to answer, so abstention has to be designed into the system, not hoped for.

The pipeline: four gates, not one prompt

I think of grounding as a sequence of gates. Each one catches something the previous one cannot, and each one has a cost in latency and complexity.

A grounded answer pipelineFive steps in a row: retrieve, gate, generate with citations, verify claims, show the answer with sources. From the gate, a dashed path leads to abstaining when the evidence is too weak. From the verifier, a dashed path leads to dropping or flagging unsupported claims. A bar below says every step is logged and feeds evaluations.A grounded answer pipelineRetrievehybrid + rerankGateenough evidence?Generatewith citationsVerifyclaim by claimShowanswer + sourcesAbstainsay what is missingDrop or flagunsupported claimsLog query, chunks, answer and verdicts: they become your eval setEach gate removes a different failure mode.
Grounding is a chain of checks. Each stage can stop or downgrade the answer, and each stage is measured.
  1. Retrieve well. Hybrid search and reranking decide whether the right passage even reaches the model. Most "hallucinations" I debug are retrieval misses where the model did what it could with the wrong material.
  2. Gate on evidence. If the best passages are weak, do not generate; abstain and say what is missing.
  3. Generate with citations. Use a native citation mechanism so every claim carries a pointer into a document you supplied.
  4. Verify claim by claim. Check that each cited passage supports the sentence it is attached to, and drop or flag what fails.

Then comes the display step, which is also a control: the interface decides whether a reader can check the answer in five seconds or has to trust it.

Step one: constrain the model to what you gave it

Anthropic's guide to reducing hallucinations lists a handful of basic techniques, and they are cheap enough that I would use all of them by default:

  • Allow "I don't know". Explicitly give the model permission to admit uncertainty. Anthropic says this simple technique can drastically reduce false information.
  • Quote first for long documents. For documents over roughly 20,000 tokens, ask the model to extract word-for-word quotes first and base its answer on those quotes only.
  • Restrict external knowledge. Instruct the model to use only the provided documents and not its general knowledge.
  • Verify after drafting. Ask the model to find a supporting quote for each claim and to retract any claim it cannot support.

The same guide lists best-of-N comparison (run the prompt several times and treat disagreement as a warning) and iterative refinement as advanced options, and it ends with a caveat I want to repeat: these techniques significantly reduce hallucinations but do not eliminate them, and critical information still needs validation.

My practical addition: treat prompt rules as a weak layer. A prompt asks the model to behave; a gate or a verifier checks that it did. Use the prompt to raise the baseline and the later stages to catch the remainder.

Step two: use a citation API instead of asking nicely

You can prompt a model to write "[1]" after sentences, and for prototypes that works. In production, the pointer is the problem: models invent plausible source names, mis-number them, or quote text that is not in the document. Native citation features move that work out of free text and into the API response.

Anthropic: citations and search results

Anthropic's citations feature works on three document types, and the citation format follows the type:

Document typeChunkingCitation points to
Plain textSentencesCharacter indices (0-indexed)
PDFSentencesPage numbers (1-indexed)
Custom contentNone added: your blocks are used as-isBlock indices (0-indexed)

For RAG, the documentation's own advice is to put each retrieved chunk into a plain text document, or to use `search_result` content blocks, which carry a source and a title and can be returned from your own search tools or placed directly in the user message. Citations then appear on the text blocks that draw on your content, without special prompting. In my experience this chunk-per-document approach is also the cleanest way to keep your own chunk IDs traceable in the answer.

The details that matter in production, all from the documentation:

  • Valid pointers. Because the API parses citations and extracts `cited_text` itself, citations are guaranteed to contain valid pointers to the documents you provided. That removes fabricated references, not misread ones.
  • Cost. Enabling citations slightly increases input tokens, but `cited_text` does not count toward output tokens, so it can be cheaper than prompting the model to quote.
  • Caching. The source documents can be cached with `cache_control`; the citation blocks in responses cannot. See my post on prompt caching and routing for when this pays off.
  • Streaming. Citations arrive as `citations_delta` events, one citation per event, attached to the current text block.
  • Limits. Citations must be enabled on all or none of the documents in a request, only text is citable (scanned PDFs without extractable text are not), and combining citations with structured outputs returns a 400 error.

When the feature launched, Anthropic reported that its internal evaluations showed built-in citations outperforming most custom implementations by up to 15% in recall accuracy, and a customer, Endex, said source hallucinations and formatting issues fell from 10% to 0%. Those are vendor-reported figures for specific setups; I would use them as a reason to test the feature, not as a number to expect.

Other providers

Anthropic is not alone. OpenAI's web search tool returns `url_citation` annotations with a URL, title and location in the text, and its documentation requires that inline citations be clearly visible and clickable in the user interface when you show web results. Cohere's chat API returns citation objects with start and end positions, the cited text and the source documents behind it. The shared idea is the same: the model's claim and its evidence travel together as structured data, and your UI has to respect that.

One structural caveat applies to all of them: the model still decides which passage to point to. A citation tells you what the model attached to a sentence, not that the sentence follows from it.

Step three: design the "I cannot answer that" path

If the model's default is to answer, you need two places to interrupt that default.

A retrieval gate before generation. Look at the retrieval result: no passages above a similarity or reranker threshold, a large gap between what was asked and what was found, or contradictory top passages. In these cases skip the model call entirely and respond with what is missing and what the user can do. This is cheaper than generating and also the most reliable abstention, because it does not depend on the model's self-assessment. Calibrate the threshold on real queries, not by feel.

A permission in the prompt. Tell the model that stating "the documents do not contain this" is a valid and preferred answer when the evidence is missing. Without that permission, the benchmark-style incentive to guess is still operating.

Design the abstention response as part of the product:

  • Say what was searched and what was not found, instead of a bare "I don't know".
  • Offer a next step: rephrase, narrow the scope, search another source, or hand over to a person.
  • Allow partial answers: answer the supported part and mark the unsupported part explicitly.
  • Count it. The abstention rate is a metric with a healthy range; zero means the system is guessing, and too high means the gate is too strict or retrieval is poor.

Step four: verify claims after generation

The most reliable pattern I know for catching misgrounded answers is to decompose the answer into claims and check each one against its evidence. It is the same idea as the faithfulness metric in Ragas, which splits a response into individual statements, checks whether each can be inferred from the retrieved context, and computes the share of supported claims.

You can build this yourself with a second model call, or use a managed checker:

OptionWhat it doesNotable limits (per documentation)
Prompted self-check (Anthropic guide)Model finds a supporting quote per claim, retracts the restSame model family judging its own draft; extra call
Google check grounding APISplits the answer into claims, returns a 0 to 1 support score, citations and optional per-claim scoresAnswer up to 4,096 tokens, up to 200 facts; partial truths count as ungrounded; documented as under 500 ms
Amazon Bedrock contextual grounding checkScores grounding and relevance against a source and query; blocks below your thresholdNot for conversational QA; source up to 100,000 characters; on streaming, irrelevance may only be flagged after the response is sent

Three design decisions come up every time:

  1. What happens on failure? Options are to regenerate with the failing claim removed, to drop the sentence, to keep it with an "unverified" marker, or to abstain on the whole answer. For high-stakes domains I prefer dropping or marking over silent regeneration, because users should see that something was removed.
  2. Where does it run? A blocking verifier adds latency before the first token is shown if you wait for it. For streaming UIs, stream the draft with citations and update each sentence's state as verdicts arrive, or verify before streaming for the few flows where wrong answers are costly. My post on streaming LLM features in Nuxt covers the transport side.
  3. Who checks the checker? A verifier is a model or a classifier and it makes mistakes. Label a few hundred verdicts by hand and track the verifier's own precision and recall, otherwise you have moved trust from one unmeasured component to another.

Technique versus effect

This is how I rank the techniques by what they actually address. I deliberately give no percentage per row: the effect depends on your corpus, your queries and your model, and the only numbers I would trust are the ones you measure on your own test set.

TechniqueFailure it targetsCostWhere it still fails
Better retrieval (hybrid, rerank)Wrong or missing passagesEngineering time; some latencyCorpus gaps, stale documents
"Only use the documents" prompt rulesAnswers from model memoryAlmost noneA prompt is a request, not a guarantee
Permission to say "I don't know"Forced guessingNoneOver-abstention if not tested
Retrieval evidence gateGenerating from weak evidenceOne threshold to calibrateStrong but irrelevant passages pass it
Quote-first extractionParaphrase drift on long documentsExtra tokens or a second stepQuotes can still be misread
Native citationsFabricated or invalid referencesSlightly more input tokensValid pointer, unsupported claim
Claim-level verificationMisgrounded claimsExtra call and latencyVerifier errors; multi-hop reasoning
Source-first UIUnchecked trustDesign and frontend workUsers who never click

Measuring faithfulness

Without measurement, every change in this article is a belief. I would track four numbers on a fixed evaluation set, and run them in CI the way you run unit tests (see evals for product features):

  • Faithfulness: the share of claims supported by the retrieved context, as in the Ragas definition. It is separate from correctness: a faithful answer from a wrong document is wrong.
  • Citation precision and recall: of the citations shown, how many truly support the sentence; of the claims made, how many have a supporting citation.
  • Abstention quality: on questions the corpus cannot answer, how often does the system decline; on answerable questions, how often does it decline wrongly.
  • Answer correctness against a reference answer, so that faithfulness is not your only signal.

Build the set from three groups: answerable questions with known passages, unanswerable questions, and adversarial ones (outdated information, near-duplicate documents, questions that tempt the model to use its own knowledge). Most teams skip the unanswerable group, and that is why their abstention behaviour is never tested.

Showing sources in the interface

The interface is the last line of defence and the only one the user sees. These are the patterns I would use:

  • Inline numbered markers next to the sentence they support, not one block of links at the end. Per-claim attachment is what makes checking possible.
  • Preview on hover or tap showing the cited passage, with the document title. This is where `cited_text` is useful, and it makes a five-second check realistic.
  • Deep links that open the source at the cited location (page number, anchor or highlighted range), not just at the top of a 60-page PDF.
  • Visible verification state per sentence or per answer: verified, unverified, or removed. If the verifier dropped something, say so.
  • A designed abstention state with next steps, as above, styled as a normal outcome rather than an error.
  • Citations that are actually clickable and visible. OpenAI's documentation makes this a requirement for web results, and it is a good rule everywhere.

One warning from the Stanford study applies to design: real, authoritative-looking citations make a wrong answer more convincing. Do not let the presence of a footnote signal more certainty than your verification supports. If your verifier did not run, do not render the "verified" badge.

Also mind the engineering side: responses with citations are no longer a plain text stream. Simon Willison pointed out when the feature launched that this forces an abstraction for responses that are annotated chunks rather than text. Plan your streaming protocol and your message storage for structured segments from the start.

Where hallucination still slips through

After all four gates, these are the failures I would still expect, and what I would do about each:

  • Retrieval misses that look like answers. A partially relevant passage passes the gate and the model fills the gap. Mitigation: reranker thresholds and an explicit "does this passage answer the question" check.
  • Wrong or stale sources. The answer is faithful to an outdated document. Mitigation: document dates and versions in the metadata, shown in the UI, and freshness filters.
  • Misgrounded citations. The valid pointer, unsupported claim problem. Mitigation: claim-level verification and sampled human review.
  • Reasoning across passages. Totals, comparisons and multi-hop conclusions are not literally in any one passage, so verifiers struggle with them. Mitigation: compute numbers with code and show the inputs.
  • Poisoned documents. If a retrieved document contains instructions, the model may follow them. Grounding on untrusted text is a security question; see prompt injection and the lethal trifecta.
  • Format constraints. Structured outputs and citations cannot be combined on the Anthropic API today, so an extraction pipeline needs a different grounding strategy, for example a verification pass over the JSON values.
  • Verifier blind spots. A verifier can be wrong in both directions. Measure it.

The longer-term direction, covered in my post on hybrid, agentic and long-context RAG, does not remove this problem either. An agent that retrieves several times has more chances to find the right passage and more chances to chain a wrong inference.

A checklist you can start with this week

  1. Add a "the documents do not contain this" instruction and test it with at least 20 unanswerable questions.
  2. Add a retrieval gate with a threshold calibrated on real queries; log every abstention.
  3. Turn on native citations; put each chunk in its own document or `search_result` block with your own ID in the source field.
  4. Store the answer, the cited passages and the retrieval scores for every response.
  5. Add a claim-level verifier on a sample first, then on the flows where wrong answers cost money or trust.
  6. Track faithfulness, citation precision and recall, and abstention quality in CI on a fixed evaluation set.
  7. Show numbered, clickable citations with passage previews and a visible verification state.
  8. Review a random sample of answers with citations by hand every week, because citations can be post-rationalised.

If you take only one thing from this: make wrong answers cheap to notice. A system that is occasionally wrong and shows its evidence is a tool; one that is occasionally wrong and sounds certain is a liability. If you need help building or auditing this kind of pipeline, see my AI engineering work.

Sources

  1. Anthropic: Citations (Claude API documentation)
  2. Anthropic: Search results (Claude API documentation)
  3. Anthropic: Reduce hallucinations (Claude API documentation)
  4. Anthropic: Introducing Citations on the Anthropic API
  5. Simon Willison: Anthropic's new Citations API (24 January 2025)
  6. OpenAI: Web search guide (url_citation annotations and display requirement)
  7. Cohere: Documents and citations
  8. Google Cloud: Check grounding API
  9. AWS: Amazon Bedrock Guardrails contextual grounding check
  10. Ragas: Faithfulness metric
  11. Kalai, Nachum, Vempala, Zhang: Why Language Models Hallucinate (arXiv 2509.04664)
  12. Magesh et al.: Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (arXiv 2405.20362)
  13. Wallat, Heuss, de Rijke, Anand: Correctness is not Faithfulness in RAG Attributions (arXiv 2412.18004)

Frequently asked questions

How do you reduce hallucinations in a RAG application?

Stack several layers instead of relying on one. Improve retrieval first, then restrict the model to the provided documents, allow it to say it does not know, make it cite the passages it used, verify each claim against those passages, and show the sources in the UI. No single layer eliminates hallucinations; Anthropic's own guidance says these techniques reduce them significantly but do not remove them.

Does RAG eliminate hallucinations?

No. In the Stanford and Yale study of commercial legal research tools, products that marketed RAG as a fix still produced hallucinated answers between 17% and 33% of the time. The study counts an answer as hallucinated when it is incorrect or misgrounded, meaning it claims a source supports something it does not.

What is the difference between faithfulness and correctness?

Faithfulness asks whether every claim in the answer is supported by the retrieved context. Correctness asks whether the claim is true in the world. A faithful answer built on an outdated document is wrong but faithful; a correct answer from the model's memory that the documents do not support is correct but unfaithful. In RAG you usually want both, and you measure them separately.

How do the Anthropic citations work?

You pass documents or search_result blocks with citations enabled, and the response text blocks carry citation objects pointing to character ranges, page numbers or content blocks in your sources. The cited_text field does not count toward output tokens, and the API guarantees the pointers are valid. Citations cannot be combined with structured outputs and currently cover text only.

When should an LLM say "I don't know"?

When the retrieved evidence does not contain the answer. Implement it in two places: a retrieval gate that stops generation when the best passages are weak, and a prompt that explicitly allows the model to state that the documents lack the information. Then test it with questions your corpus cannot answer, otherwise the behaviour is never verified.

How should a chatbot show its sources?

Put numbered markers next to the claims they support, show the cited passage on hover or tap, link to the original document at the right location, and mark answers or sentences that could not be verified. Never show a citation that has not been checked to point at real text, because a confident-looking footnote makes a wrong answer more believable.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.