What you will be able to do
- Recognise the two common shapes of a Claude hallucination: stale knowledge and convincing but ungrounded specifics
- Use best-of-N runs, step-by-step reasoning and follow-up prompts to detect a suspected hallucination
- Choose the right grounding fix: permission to say "I don't know", quote-first extraction, or restricting Claude to external knowledge
- Build a prompt that makes Claude retract any claim it cannot support with a quote
- Explain why hallucinations are a system-level reliability problem rather than a single bad answer, and where in a system they cluster
Key concept
Diagnose before you fix — A misbehaving LLM system fails for a nameable reason: the model filled a gap with invented specifics, the prompt did not say what was needed, or the model is wrong for the task. Each cause has its own remedy, so name the cause first. A fix aimed at the wrong cause leaves the failure in place.
1.What a hallucination actually is
Anthropic's guardrails documentation defines hallucination as a model producing text that is "factually incorrect or inconsistent with the given context." There are two halves to that definition. An answer can be wrong about the world, or it can contradict the documents you gave Claude in the prompt. The second kind matters most in retrieval and document-analysis systems. There the model has the right material in front of it and still produces something the material does not say.
Anthropic's support guidance calls hallucination a byproduct of current limitations in frontier generative models, and names two recognisable shapes. The first is stale knowledge. In some subject areas Claude may not have been trained on the most recent information, so questions about current events can confuse it. The second is convincing fabrication. Claude can produce quotes that look authoritative or sound convincing but are not grounded in fact. The second shape is more dangerous because nothing on the surface of the output looks wrong.
Anthropic Academy's course on AI capabilities and limitations explains why fabrication clusters around specifics. It treats most real failures as two model properties interacting. The pairing it names for this case is Next Token Prediction + Knowledge, which produces *hallucinated specifics*. The model is built to produce a fluent next token. When its knowledge runs out partway through a name, date, figure or citation, fluency still wins. The diagnostic lesson is that the specific details of an answer (numbers, names, quotations) are where hallucination is most likely. The general shape of the answer is less likely to be wrong.
Why do hallucinations count as a system issue and not just an unlucky answer? Anthropic's guide is explicit that even the most advanced models, Claude included, can sometimes produce them, and that hallucinations "can undermine the reliability of your AI-driven solutions." That is the diagnostic framing to carry into the exam. Hallucinations are a property of the whole deployment: the model, the prompt that did or did not permit uncertainty, and the context that was or was not supplied. When you triage a failing system, hallucinations sit alongside prompt failure and model mismatch as one of three distinct causes, and the rest of this lesson covers the techniques the guide offers to minimize hallucinations and keep outputs accurate and trustworthy. Diagnosing hallucinations therefore means asking three questions in order: Is the wrong detail a specific (a number, name, date or quote)? Does the prompt supply that detail, or leave a gap? Was the model allowed to abstain? A "yes, no, no" pattern points squarely at hallucinations rather than at a prompt or model problem.
2.Detecting a suspected hallucination
You cannot fix what you have not confirmed. Anthropic's documentation describes three techniques that show a hallucination without needing ground truth. Each one exposes a different symptom.
| Technique | What you do | What it reveals |
|---|---|---|
| Best-of-N verification | Run the same prompt multiple times and compare the outputs | Inconsistency: details that change between runs are likely invented |
| Chain-of-thought verification | Ask Claude to explain its reasoning step by step before the final answer | Faulty logic or hidden assumptions behind a conclusion |
| Iterative refinement | Feed Claude's output back in a follow-up prompt asking it to verify or expand on it | Inconsistencies Claude can catch and correct on a second pass |
Best-of-N is the cheapest signal, and it follows directly from the Next Token Prediction + Knowledge diagnosis. Something Claude actually knows tends to come back the same across runs. A gap filled with a fluent guess can come back different each time. Chain-of-thought verification works on another layer. It does not test the facts. It surfaces the reasoning, so you can see whether a correct-sounding conclusion rests on an assumption nobody checked.
The setup steps are probably grounded, but the driver version is probably hallucinated. It is a specific detail that changes between runs, which is exactly the inconsistency best-of-N is meant to catch. Next, check whether the prompt actually supplies the version (if it does not, Claude is guessing), then apply a grounding fix to that detail.
A retrieval-augmented HR assistant answers employee questions using uploaded policy PDFs. Reviewers find it confidently states a specific vacation-accrual percentage that does not appear anywhere in the source documents. Which change to the assistant's instructions best addresses the root cause of this failure?
Correct answer: A — Require the assistant to quote exact policy text before answering, admitting uncertainty when none exists.
- A. Correct. Forcing the assistant to ground claims in extracted quotes and explicitly permitting an admission of uncertainty directly targets the mechanism behind the fabrication: the model filled a gap with a plausible-sounding invented figure instead of acknowledging the documents didn't cover it.
- B. Incorrect. Explicitly encouraging the model to fall back on general training knowledge is the opposite of the fix — it legitimizes exactly the behavior that produced the fabricated percentage.
- C. Incorrect. Forcing a numeric answer in every case removes the model's ability to say it doesn't know, making fabrication more likely, not less.
- D. Incorrect. Restating the full document doesn't stop the model from inventing details for questions the document never addresses, and it wastes context on unrelated sections.
Sources1
3.Grounding fixes: uncertainty, quotes, and a closed knowledge boundary
Once a hallucination is confirmed, the fix should match what caused it. Anthropic's guide gives three prompt-level remedies, and each one takes away a different reason for the model to guess.
Give Claude permission to say "I don't know." A model asked a question tends to produce an answer. If the prompt explicitly allows uncertainty, abstaining becomes an acceptable output rather than a failure. The documentation says this simple technique can drastically reduce false information.
As our M&A advisor, analyze this report on the potential acquisition of AcmeCo by ExampleCorp.
<report>
{{REPORT}}
</report>
Focus on financial projections, integration risks, and regulatory hurdles. If you're unsure about any aspect or if the report lacks necessary information, say "I don't have enough information to confidently assess this."Extract direct quotes first. For long documents, which the guide puts at over 20k tokens, have Claude pull word-for-word quotes before it does the actual task. The later analysis is then anchored to text that exists in the document, not to Claude's impression of it. This targets the "inconsistent with the given context" half of the definition.
Restrict Claude to the documents you provided. If an answer blends retrieved material with the model's general knowledge, stale knowledge can slip in. Explicitly telling Claude to use only the provided documents closes that route.
Sources1
4.Making claims auditable with citation checks
The strongest fix combines grounding with verification. Claude cites a quote and source for each claim, which makes the response auditable. It then checks its own draft: every claim must be backed by a supporting quote, and any claim without one is retracted. This turns a hallucination from invisible fabricated text into a visible gap you can see.
Draft a press release for our new cybersecurity product, AcmeSecurity Pro, using only information from these product briefs and market reports.
<documents>
{{DOCUMENTS}}
</documents>
After drafting, review each claim in your press release. For each claim, find a direct quote from the documents that supports it. If you can't find a supporting quote for a claim, remove that claim from the press release and mark where it was removed with empty [] brackets.The empty [] markers are a useful diagnostic in their own right. Each one marks a claim the model wanted to make but the sources did not support. That is exactly where a hallucination would otherwise have appeared. Verification also has a human side. Anthropic warns against relying on Claude as a singular source of truth. When Claude works from web search results, reviewers should read the cited originals, because the synthesis can drop context that changes what a source means.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A response that quotes a source in authoritative language is grounded in that source.Why is that wrong?
Fabricated quotes are one of the recognised forms of hallucination. Looking authoritative is not evidence of grounding. Only checking the quote against the actual source is.
Covered in What a hallucination actually is
2.Hallucinations only happen with weaker or older models, so upgrading to the most capable model removes them.Why is that wrong?
Anthropic states that even the most advanced models, including Claude, can sometimes hallucinate. Hallucinations are a system-level reliability issue addressed with prompt and verification techniques, not solely by model choice.
Covered in What a hallucination actually is
3.If one run of a prompt gives a plausible answer, the system is not hallucinating on that input.Why is that wrong?
A single plausible run shows nothing about consistency. Running the same prompt several times and comparing the outputs exposes invented details, because they tend to change between runs.
Covered in Detecting a suspected hallucination
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-hallucinationsOfficial docs
“can sometimes generate text that is factually incorrect or inconsistent with the given context”
↩︎ What a hallucination actually is“This phenomenon, known as "hallucination," can undermine the reliability of your AI-driven solutions.”
↩︎ What a hallucination actually is“This guide will explore techniques to minimize hallucinations and ensure Claude's outputs are accurate and trustworthy.”
↩︎ What a hallucination actually is“Inconsistencies across outputs could indicate hallucinations.”
↩︎ Detecting a suspected hallucination“Ask Claude to explain its reasoning step-by-step before giving a final answer. This can reveal faulty logic or assumptions.”
↩︎ Detecting a suspected hallucination“Use Claude's outputs as inputs for follow-up prompts, asking it to verify or expand on previous statements.”
↩︎ Detecting a suspected hallucination“Explicitly give Claude permission to admit uncertainty. This simple technique can drastically reduce false information.”
↩︎ Grounding fixes: uncertainty, quotes, and a closed knowledge boundary“For tasks involving long documents (>20k tokens), ask Claude to extract word-for-word quotes first before performing its task.”
↩︎ Grounding fixes: uncertainty, quotes, and a closed knowledge boundary“Explicitly instruct Claude to only use information from provided documents and not its general knowledge.”
↩︎ Grounding fixes: uncertainty, quotes, and a closed knowledge boundary“If it can't find a quote, it must retract the claim.”
↩︎ Making claims auditable with citation checks“Even the most advanced language models, like Claude, can sometimes generate text that is factually incorrect or inconsistent with the given context.”
↩︎ Exam trap 2“Run Claude through the same prompt multiple times and compare the outputs.”
↩︎ Exam trap 3 - 2.https://support.claude.com/en/articles/8525154-claude-is-providing-incorrect-or-misleading-responses-what-s-going-onOfficial docs
“Claude might not have been trained on the most-up-to-date information and may get confused when prompted about current events”
↩︎ What a hallucination actually is“Users should not rely on Claude as a singular source of truth”
↩︎ Making claims auditable with citation checks“Claude can display quotes that may look authoritative or sound convincing, but are not grounded in fact”
↩︎ Exam trap 1 - 3.https://academy.claude.com/courses/ai-capabilities-and-limitations/when-properties-collideOfficial docs
“Next Token Prediction + Knowledge (hallucinated specifics)”
↩︎ What a hallucination actually is“Naming the properties at play points you straight to the fix: verify specifics, re-supply context, offload to code execution, or invite pushback.”
↩︎ Key concept