What you will be able to do
- Explain why a fluent, confident Claude output is not evidence that it is accurate or complete
- Turn a vague sense of quality into specific, measurable criteria that different reviewers can apply the same way
- Check an output's completeness against the original request, not against how finished it sounds
Key concept
Criteria-first evaluation — You judge a Claude output against criteria written down before you read it: what the request required and what counts as correct. The judgement is not how convincing the output feels. Without stated criteria, two reviewers reading the same draft will reach two different verdicts.
1.Why "it sounds right" is not the test
Anthropic's guidance on building with Claude says it openly: even the most capable models can produce text that is wrong, or that contradicts the material they were given. This is called hallucination. What makes it hard to catch is that it doesn't look any different from a correct answer. The prose is just as fluent, the structure just as tidy and the tone just as confident.
Anthropic's agent guidance says the same thing about finished work. Agents are good at producing things that look done. A research brief can arrive with footnotes and headings and still have a thin section, a quote that drifts from its source, or a citation that points to a press release when the original filing was needed. Each of those flaws is invisible if your only question is whether the document reads well.
This example comes from Anthropic's own evaluation cookbook. The golden answer says the Kansas City Chiefs beat the San Francisco 49ers. The output was coherent, confident and reasonable on its face, and it was wrong. You only find out by comparing it with a known correct answer. Reading it for plausibility tells you nothing. The rest of this lesson follows from that: before you can say an output is accurate and complete, you need something outside the output to measure it against.
2.Decide what "good" means before you read the draft
Say a team keeps arguing about whether Claude's drafts are good enough, and every reviewer comes to a different verdict. That usually doesn't mean the reviewers are careless. It means each one is using private criteria. Anthropic's evaluation docs start from this point: define success criteria first, then measure against them. Good criteria are specific, measurable and relevant to what the output is for. "The report should be good" gives a reviewer nothing to check. "The report states the quarter's figure for each of the five service lines and flags any that missed target" gives every reviewer the same checklist.
A single overall impression also hides problems, because quality has several independent parts. The docs say most use cases need evaluation along several criteria at once. An output can be accurate and irrelevant. It can be relevant and inconsistent. The table turns the docs' list of example criteria into questions you can ask of any output.
| Dimension | Question to ask of the output |
|---|---|
| Task fidelity | Does it do the task correctly, including on rare or challenging inputs? |
| Consistency | Would a similar question get a semantically similar answer? |
| Relevance and coherence | Does it directly address the question or instruction, in a logical, easy-to-follow order? |
| Tone and style | Is the language appropriate for the target audience? |
| Privacy | Does it follow instructions not to use or share certain details? |
| Context use | Does it use and build on the information it was given? |
Your criteria need to be more precise than the request was. Anthropic's rubric guidance puts it this way: the rubric should always be more specific than the task. If the task says "cover demand charges", a reviewer can skim, spot a paragraph on the topic and tick the box. If the rubric says the section must state a $/kW figure or a percentage of operating cost, the reviewer has to find evidence. Anthropic notes that the default failure mode is a reviewer who approves everything, and vague criteria are what allow it.
If you can't write criteria from scratch, start from an example. Hand Claude a known-good example and ask it to explain what makes it good. Then turn that explanation into criteria your team agrees on.
A team is evaluating a customer support assistant and only measures whether responses are factually correct, ignoring tone, relevance, and consistency. Following recommended evaluation practice, what is the main risk of this approach?
Correct answer: D — Single-dimension evaluation misses important quality issues, so complex tasks should be scored across multiple criteria.
- A. Incorrect. Factual correctness alone is insufficient for complex tasks like customer support; ignoring tone, relevance, and consistency can lead to poor user experiences. Recommended practice emphasizes multidimensional evaluation to fully assess quality.
- B. Incorrect. Adding more evaluation dimensions does not reduce reliability; rather, it provides a more comprehensive and accurate assessment. Multidimensional evaluation is encouraged for complex tasks to capture quality aspects beyond factual accuracy.
- C. Incorrect. Multidimensional evaluation is recommended broadly for complex tasks, not only safety-critical systems. A customer support assistant benefits from evaluating tone and relevance, as these influence user satisfaction.
- D. Correct. Single-dimension evaluation misses important quality issues like tone, relevance, and consistency. Recommended practice calls for multidimensional scoring to ensure comprehensive assessment of complex tasks.
3.Completeness is measured against the request
Completeness is measured against what was asked, not against how long or finished the output looks. In Anthropic's evaluation cookbook, every test item has a golden answer, which is the reference the output is compared with. For open-ended requests, the golden answer is a list of what a correct response must contain and what it must not contain.
One cookbook request asks for a workout with at least 50 reps of pulling leg exercises, at least 50 reps of pulling arm exercises, and ten minutes of core. Its golden answer spells out the limits. Squats don't count as pulling legs and presses don't count as pulling arms. Stretching or a warm-up is allowed, and no other meaningful exercises are. Every part of the request becomes a separate item to check.
No. Restating the request is not evidence that it was met. Check the parts. Leg exercises: 36 hamstring curls plus 40 single-leg Romanian deadlifts is 76. Arm exercises: 30 rows plus 24 chin-ups is 54. Both totals clear 50. The core section is labelled "(10 minutes)", but only the plank has a stated duration. The twists and crunches are given in reps, so the ten minutes has to be judged, not taken from the label. The warm-up and cool-down are allowed. Completeness is established item by item, never by the output's own summary.
The most common completeness failure is an output that answers a nearby question instead of the one asked. Anthropic's customer-support criteria measure how well Claude directly addresses the specific question or issue. Suppose a retention lead asks why cancellations rose last quarter and gets a careful explanation of how cancellation rates are defined and measured. Every sentence may be correct and the answer still fails. The question was about causes, and the response contains none. Grade that output as incomplete, however accurate its background material is.
A user asks Claude about a company earnings announcement that happened yesterday, and Claude responds confidently with financial figures. Given that Claude's knowledge comes from training data with a fixed cutoff date, what should the user do before trusting this answer?
Correct answer: B — Verify the figures against a current source, since recent events fall outside Claude's training cutoff
- A. Claude's training data has a fixed cutoff and is not updated in real time between releases, so events occurring after that cutoff are not reliably reflected in its knowledge.
- B. Correct. Because a model's knowledge is only reliable up to its training cutoff, users should verify claims about very recent events against a current, authoritative source before trusting them.
- C. Claude does not always decline questions about recent events; it may respond confidently even when the information falls outside its reliable knowledge, which is exactly why independent verification is needed.
- D. Ignoring the response entirely overstates the limitation; the concern is specifically about information beyond the training cutoff, not a blanket inability to discuss financial topics.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A fluent, well-structured, confident output can be trusted as accurate.Why is that wrong?
Hallucinated content is exactly as fluent as correct content, and polished work often hides thin coverage or drifting quotes. Accuracy has to be checked against something outside the output.
Covered in Why "it sounds right" is not the test
2.Checking the draft against the wording of the original request is precise enough for reviewers to agree.Why is that wrong?
Criteria copied from the request let a reviewer tick a box after skimming. Criteria that name the specific evidence required, such as a figure, a source type or a threshold, are what make reviewers converge.
Covered in Decide what "good" means before you read the draft
3.An answer made entirely of correct, relevant-sounding background information is a good answer.Why is that wrong?
Relevance is judged against the specific question asked. Correct material that answers a different question is incomplete.
Covered in Completeness is measured against the request
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-hallucinationsOfficial docs
“can sometimes generate text that is factually incorrect or inconsistent with the given context”
↩︎ Why "it sounds right" is not the test“can sometimes generate text that is factually incorrect or inconsistent with the given context”
↩︎ Exam trap 1 - 2.
“Agents are good at producing things that look done.”
↩︎ Why "it sounds right" is not the test“The default failure mode is a grader that approves everything.”
↩︎ Decide what "good" means before you read the draft“Hand Claude a known-good example of the artifact and ask it to analyze what makes it good, then turn that analysis into criteria.”
↩︎ Decide what "good" means before you read the draft“The rubric should always be more specific than the task.”
↩︎ Exam trap 2 - 3.https://platform.claude.com/cookbook/misc-building-evalsSecondary source
“A correct answer states that the Kansas City Chiefs defeated the San Francisco 49ers.”
↩︎ Why "it sounds right" is not the test“A "golden answer" to which we compare the model output.”
↩︎ Completeness is measured against the request“It can but does not have to include stretching or a dynamic warmup, but it cannot include any other meaningful exercises.”
↩︎ Completeness is measured against the request - 4.
“Specific: Clearly define what you want to achieve. Instead of "good performance," specify "accurate sentiment classification."”
↩︎ Decide what "good" means before you read the draft“Most use cases need multidimensional evaluation along several success criteria.”
↩︎ Decide what "good" means before you read the draft“How well does the model directly address the user's questions or instructions?”
↩︎ Completeness is measured against the request“Specific: Clearly define what you want to achieve. Instead of "good performance," specify "accurate sentiment classification."”
↩︎ Key concept - 5.
“This assesses how well Claude's response addresses the customer's specific question or issue.”
↩︎ Completeness is measured against the request“This assesses how well Claude's response addresses the customer's specific question or issue.”
↩︎ Exam trap 3