CertSafari
    CLAUDE-CERTIFIED-ASSOCIATE-FOUNDATIONS-CCAO-F-VAR5 · Lessons

    Domain 2 · Lesson 5/30

    Evaluating Claude outputs for accuracy and completeness

    Evaluate Claude-generated outputs for accuracy and completeness

    8 min read
    3.5% of exam
    5 sources
    Published 28 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Explain why a fluent, confident Claude output is not evidence that it is accurate or complete
    • Turn a vague sense of quality into specific, measurable criteria that different reviewers can apply the same way
    • Check an output's completeness against the original request, not against how finished it sounds

    Key concept

    Criteria-first evaluation — You judge a Claude output against criteria written down before you read it: what the request required and what counts as correct. The judgement is not how convincing the output feels. Without stated criteria, two reviewers reading the same draft will reach two different verdicts.

    1.Why "it sounds right" is not the test

    Anthropic's guidance on building with Claude says it openly: even the most capable models can produce text that is wrong, or that contradicts the material they were given. This is called hallucination. What makes it hard to catch is that it doesn't look any different from a correct answer. The prose is just as fluent, the structure just as tidy and the tone just as confident.

    Anthropic's agent guidance says the same thing about finished work. Agents are good at producing things that look done. A research brief can arrive with footnotes and headings and still have a thin section, a quote that drifts from its source, or a citation that points to a press release when the original filing was needed. Each of those flaws is invisible if your only question is whether the document reads well.

    This example comes from Anthropic's own evaluation cookbook. The golden answer says the Kansas City Chiefs beat the San Francisco 49ers. The output was coherent, confident and reasonable on its face, and it was wrong. You only find out by comparing it with a known correct answer. Reading it for plausibility tells you nothing. The rest of this lesson follows from that: before you can say an output is accurate and complete, you need something outside the output to measure it against.

    Sources123

    2.Decide what "good" means before you read the draft

    Say a team keeps arguing about whether Claude's drafts are good enough, and every reviewer comes to a different verdict. That usually doesn't mean the reviewers are careless. It means each one is using private criteria. Anthropic's evaluation docs start from this point: define success criteria first, then measure against them. Good criteria are specific, measurable and relevant to what the output is for. "The report should be good" gives a reviewer nothing to check. "The report states the quarter's figure for each of the five service lines and flags any that missed target" gives every reviewer the same checklist.

    A single overall impression also hides problems, because quality has several independent parts. The docs say most use cases need evaluation along several criteria at once. An output can be accurate and irrelevant. It can be relevant and inconsistent. The table turns the docs' list of example criteria into questions you can ask of any output.

    Dimensions to check separately, drawn from Anthropic's example success criteria
    DimensionQuestion to ask of the output
    Task fidelityDoes it do the task correctly, including on rare or challenging inputs?
    ConsistencyWould a similar question get a semantically similar answer?
    Relevance and coherenceDoes it directly address the question or instruction, in a logical, easy-to-follow order?
    Tone and styleIs the language appropriate for the target audience?
    PrivacyDoes it follow instructions not to use or share certain details?
    Context useDoes it use and build on the information it was given?

    Your criteria need to be more precise than the request was. Anthropic's rubric guidance puts it this way: the rubric should always be more specific than the task. If the task says "cover demand charges", a reviewer can skim, spot a paragraph on the topic and tick the box. If the rubric says the section must state a $/kW figure or a percentage of operating cost, the reviewer has to find evidence. Anthropic notes that the default failure mode is a reviewer who approves everything, and vague criteria are what allow it.

    If you can't write criteria from scratch, start from an example. Hand Claude a known-good example and ask it to explain what makes it good. Then turn that explanation into criteria your team agrees on.

    A team is evaluating a customer support assistant and only measures whether responses are factually correct, ignoring tone, relevance, and consistency. Following recommended evaluation practice, what is the main risk of this approach?

    Sources42

    3.Completeness is measured against the request

    Completeness is measured against what was asked, not against how long or finished the output looks. In Anthropic's evaluation cookbook, every test item has a golden answer, which is the reference the output is compared with. For open-ended requests, the golden answer is a list of what a correct response must contain and what it must not contain.

    One cookbook request asks for a workout with at least 50 reps of pulling leg exercises, at least 50 reps of pulling arm exercises, and ten minutes of core. Its golden answer spells out the limits. Squats don't count as pulling legs and presses don't count as pulling arms. Stretching or a warm-up is allowed, and no other meaningful exercises are. Every part of the request becomes a separate item to check.

    The most common completeness failure is an output that answers a nearby question instead of the one asked. Anthropic's customer-support criteria measure how well Claude directly addresses the specific question or issue. Suppose a retention lead asks why cancellations rose last quarter and gets a careful explanation of how cancellation rates are defined and measured. Every sentence may be correct and the answer still fails. The question was about causes, and the response contains none. Grade that output as incomplete, however accurate its background material is.

    A user asks Claude about a company earnings announcement that happened yesterday, and Claude responds confidently with financial figures. Given that Claude's knowledge comes from training data with a fixed cutoff date, what should the user do before trusting this answer?

    Sources345

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A fluent, well-structured, confident output can be trusted as accurate.Why is that wrong?

      Hallucinated content is exactly as fluent as correct content, and polished work often hides thin coverage or drifting quotes. Accuracy has to be checked against something outside the output.

      Covered in Why "it sounds right" is not the test

    2. 2.Checking the draft against the wording of the original request is precise enough for reviewers to agree.Why is that wrong?

      Criteria copied from the request let a reviewer tick a box after skimming. Criteria that name the specific evidence required, such as a figure, a source type or a threshold, are what make reviewers converge.

      Covered in Decide what "good" means before you read the draft

    3. 3.An answer made entirely of correct, relevant-sounding background information is a good answer.Why is that wrong?

      Relevance is judged against the specific question asked. Correct material that answers a different question is incomplete.

      Covered in Completeness is measured against the request

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “can sometimes generate text that is factually incorrect or inconsistent with the given context”
      ↩︎ Why "it sounds right" is not the test
      “can sometimes generate text that is factually incorrect or inconsistent with the given context”
      ↩︎ Exam trap 1
    2. 2.
      “Agents are good at producing things that look done.”
      ↩︎ Why "it sounds right" is not the test
      “The default failure mode is a grader that approves everything.”
      ↩︎ Decide what "good" means before you read the draft
      “Hand Claude a known-good example of the artifact and ask it to analyze what makes it good, then turn that analysis into criteria.”
      ↩︎ Decide what "good" means before you read the draft
      “The rubric should always be more specific than the task.”
      ↩︎ Exam trap 2
    3. 3.
      “A correct answer states that the Kansas City Chiefs defeated the San Francisco 49ers.”
      ↩︎ Why "it sounds right" is not the test
      “A "golden answer" to which we compare the model output.”
      ↩︎ Completeness is measured against the request
      “It can but does not have to include stretching or a dynamic warmup, but it cannot include any other meaningful exercises.”
      ↩︎ Completeness is measured against the request
    4. 4.
      “Specific: Clearly define what you want to achieve. Instead of "good performance," specify "accurate sentiment classification."”
      ↩︎ Decide what "good" means before you read the draft
      “Most use cases need multidimensional evaluation along several success criteria.”
      ↩︎ Decide what "good" means before you read the draft
      “How well does the model directly address the user's questions or instructions?”
      ↩︎ Completeness is measured against the request
      “Specific: Clearly define what you want to achieve. Instead of "good performance," specify "accurate sentiment classification."”
      ↩︎ Key concept
    5. 5.
      “This assesses how well Claude's response addresses the customer's specific question or issue.”
      ↩︎ Completeness is measured against the request
      “This assesses how well Claude's response addresses the customer's specific question or issue.”
      ↩︎ Exam trap 3

    Continue to page 2 of 2

    Checking Claude's claims against sources and outcomes