CertSafari
    CLAUDE-CERTIFIED-ASSOCIATE-FOUNDATIONS-CCAO-F-VAR5 · Lessons

    Domain 2 · Lesson 9/30

    Comparing Claude Output Versions Against Stated Criteria

    Edit, adapt, refine, and compare outputs for the intended audience

    7 min read
    3.5% of exam
    4 sources
    Published 28 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Set explicit, audience-based criteria before comparing two or more Claude drafts
    • Compare candidate outputs side by side against a reader's stated priorities, not general polish
    • Check that versions adapted for different audiences still carry the same substance
    • Take ownership of the final choice and state what you changed

    1.Set the criteria before you read the drafts

    When you have two versions of the same output, such as two pitches, two announcements or two strategy updates, the first step is not to read them. It is to write down what 'better' means for this audience. Without that, you will pick the draft that sounds more polished to you, and that may not be the one the reader needs.

    Anthropic's evaluation guidance makes this point for developers, and it transfers directly. Criteria should be tied to purpose and user needs, and they should be specific. Its own contrast shows what specific means: 'the model should classify sentiments well' is a bad criterion, while a measurable threshold on a defined test set is a good one. For a board pack, 'which is clearer' is the weak version. 'Which states the decision we need from the board, the cost, and the risk within the first page' is the strong one.

    Criteria from Anthropic's evaluation guidance, reframed as questions for comparing two drafts
    CriterionQuestion from the guidanceApplied to a draft comparison
    RelevanceHow well does the model directly address the user's questions or instructions?Which draft answers what this reader actually asked or cares about?
    CoherenceHow important is it for the information to be presented in a logical, easy to follow manner?Which draft can the reader follow without re-reading?
    Tone and styleHow appropriate is its language for the target audience?Which register matches this reader: formal, conversational, technical?
    ConsistencyHow similar do the model's responses need to be for similar types of input?Do both drafts make the same core claims, or does one add or drop something?

    A support lead finds a Claude-drafted troubleshooting article accurate but written at a reading level too advanced for the average customer. Which follow-up instruction best refines the draft without sacrificing accuracy?

    Sources1

    2.Comparing candidates against the reader's stated priorities

    With criteria in hand, compare the drafts one criterion at a time instead of forming an overall impression. Suppose one pitch leads with cost savings and another with reliability. The useful question is not which pitch is more persuasive in general. It is which one addresses the priorities the client has actually stated. That is the relevance criterion applied literally.

    Comparing versions also has a verification benefit. Anthropic's hallucination guidance describes best-of-N verification: running the same prompt more than once and comparing the outputs, because inconsistencies between them can signal a problem. If two drafts of the same brief disagree on a figure or a commitment, that difference is a flag to check before you send either one, not a stylistic choice between them.

    Once you have picked a version, you do not have to take it as is. The same guidance describes iterative refinement: feeding the chosen output back to Claude and asking it to verify or expand on specific statements. You can also borrow a strong section from the draft you rejected and ask Claude to fold it in.

    Sources12

    3.Adapted versions: change the register, keep the substance

    A different comparison comes up when one piece of content has been adapted for two audiences: a headquarters announcement and a regional version, or an integration guide for experienced developers and one for low-code users. Here the versions are supposed to differ in vocabulary, depth, examples and ordering. They are not supposed to differ in the facts, requirements, dates or obligations they convey.

    The consistency criterion asks how important it is that similar inputs produce semantically similar answers. For an adapted pair, the answer is: very important for the substance, and not at all for the wording. So check the pair claim by claim. Every requirement in the expert version should appear in the simplified version, even if it is phrased more simply. Nothing in the simplified version should contradict the expert version or quietly soften a rule. Simplifying is where content most often goes missing.

    A practical way to run this check is to ask Claude to list the claims in each version and flag any that appear in only one. This follows the iterative-refinement pattern of turning outputs back into inputs. You still confirm the result yourself.

    A marketing analyst asks Claude for a sales summary for executives who want only key trends and decisions, not methodology, but the first draft includes several paragraphs on data cleaning. How should the analyst refine the prompt for this audience?

    Sources12

    4.The final choice is yours

    Comparing and refining can be helped along by Claude, but not handed over to it. The AI Fluency material is clear that the person stays the decision-maker and that the work is not done when the AI finishes. Its end-to-end exercise ends with a test worth using on any adapted output: would you put your name on it? Along the way it asks you to note what you kept, changed or threw out, and what you added that only you could provide. That record of your own edits is what makes the output yours.

    Sources34

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.When choosing between two Claude drafts, start by reading both and picking the one that reads best.Why is that wrong?

      Define criteria based on the audience's purpose and needs first. Otherwise the choice reflects your taste rather than what the reader needs.

      Covered in Set the criteria before you read the drafts

    2. 2.Two versions adapted for different audiences only need to be checked for tone, because they are meant to be different.Why is that wrong?

      Adapted versions should differ in register and depth but convey the same substance. Check them claim by claim for missing, added or contradictory content.

      Covered in Adapted versions: change the register, keep the substance

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Align your criteria with your application's purpose and user needs.”
      ↩︎ Set the criteria before you read the drafts
      “How well does the model's output style match expectations? How appropriate is its language for the target audience?”
      ↩︎ Set the criteria before you read the drafts
      “How well does the model directly address the user's questions or instructions?”
      ↩︎ Comparing candidates against the reader's stated priorities
      “How similar do the model's responses need to be for similar types of input?”
      ↩︎ Adapted versions: change the register, keep the substance
      “Align your criteria with your application's purpose and user needs.”
      ↩︎ Exam trap 1
      “If a user asks the same question twice, how important is it that they get semantically similar answers?”
      ↩︎ Exam trap 2
    2. 2.
      “Run Claude through the same prompt multiple times and compare the outputs.”
      ↩︎ Comparing candidates against the reader's stated priorities
      “Use Claude's outputs as inputs for follow-up prompts, asking it to verify or expand on previous statements.”
      ↩︎ Comparing candidates against the reader's stated priorities
      “This can catch and correct inconsistencies.”
      ↩︎ Adapted versions: change the register, keep the substance