What you will be able to do
- Set explicit, audience-based criteria before comparing two or more Claude drafts
- Compare candidate outputs side by side against a reader's stated priorities, not general polish
- Check that versions adapted for different audiences still carry the same substance
- Take ownership of the final choice and state what you changed
1.Set the criteria before you read the drafts
When you have two versions of the same output, such as two pitches, two announcements or two strategy updates, the first step is not to read them. It is to write down what 'better' means for this audience. Without that, you will pick the draft that sounds more polished to you, and that may not be the one the reader needs.
Anthropic's evaluation guidance makes this point for developers, and it transfers directly. Criteria should be tied to purpose and user needs, and they should be specific. Its own contrast shows what specific means: 'the model should classify sentiments well' is a bad criterion, while a measurable threshold on a defined test set is a good one. For a board pack, 'which is clearer' is the weak version. 'Which states the decision we need from the board, the cost, and the risk within the first page' is the strong one.
| Criterion | Question from the guidance | Applied to a draft comparison |
|---|---|---|
| Relevance | How well does the model directly address the user's questions or instructions? | Which draft answers what this reader actually asked or cares about? |
| Coherence | How important is it for the information to be presented in a logical, easy to follow manner? | Which draft can the reader follow without re-reading? |
| Tone and style | How appropriate is its language for the target audience? | Which register matches this reader: formal, conversational, technical? |
| Consistency | How similar do the model's responses need to be for similar types of input? | Do both drafts make the same core claims, or does one add or drop something? |
A support lead finds a Claude-drafted troubleshooting article accurate but written at a reading level too advanced for the average customer. Which follow-up instruction best refines the draft without sacrificing accuracy?
Correct answer: A — Rewrite the article with shorter sentences and everyday words, keeping every step accurate
- A. Correct. Simplifying sentence structure and vocabulary while preserving the accurate steps directly targets the readability problem without losing correctness.
- B. Incorrect. Removing steps changes the article's completeness and could omit cases some customers need, which does not fix a reading-level problem.
- C. Incorrect. Adding more technical detail moves the article further from, not closer to, the audience's reading level.
- D. Incorrect. A formal specification format is written for engineers and would make the article harder for typical customers to follow.
Sources1
2.Comparing candidates against the reader's stated priorities
With criteria in hand, compare the drafts one criterion at a time instead of forming an overall impression. Suppose one pitch leads with cost savings and another with reliability. The useful question is not which pitch is more persuasive in general. It is which one addresses the priorities the client has actually stated. That is the relevance criterion applied literally.
Comparing versions also has a verification benefit. Anthropic's hallucination guidance describes best-of-N verification: running the same prompt more than once and comparing the outputs, because inconsistencies between them can signal a problem. If two drafts of the same brief disagree on a figure or a commitment, that difference is a flag to check before you send either one, not a stylistic choice between them.
Once you have picked a version, you do not have to take it as is. The same guidance describes iterative refinement: feeding the chosen output back to Claude and asking it to verify or expand on specific statements. You can also borrow a strong section from the draft you rejected and ask Claude to fold it in.
3.Adapted versions: change the register, keep the substance
A different comparison comes up when one piece of content has been adapted for two audiences: a headquarters announcement and a regional version, or an integration guide for experienced developers and one for low-code users. Here the versions are supposed to differ in vocabulary, depth, examples and ordering. They are not supposed to differ in the facts, requirements, dates or obligations they convey.
The consistency criterion asks how important it is that similar inputs produce semantically similar answers. For an adapted pair, the answer is: very important for the substance, and not at all for the wording. So check the pair claim by claim. Every requirement in the expert version should appear in the simplified version, even if it is phrased more simply. Nothing in the simplified version should contradict the expert version or quietly soften a rule. Simplifying is where content most often goes missing.
A practical way to run this check is to ask Claude to list the claims in each version and flag any that appear in only one. This follows the iterative-refinement pattern of turning outputs back into inputs. You still confirm the result yourself.
No. The register change is fine, but the substance has changed: the ten-day deadline and the required channel are gone. A reader following the simplified version could break the policy. Restore both facts in plain language, for example 'request leave in the HR system at least ten working days ahead'.
A marketing analyst asks Claude for a sales summary for executives who want only key trends and decisions, not methodology, but the first draft includes several paragraphs on data cleaning. How should the analyst refine the prompt for this audience?
Correct answer: A — Ask Claude to lead with the trends and decisions and drop the data cleaning explanation
- A. Correct. Executives want trends and decisions, so restructuring the summary to foreground that content and dropping methodology detail matches the stated audience need.
- B. Incorrect. Moving the methodology to a footnote still includes content the executives do not want instead of leading with what they need first.
- C. Incorrect. Expanding the methodology section moves the summary further from what a time-constrained executive audience wants.
- D. Incorrect. Uniformly shortening every paragraph keeps the irrelevant methodology section in the summary instead of removing it.
4.The final choice is yours
Comparing and refining can be helped along by Claude, but not handed over to it. The AI Fluency material is clear that the person stays the decision-maker and that the work is not done when the AI finishes. Its end-to-end exercise ends with a test worth using on any adapted output: would you put your name on it? Along the way it asks you to note what you kept, changed or threw out, and what you added that only you could provide. That record of your own edits is what makes the output yours.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.When choosing between two Claude drafts, start by reading both and picking the one that reads best.Why is that wrong?
Define criteria based on the audience's purpose and needs first. Otherwise the choice reflects your taste rather than what the reader needs.
Covered in Set the criteria before you read the drafts
2.Two versions adapted for different audiences only need to be checked for tone, because they are meant to be different.Why is that wrong?
Adapted versions should differ in register and depth but convey the same substance. Check them claim by claim for missing, added or contradictory content.
Covered in Adapted versions: change the register, keep the substance
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Align your criteria with your application's purpose and user needs.”
↩︎ Set the criteria before you read the drafts“How well does the model's output style match expectations? How appropriate is its language for the target audience?”
↩︎ Set the criteria before you read the drafts“How well does the model directly address the user's questions or instructions?”
↩︎ Comparing candidates against the reader's stated priorities“How similar do the model's responses need to be for similar types of input?”
↩︎ Adapted versions: change the register, keep the substance“Align your criteria with your application's purpose and user needs.”
↩︎ Exam trap 1“If a user asks the same question twice, how important is it that they get semantically similar answers?”
↩︎ Exam trap 2 - 2.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-hallucinationsOfficial docs
“Run Claude through the same prompt multiple times and compare the outputs.”
↩︎ Comparing candidates against the reader's stated priorities“Use Claude's outputs as inputs for follow-up prompts, asking it to verify or expand on previous statements.”
↩︎ Comparing candidates against the reader's stated priorities“This can catch and correct inconsistencies.”
↩︎ Adapted versions: change the register, keep the substance - 3.
“Would you put your name on the final output?”
↩︎ The final choice is yours - 4.
“AI accelerates, but doesn't replace expertise. You're still the decision-maker.”
↩︎ The final choice is yours