CertSafari
    CLAUDE-CERTIFIED-ASSOCIATE-FOUNDATIONS-CCAO-F-VAR5 · Lessons

    Domain 1 · Lesson 3/30

    Checking a Refined Prompt: Accuracy, Criteria and Varied Inputs

    Iterate prompts to improve output quality

    6 min read
    3.5% of exam
    5 sources
    Published 28 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Define what 'better' means for a prompt before judging each new version
    • Treat factual errors as a grounding problem to check before rewording
    • Test a refined prompt on varied inputs before adopting it as a standard

    1.Decide what 'better' means before you iterate

    Every round of prompt iteration asks whether the new version is better than the last. Without a definition of better, that question comes down to impression, and impressions drift from round to round. Anthropic's evaluation guidance puts the definition first: good work starts with clearly defining your success criteria and then designing evaluations to measure performance against them, and it calls this cycle central to prompt engineering. Its prompt-engineering overview assumes you already have success criteria, some way to test against them, and a first draft prompt you want to improve.

    Good criteria are specific and measurable. The guide contrasts a bad criterion ('The model should classify sentiments well') with a good one that names a score, a test set and a baseline. You don't need a formal test harness to use the idea. Deciding in advance that 'every figure matches the source' and 'fits on one page for the finance director' are the bar gives you a way to judge each round. The guidance also notes that most use cases need evaluation along several criteria at once. One draft can be well written and wrong at the same time.

    Quality dimensions Anthropic's guide suggests checking, with the question each one asks of an output
    DimensionQuestion to ask of each output
    Task fidelityHow well does it perform the task, including on rare or challenging inputs?
    ConsistencyDo similar inputs get similar answers?
    Relevance and coherenceDoes it directly address the instructions, in a logical, easy-to-follow order?
    Tone and styleIs the language appropriate for the target audience?
    Context useHow well does it use and build on the information it was given?

    Sources12

    2.Wrong facts: check the grounding before the wording

    Some drafts read well but contain wrong figures or claims. That is an accuracy failure, not a style failure, and rewording the request for tone or clarity won't address it. First find where the error came from. Compare the wrong figures against the source material. Check whether the actual source was in the prompt at all, or only described. If Claude didn't have the real numbers, no amount of rephrasing will produce them.

    Anthropic's guide to reducing hallucinations lists prompt changes that target accuracy directly. One is external knowledge restriction: explicitly instruct Claude to only use information from provided documents and not its general knowledge. Another is to explicitly give Claude permission to admit uncertainty. A third is to make the response auditable, so every claim can be traced to a quote. The guide's press-release example builds that check into the prompt:

    A verification step from Anthropic's guide: each claim must be backed by a quote from the supplied documents or removedtext
    After drafting, review each claim in your press release. For each claim, find a direct quote from the documents that supports it. If you can't find a supporting quote for a claim, remove that claim from the press release and mark where it was removed with empty [] brackets.

    The same guide describes iterative refinement as its own technique: use Claude's outputs as inputs to follow-up prompts, asking it to verify or expand on earlier statements. A follow-up turn can be used to audit the draft, not only to restyle it.

    A marketing coordinator's first attempt at a campaign brief comes back disappointing. She wants to improve the result on her next attempt. What should she do first?

    Sources3

    3.Test on varied inputs before you standardise

    Refinement has a hidden risk. If you tune a prompt over many rounds against the same input, such as one supplier's file or one report, you can end up with wording that fits that input and nothing else. It works because every round was judged on the same case. Anthropic's evaluation guidance warns against this: be task-specific, design evals that mirror your real-world task distribution, and don't forget to factor in edge cases. Its examples of edge cases include irrelevant or missing input data, overly long input, and ambiguous cases where even humans would disagree.

    Before you share a refined prompt as a team standard, run it on a spread of inputs that looks like the real work: different suppliers, a messy file, a short one, a borderline case. Then judge the results against the criteria you set. The same thinking applies to any examples inside the prompt. The best-practice guide asks for diverse examples that cover edge cases, so Claude doesn't pick up unintended patterns. For consistency, the hallucination guide describes Best-of-N verification: run the same prompt several times and compare the outputs, because inconsistencies can point to hallucinations. Rerunning a prompt this way checks it. It doesn't improve it.

    Even a machine-generated starting point is meant to be iterated on. Anthropic's metaprompt notebook produces a prompt template to solve the blank-page problem. Its own caveat says the result is not guaranteed to be optimal and invites you to change it. Whether you start from a generated draft or your own, the test is the same: does it meet your criteria across the inputs it will really see?

    Sources435

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A prompt refined over many rounds on one document is proven and ready to become the team standard.Why is that wrong?

      It has only been shown to work on that one input. Test it on inputs that mirror the real task distribution, including edge cases, before relying on it.

      Covered in Test on varied inputs before you standardise

    2. 2.If a well-written draft quotes wrong figures, the fix is to polish the wording of the request.Why is that wrong?

      Wrong facts are a grounding problem. Check the figures against the source, make sure the real material is in the prompt, and restrict Claude to it, with claims traceable to quotes.

      Covered in Wrong facts: check the grounding before the wording

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “clearly defining your success criteria and then designing evaluations to measure performance against them. This cycle is central to prompt engineering.”
      ↩︎ Decide what 'better' means before you iterate
      “Most use cases need multidimensional evaluation along several success criteria.”
      ↩︎ Decide what 'better' means before you iterate
      “Be task-specific: Design evals that mirror your real-world task distribution. Don't forget to factor in edge cases!”
      ↩︎ Exam trap 1
    2. 3.
      “Explicitly give Claude permission to admit uncertainty.”
      ↩︎ Wrong facts: check the grounding before the wording
      “Iterative refinement: Use Claude's outputs as inputs for follow-up prompts, asking it to verify or expand on previous statements.”
      ↩︎ Wrong facts: check the grounding before the wording
      “Best-of-N verification: Run Claude through the same prompt multiple times and compare the outputs.”
      ↩︎ Test on varied inputs before you standardise
      “External knowledge restriction: Explicitly instruct Claude to only use information from provided documents and not its general knowledge.”
      ↩︎ Exam trap 2
    3. 4.
      “Diverse: Cover edge cases and vary enough that Claude doesn't pick up unintended patterns.”
      ↩︎ Test on varied inputs before you standardise
    4. 5.
      “The prompt you'll get at the end is not guaranteed to be optimal by any means, so don't be afraid to change it!”
      ↩︎ Test on varied inputs before you standardise

    Ready to test yourself?

    Practise the 22 questions on this subdomain.