CertSafari
    CLAUDE-CERTIFIED-ASSOCIATE-FOUNDATIONS-CCAO-F-VAR5 · Lessons

    Domain 7 · Lesson 29/30

    Turning Feedback Into Testable Success Criteria

    Adjust approach based on feedback and results

    8 min read
    3.33% of exam
    3 sources
    Published 28 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Turn vague feedback into specific, measurable criteria before changing a prompt
    • Explain why one improved output does not prove a new approach is better
    • Choose a way to check whether a working setup is really succeeding when feedback is sparse

    Key concept

    Success criteria before changes — Before you change your approach because of feedback, turn that feedback into specific, measurable criteria. Then judge each change against those criteria, not against a single impression or a single output.

    1.Raw feedback is a symptom, not a specification

    Adjusting your approach starts with feedback. A colleague says the summaries are 'off', a client flags a draft, or a result just disappoints you. The natural reaction is to rewrite the prompt straight away. But 'off' doesn't say what is wrong. The summary could be too long, missing a figure, in the wrong tone or factually shaky. Any change you make is then a guess, and afterwards you can't tell whether the guess helped.

    Anthropic's testing guidance builds the whole prompt-engineering cycle around this step: first define what success means, then measure against it. It calls this cycle 'central to prompt engineering'. The first property of a good criterion is that it is specific: 'Clearly define what you want to achieve.' The documentation's example replaces a vague goal ('good performance') with a precise one ('accurate sentiment classification').

    So the first response to vague feedback is to get the concrete detail behind it: examples of the outputs people judged poor, and what the reader expected instead. Then restate that as criteria. 'The summaries are off' might become 'each summary names the decision taken and who owns it', or 'no summary runs past one paragraph'. Now you have something to fix and something to check.

    The same logic applies where Anthropic has automated the checking. In Claude Managed Agents, a grader scores work against a rubric. The documentation warns that the grader scores each criterion independently, 'so vague criteria produce noisy evaluations.' A human reviewer has the same problem. Vague criteria produce noisy judgements, and noisy judgements can't tell you whether a change worked.

    Several colleagues tell a project coordinator that Claude's meeting summaries 'miss the decisions'. Nobody has said which decisions or which meetings. What should she do first?

    Sources12

    2.One better answer is not a validated change

    Suppose a rewritten prompt gives a clearly better answer on the next try. It is tempting to declare victory and roll the new wording out to everyone. But you have seen one output for one input. You don't know whether the new wording is better on the other kinds of input your team will send. You also don't know whether it quietly broke something the old wording handled.

    Anthropic's evals guide says getting the best accuracy from Claude 'is an empirical science, and a process of continuous improvement'. It names the question you are really asking: whether 'a change to your prompt made the model perform better on a key metric'. One sample can't answer that. The testing documentation lists A/B testing as a method: 'Compare performance against a baseline model or earlier version.' For a prompt change, that means running the old and new wording over the same set of representative inputs and scoring both against your criteria, before the new wording becomes the team standard.

    The four parts of an eval, as described in Anthropic's evals guide
    PartWhat it is
    Input promptWhat is fed to the model, often a set of variable inputs placed into a prompt template at test time
    OutputWhat the model produces when that input prompt is run
    Golden answerWhat the output is compared to: either a required exact match or an example of a perfect answer
    ScoreA result from a grading method showing how the model did on that question

    The guide also points out where the ongoing cost is. Writing the test questions is usually a one-time cost. Grading 'is a cost you will incur every time you re-run your eval, in perpetuity', so it recommends checks that are quick to grade. Code-based checks, such as an exact match or looking for a key phrase, are 'by far the best grading method if you can design an eval that allows for it'. Human grading works for almost any task but is slow and expensive.

    Half a team says a Claude-assisted drafting pilot saves hours; the other half says the drafts need so much editing they are not worth it. What should the pilot lead do first?

    Sources13

    3.When the only feedback is 'it's fine'

    The opposite problem is silence. A workflow has run for weeks and the only comment is that it's fine. That is not evidence of success. It may mean nobody has looked closely, or that people quietly stopped using the output. The fix is the same as for vague complaints: decide what success looks like, then measure it.

    Anthropic's testing documentation says good criteria are specific, measurable, achievable and relevant. On measurability: 'Numbers provide clarity and scalability, but qualitative measures can be valuable if consistently applied along with quantitative measures.' Measuring doesn't need a complex system. The same page lists user feedback in the form of 'Implicit measures like task completion rates', meaning whether people actually get the job done with the output. It also lists Likert scales, such as rating coherence from 1 to 5.

    The four properties of good success criteria in Anthropic's testing guidance
    PropertyWhat it asksExample from the documentation
    SpecificClearly define what you want to achieve'accurate sentiment classification' instead of 'good performance'
    MeasurableQuantitative metrics, or qualitative scales applied consistentlyTask completion rates; rating coherence from 1 to 5
    AchievableBase targets on industry benchmarks, prior experiments, AI research or expert knowledgeTargets should not be unrealistic for current frontier model capabilities
    RelevantMatch the criteria to the application's purpose and user needsCitation accuracy critical for medical apps, less so for casual chatbots

    The last two properties keep the loop honest. A target that ignores what users need produces feedback that is technically met but useless. A target set beyond what the model can do produces 'failure' feedback that no prompt change will fix, and you end up adjusting in circles.

    Sources1

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.The quickest response to vague feedback is to rewrite the prompt immediately.Why is that wrong?

      Without concrete examples and specific criteria, a rewrite is a guess you can't check. Vague criteria produce noisy judgements, so first pin down what 'off' actually means.

      Covered in Raw feedback is a symptom, not a specification

    2. 2.A rewrite that fixes the one bad answer you saw is ready to become the team standard.Why is that wrong?

      One output is not a comparison. Run the new and earlier versions over representative inputs and compare them against your criteria before you standardise.

      Covered in One better answer is not a validated change

    3. 3.If nobody complains about an output, it is working.Why is that wrong?

      Silence is not a measurement. Define criteria and check them, including implicit signals such as whether people actually complete their task with the output.

      Covered in When the only feedback is 'it's fine'

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “This cycle is central to prompt engineering.”
      ↩︎ Raw feedback is a symptom, not a specification
      “Clearly define what you want to achieve.”
      ↩︎ Raw feedback is a symptom, not a specification
      “A/B testing: Compare performance against a baseline model or earlier version.”
      ↩︎ One better answer is not a validated change
      “Numbers provide clarity and scalability, but qualitative measures can be valuable if consistently applied along with quantitative measures.”
      ↩︎ When the only feedback is 'it's fine'
      “Your success metrics should not be unrealistic to current frontier model capabilities.”
      ↩︎ When the only feedback is 'it's fine'
      “Building a successful LLM-based application starts with clearly defining your success criteria and then designing evaluations to measure performance against them.”
      ↩︎ Key concept
      “A/B testing: Compare performance against a baseline model or earlier version.”
      ↩︎ Exam trap 2
      “User feedback: Implicit measures like task completion rates.”
      ↩︎ Exam trap 3
    2. 2.
      “The grader scores each criterion independently, so vague criteria produce noisy evaluations.”
      ↩︎ Raw feedback is a symptom, not a specification
      “The grader scores each criterion independently, so vague criteria produce noisy evaluations.”
      ↩︎ Exam trap 1
    3. 3.
      “Whether you are trying to know if a change to your prompt made the model perform better on a key metric”
      ↩︎ One better answer is not a validated change
      “Grading on the other hand is a cost you will incur every time you re-run your eval, in perpetuity”
      ↩︎ One better answer is not a validated change

    Continue to page 2 of 2

    Iterating on Feedback and Making Fixes Stick