CertSafari
    CLAUDE-CERTIFIED-ASSOCIATE-FOUNDATIONS-CCAO-F-VAR5 · Lessons

    Domain 7 · Lesson 29/30

    Iterating on Feedback and Making Fixes Stick

    Adjust approach based on feedback and results

    7 min read
    3.33% of exam
    5 sources
    Published 28 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Run a generate, evaluate and revise loop with explicit criteria and a stopping point
    • Decide when to follow feedback and when evidence should override it
    • Turn a correction you keep repeating into a persistent check, instruction or rubric

    1.A loop: generate, evaluate against criteria, revise

    Once feedback is written as criteria, adjusting your approach becomes a loop instead of a one-off rewrite. Anthropic's evaluator-optimizer pattern makes the loop explicit: one call generates a response, another evaluates it and gives feedback, and the cycle repeats. The pattern fits when 'LLM responses can be demonstrably improved when feedback is provided' and when the model can give meaningful feedback itself. It also relies on 'Clear evaluation criteria'.

    In the cookbook's coding example, the evaluator marks a first draft NEEDS_IMPROVEMENT and lists concrete gaps: error handling, type hints, documentation and edge cases. In its next turn, the generator turns each point into a planned change:

    The generator turns itemised evaluator feedback into specific revisionstext
    Based on the feedback, I'll improve the implementation by:
    1. Adding proper error handling with exceptions
    2. Including type hints and docstrings
    3. Adding input validation
    4. Maintaining O(1) time complexity for all operations

    The feedback was itemised and actionable, so the revision could address each point. Feedback like 'could be better' would have left the generator guessing. That is the same problem as vague feedback from people.

    Claude Managed Agents build this loop into the platform. You define an outcome, and 'The agent works toward that target, self-evaluating and iterating until the outcome is met.' A grader checks the artifact against your rubric, and 'That feedback is handed back to the agent for the next iteration.' The loop has a limit: the event that starts an outcome sets a maximum number of iterations.

    Starting an outcome: a description, a rubric and an iteration capjson
    {
      "type": "user.define_outcome",
      "description": "Build a DCF model for Costco in .xlsx",
      "rubric": { "type": "file", "file_id": "file_01..." },
      "max_iterations": 5
    }
    How each evaluation result decides what happens next in an outcome loop
    resultWhat happens next
    satisfiedSession transitions to idle
    needs_revisionAgent starts a new iteration cycle
    max_iterations_reachedOne final acknowledgment turn, then idle; no further evaluation runs
    failedSession transitions to idle; returned when the rubric does not apply to the deliverables

    The failed result carries a broader lesson. It is returned when, for example, 'the description and rubric contradict each other'. When the goal and the criteria disagree, more iterations won't converge. Fix the criteria, not the output.

    Sources12

    2.Weighing feedback against evidence

    Feedback is an input, not an order. Anthropic's guidance for its advisor tool, where a model doing the work can consult a stronger reviewer model, describes a balance that holds for any reviewer. First: 'Give the advice serious weight.' But adapt 'If you follow a step and it fails empirically, or you have primary-source evidence that contradicts a specific claim'. And don't dismiss feedback just because your own check passed: 'A passing self-test is not evidence the advice is wrong'. It may only mean your test doesn't check what the reviewer is checking.

    When feedback and evidence point in different directions, the guidance is neither to obey quietly nor to ignore quietly: 'Surface the conflict in one more advisor call', because 'a reconcile call is cheaper than committing to the wrong branch.' It also lists moments when a review is especially worth asking for: when an approach is not converging, and 'When considering a change of approach.'

    The reviewer can be the weak link too. The Managed Agents cookbook warns that 'The default failure mode is a grader that approves everything.' A check that only confirms a topic is mentioned passes almost anything. Its advice is to make the grader earn a pass: 'Require concrete evidence (a fetched page, a traced formula, a file:line reference) before the grader passes anything.' A stream of approvals is only worth as much as the check behind it.

    An assistant notices that Claude's drafts in a shared project have become weaker since a teammate edited the project instructions last week. What should she do first?

    Sources34

    3.Closing the loop: make the fix persistent

    The last step in adjusting to feedback is making sure you don't have to make the same adjustment again. The Managed Agents cookbook observes that in manual review rounds, 'most of what you say in those rounds is feedback you could have written down before the agent started.' If you give the same correction every week, it belongs somewhere persistent, not in your next message.

    Anthropic's guidance on verification loops in Claude Code describes how to do this. When you keep making the same small corrections, 'The first step is to write down everything that you find yourself doing every time'. Those written checks can be packaged as skills, 'so every session applies the same checks automatically instead of relying on a human to remember them.' Stable facts go in instruction files, for example: 'list your exact build and test commands in CLAUDE.md'. The rule of thumb: 'Anything you keep having to enforce by hand as a manual check qualifies for capture as a loop.'

    If writing the criteria is the hard part, start from a good example. The outcomes documentation suggests 'giving Claude an example of a known-good artifact and asking it to analyze what makes that content good', and then turning that analysis into a rubric. 'This middle-ground approach often produces better results than writing criteria from scratch.' Once a rubric is written, you can upload it 'for reuse across sessions'. The Claude Code guidance takes a similar approach: ask Claude for best practices, then edit them. 'Your version probably differs on a few specific points, and those differences are exactly what you want to capture.'

    Sources145

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Restating a correction each time it comes up is an acceptable way to handle a recurring problem.Why is that wrong?

      A recurring correction should be written down once and encoded so every session applies it automatically. Otherwise the fix depends on someone remembering it.

      Covered in Closing the loop: make the fix persistent

    2. 2.If my own test passes, the reviewer's objection must be wrong.Why is that wrong?

      A passing check may just not test what the reviewer is looking at. Take the feedback seriously, and bring the conflict into the open instead of silently picking a side.

      Covered in Weighing feedback against evidence

    3. 3.If the reviewer approves the output, it meets the standard.Why is that wrong?

      Lenient reviewing is the default failure mode. An approval only means something if the check required concrete evidence.

      Covered in Weighing feedback against evidence

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “The agent works toward that target, self-evaluating and iterating until the outcome is met.”
      ↩︎ A loop: generate, evaluate against criteria, revise
      “That feedback is handed back to the agent for the next iteration.”
      ↩︎ A loop: generate, evaluate against criteria, revise
      “Returned when the rubric does not apply to the deliverables, for example if the description and rubric contradict each other.”
      ↩︎ A loop: generate, evaluate against criteria, revise
      “giving Claude an example of a known-good artifact and asking it to analyze what makes that content good”
      ↩︎ Closing the loop: make the fix persistent
      “or upload it through the Files API for reuse across sessions.”
      ↩︎ Closing the loop: make the fix persistent
    2. 2.
      “LLM responses can be demonstrably improved when feedback is provided”
      ↩︎ A loop: generate, evaluate against criteria, revise
    3. 3.
      “If you follow a step and it fails empirically, or you have primary-source evidence that contradicts a specific claim”
      ↩︎ Weighing feedback against evidence
      “a reconcile call is cheaper than committing to the wrong branch.”
      ↩︎ Weighing feedback against evidence
      “A passing self-test is not evidence the advice is wrong”
      ↩︎ Exam trap 2
    4. 4.
      “Require concrete evidence (a fetched page, a traced formula, a file:line reference) before the grader passes anything.”
      ↩︎ Weighing feedback against evidence
      “And most of what you say in those rounds is feedback you could have written down before the agent started.”
      ↩︎ Closing the loop: make the fix persistent
      “The default failure mode is a grader that approves everything.”
      ↩︎ Exam trap 3
    5. 5.
      “Anything you keep having to enforce by hand as a manual check qualifies for capture as a loop.”
      ↩︎ Closing the loop: make the fix persistent
      “so every session applies the same checks automatically instead of relying on a human to remember them.”
      ↩︎ Exam trap 1

    Ready to test yourself?

    Practise the 26 questions on this subdomain.