CertSafari
    CLAUDE-CERTIFIED-ASSOCIATE-FOUNDATIONS-CCAO-F-VAR5 · Lessons

    Domain 4 · Lesson 17/30

    Iterating on a Claude Solution with Evals and Rubrics

    Use Claude to support solution design, development, and iteration

    7 min read
    3.2% of exam
    5 sources
    Published 28 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Start from a minimal version and change it based on the failures you observe
    • Measure every revision against the same fixed evaluation set so you can tell which change helped or hurt
    • Decide between keeping a new approach and returning to the old one by comparing both on the same criteria
    • Write rubric criteria that a grader can score, and use steering to redirect long-running work

    1.Start minimal and let failures drive the changes

    Iterating on a Claude solution works best from a small starting point. Anthropic's context-engineering guidance suggests first testing a minimal prompt with the best available model to see how it does on your task, then adding clear instructions and examples to fix the failure modes that initial testing reveals. Each addition should answer a specific observed failure. Adding rules because they might help makes the prompt harder to maintain and does not tell you whether any of them worked.

    Before you edit anything, work out what kind of failure you are looking at. The prompt engineering guide points out that prompting cannot fix every failing criterion. If the problem is latency or cost, choosing a different model may be the easier fix. Treating every miss as a wording problem wastes cycles on the wrong lever.

    A project lead is demonstrating a new Claude-assisted reporting workflow to executives who could fund its wider use. What should the demonstration make clear alongside the benefits?

    Sources12

    2.Measure every revision against the same evaluation set

    Iteration only works if you can see what each change did. Build the evaluation set once, from the success criteria, and run it again after every revision. Anthropic's advice on building that set: make it task-specific so it mirrors the real distribution of work, and include edge cases such as irrelevant or missing input, very long input, and ambiguous cases where even humans would disagree. Automate grading where you can (exact match, code-graded, LLM-graded). Anthropic prefers more test cases with slightly noisier automated grading over a few hand-graded ones.

    The simplest automated grader: normalise the output and compare it with the labelled answerpython
    def evaluate_exact_match(model_output, correct_answer):
        return model_output.strip().lower() == correct_answer.lower()

    Consider a team that changes its arrangement several times without a fixed test set. When one version gets worse on some reports, nobody can tell which change caused it, because nothing recorded how each version performed on the same inputs. Running the same set after every change makes each result traceable to that change. The same discipline answers the keep-or-revert question. Anthropic lists A/B testing, comparing performance against a baseline or an earlier version, as a quantitative method. The fair comparison is the new arrangement against the previous one, on the same inputs, scored on the criteria you defined before the pilot.

    Measurement methods Anthropic lists, and the iteration question each helps answer
    MethodWhat it measuresIteration use
    A/B testingPerformance against a baseline model or earlier versionKeep the new version, or go back to the old one?
    User feedbackImplicit measures like task completion ratesIs the change helping real users finish their work?
    Edge case analysisPercentage of edge cases handled without errorsDid a fix for common inputs break the rare ones?
    Likert scales / expert rubricsQualitative quality on a defined scaleDid tone, coherence or domain quality move?

    In a pilot of a new Claude-assisted briefing arrangement, three of eight testers report that the output is "not quite right" but say nothing more specific. What should the pilot's owner do first?

    Sources3

    3.Turn "good" into a rubric, and steer long-running work

    For outputs that no single metric captures, such as a report, a model or a proposal, write the definition of good as a rubric. Claude Managed Agents make this explicit. An outcome tells the session what the end result should look like and how to measure its quality. A grader then scores the artifact against a required markdown rubric and passes its feedback back to the agent for the next iteration. The grader works in a separate context window so the main agent's implementation choices do not sway it.

    Part of Anthropic's example rubric: each line is a criterion that can be checked on its ownmarkdown
    # DCF Model Rubric
    
    ## Revenue Projections
    - Uses historical revenue data from the last 5 fiscal years
    - Projects revenue for at least 5 years forward
    - Growth rate assumptions are explicitly stated and reasonable

    Each criterion is scored on its own, so write them to be gradeable: "The CSV contains a price column with numeric values", not "The data looks good." If you do not have a rubric, Anthropic suggests a middle path. Give Claude a known-good artifact, ask it to analyse what makes it good, and turn that analysis into criteria. This is often better than writing criteria from scratch, and it is a practical way to use Claude on the design problem itself.

    Long-running work also needs checkpoints where a person can step in. In Managed Agents you can send more user events to guide the agent while it is working, or interrupt it to change direction. The sources for this lesson do not include the Claude Cowork articles, so this lesson does not describe how Cowork checks in for judgment or approval. Check that in the Cowork documentation.

    Sources45

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.When the solution misses a target, the fix is always to rewrite the prompt.Why is that wrong?

      Some failing criteria are better fixed another way. Latency and cost, for example, can often be improved more easily by choosing a different model.

      Covered in Start minimal and let failures drive the changes

    2. 2.A handful of carefully hand-graded examples is a better basis for iteration than a large automated test set.Why is that wrong?

      Anthropic recommends more questions with slightly lower-signal automated grading over fewer hand-graded ones, so every revision can be checked cheaply against the same broad set.

      Covered in Measure every revision against the same evaluation set

    3. 3.A rubric line like "the output looks professional" is enough for a grader to judge quality.Why is that wrong?

      The grader scores each criterion independently, so criteria have to be explicit and gradeable. Vague ones give inconsistent scores.

      Covered in Turn "good" into a rubric, and steer long-running work

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “start by testing a minimal prompt with the best model available to see how it performs on your task”
      ↩︎ Start minimal and let failures drive the changes
      “add clear instructions and examples to improve performance based on failure modes found during initial testing”
      ↩︎ Start minimal and let failures drive the changes
    2. 2.
      “For example, you can sometimes improve latency and cost more easily by selecting a different model.”
      ↩︎ Start minimal and let failures drive the changes
      “Not every success criteria or failing eval is best solved by prompt engineering.”
      ↩︎ Exam trap 1
    3. 3.
      “Design evals that mirror your real-world task distribution.”
      ↩︎ Measure every revision against the same evaluation set
      “A/B testing: Compare performance against a baseline model or earlier version.”
      ↩︎ Measure every revision against the same evaluation set
      “Automate when possible: Structure questions to allow for automated grading”
      ↩︎ Measure every revision against the same evaluation set
      “Prioritize volume over quality”
      ↩︎ Exam trap 2
    4. 4.
      “An outcome tells the session what the end result should look like and how to measure its quality.”
      ↩︎ Turn "good" into a rubric, and steer long-running work
      “The grader uses a separate context window to avoid being influenced by the main agent's implementation choices.”
      ↩︎ Turn "good" into a rubric, and steer long-running work
      “giving Claude an example of a known-good artifact and asking it to analyze what makes that content good”
      ↩︎ Turn "good" into a rubric, and steer long-running work
      “vague criteria produce noisy evaluations.”
      ↩︎ Exam trap 3
    5. 5.
      “Send additional user events to guide the agent mid-execution, or interrupt it to change direction.”
      ↩︎ Turn "good" into a rubric, and steer long-running work