What you will be able to do
- Start from a minimal version and change it based on the failures you observe
- Measure every revision against the same fixed evaluation set so you can tell which change helped or hurt
- Decide between keeping a new approach and returning to the old one by comparing both on the same criteria
- Write rubric criteria that a grader can score, and use steering to redirect long-running work
1.Start minimal and let failures drive the changes
Iterating on a Claude solution works best from a small starting point. Anthropic's context-engineering guidance suggests first testing a minimal prompt with the best available model to see how it does on your task, then adding clear instructions and examples to fix the failure modes that initial testing reveals. Each addition should answer a specific observed failure. Adding rules because they might help makes the prompt harder to maintain and does not tell you whether any of them worked.
Before you edit anything, work out what kind of failure you are looking at. The prompt engineering guide points out that prompting cannot fix every failing criterion. If the problem is latency or cost, choosing a different model may be the easier fix. Treating every miss as a wording problem wastes cycles on the wrong lever.
A project lead is demonstrating a new Claude-assisted reporting workflow to executives who could fund its wider use. What should the demonstration make clear alongside the benefits?
Correct answer: C — Where the workflow still needs a person to check the output, and the kinds of error it can produce.
- A. Design effort is a cost already incurred and tells executives nothing about whether the workflow is dependable. It answers a question about the team rather than about the decision being funded.
- B. A tool comparison explains how the team arrived here, which may be of passing interest. It does not tell executives what they will be relying on or where it can go wrong.
- C. Correct. Honest communication of capability and limitation is what makes the funding decision sound: executives need to know which checks remain human and what kinds of error to expect, or they will plan for a level of automation the workflow does not deliver.
- D. Showing the instructions is transparent about the mechanics but not about the risk. Executives are not well placed to infer failure modes from wording, and detail here can create false reassurance.
2.Measure every revision against the same evaluation set
Iteration only works if you can see what each change did. Build the evaluation set once, from the success criteria, and run it again after every revision. Anthropic's advice on building that set: make it task-specific so it mirrors the real distribution of work, and include edge cases such as irrelevant or missing input, very long input, and ambiguous cases where even humans would disagree. Automate grading where you can (exact match, code-graded, LLM-graded). Anthropic prefers more test cases with slightly noisier automated grading over a few hand-graded ones.
def evaluate_exact_match(model_output, correct_answer):
return model_output.strip().lower() == correct_answer.lower()Consider a team that changes its arrangement several times without a fixed test set. When one version gets worse on some reports, nobody can tell which change caused it, because nothing recorded how each version performed on the same inputs. Running the same set after every change makes each result traceable to that change. The same discipline answers the keep-or-revert question. Anthropic lists A/B testing, comparing performance against a baseline or an earlier version, as a quantitative method. The fair comparison is the new arrangement against the previous one, on the same inputs, scored on the criteria you defined before the pilot.
| Method | What it measures | Iteration use |
|---|---|---|
| A/B testing | Performance against a baseline model or earlier version | Keep the new version, or go back to the old one? |
| User feedback | Implicit measures like task completion rates | Is the change helping real users finish their work? |
| Edge case analysis | Percentage of edge cases handled without errors | Did a fix for common inputs break the rare ones? |
| Likert scales / expert rubrics | Qualitative quality on a defined scale | Did tone, coherence or domain quality move? |
In a pilot of a new Claude-assisted briefing arrangement, three of eight testers report that the output is "not quite right" but say nothing more specific. What should the pilot's owner do first?
Correct answer: B — Ask the three testers for specific examples, with what the output said and what it should have said.
- A. Rewriting on a hunch changes the arrangement without knowing what is wrong. If the guess misses, the pilot now has two versions and still no diagnosis, and the testers' original complaint remains unexplained.
- B. Correct. "Not quite right" is not actionable, and the cheapest way to make it actionable is to ask for cases. A concrete pair of what was produced and what was wanted turns vague dissatisfaction into a specific gap the instructions can address.
- C. Learning from the satisfied testers may eventually help, but it assumes the difference lies in how the tool was used rather than in the briefings themselves. It also leaves the three unresolved complaints uninvestigated.
- D. Direct observation is thorough but slow, and pausing the pilot stops the flow of evidence just when it is needed. Asking for examples costs an email and usually locates the problem immediately.
Sources3
3.Turn "good" into a rubric, and steer long-running work
For outputs that no single metric captures, such as a report, a model or a proposal, write the definition of good as a rubric. Claude Managed Agents make this explicit. An outcome tells the session what the end result should look like and how to measure its quality. A grader then scores the artifact against a required markdown rubric and passes its feedback back to the agent for the next iteration. The grader works in a separate context window so the main agent's implementation choices do not sway it.
# DCF Model Rubric
## Revenue Projections
- Uses historical revenue data from the last 5 fiscal years
- Projects revenue for at least 5 years forward
- Growth rate assumptions are explicitly stated and reasonableEach criterion is scored on its own, so write them to be gradeable: "The CSV contains a price column with numeric values", not "The data looks good." If you do not have a rubric, Anthropic suggests a middle path. Give Claude a known-good artifact, ask it to analyse what makes it good, and turn that analysis into criteria. This is often better than writing criteria from scratch, and it is a practical way to use Claude on the design problem itself.
Long-running work also needs checkpoints where a person can step in. In Managed Agents you can send more user events to guide the agent while it is working, or interrupt it to change direction. The sources for this lesson do not include the Claude Cowork articles, so this lesson does not describe how Cowork checks in for judgment or approval. Check that in the Cowork documentation.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.When the solution misses a target, the fix is always to rewrite the prompt.Why is that wrong?
Some failing criteria are better fixed another way. Latency and cost, for example, can often be improved more easily by choosing a different model.
2.A handful of carefully hand-graded examples is a better basis for iteration than a large automated test set.Why is that wrong?
Anthropic recommends more questions with slightly lower-signal automated grading over fewer hand-graded ones, so every revision can be checked cheaply against the same broad set.
Covered in Measure every revision against the same evaluation set
3.A rubric line like "the output looks professional" is enough for a grader to judge quality.Why is that wrong?
The grader scores each criterion independently, so criteria have to be explicit and gradeable. Vague ones give inconsistent scores.
Covered in Turn "good" into a rubric, and steer long-running work
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“start by testing a minimal prompt with the best model available to see how it performs on your task”
↩︎ Start minimal and let failures drive the changes“add clear instructions and examples to improve performance based on failure modes found during initial testing”
↩︎ Start minimal and let failures drive the changes - 2.
“For example, you can sometimes improve latency and cost more easily by selecting a different model.”
↩︎ Start minimal and let failures drive the changes“Not every success criteria or failing eval is best solved by prompt engineering.”
↩︎ Exam trap 1 - 3.
“Design evals that mirror your real-world task distribution.”
↩︎ Measure every revision against the same evaluation set“A/B testing: Compare performance against a baseline model or earlier version.”
↩︎ Measure every revision against the same evaluation set“Automate when possible: Structure questions to allow for automated grading”
↩︎ Measure every revision against the same evaluation set“Prioritize volume over quality”
↩︎ Exam trap 2 - 4.
“An outcome tells the session what the end result should look like and how to measure its quality.”
↩︎ Turn "good" into a rubric, and steer long-running work“The grader uses a separate context window to avoid being influenced by the main agent's implementation choices.”
↩︎ Turn "good" into a rubric, and steer long-running work“giving Claude an example of a known-good artifact and asking it to analyze what makes that content good”
↩︎ Turn "good" into a rubric, and steer long-running work“vague criteria produce noisy evaluations.”
↩︎ Exam trap 3 - 5.
“Send additional user events to guide the agent mid-execution, or interrupt it to change direction.”
↩︎ Turn "good" into a rubric, and steer long-running work