What you will be able to do
- Run the observe-refine-test cycle to improve prompts and instructions from observed failures
- Check each candidate change against a fixed test set so regressions show up before release
- Fix the general rule behind a failure rather than patching the single failing case
- Keep improving after launch using production data
1.The loop: observe, refine, test again
The Building Evals cookbook says optimizing Claude for accuracy 'is an empirical science, and a process of continuous improvement'. No single A/B test finishes the job. Each result shows you something to change next. Anthropic's Agent Skills best practices spell out the loop. Give the agent real tasks. Note where it struggles, succeeds or makes unexpected choices. Revise the instructions based on what you saw. Then test again on similar requests.
The same guidance gives a worked observation. An agent wrote a regional sales query but forgot to filter out test accounts, even though its instructions contained that rule. The fix went after the cause, not the one output: make the rule more prominent, use stronger wording such as 'MUST filter', or restructure the workflow. The next test showed whether the change worked. The guidance says to 'Continue this observe-refine-test cycle as you encounter new scenarios', and that each iteration improves the result 'based on real agent behavior, not assumptions.'
Check how the rule is presented, not only whether it is there. The suggested fixes are to make it more prominent, use stronger language ('MUST filter' instead of 'always filter'), or restructure the workflow section. Then test again on similar requests to confirm the change worked.
2.Test each change against a fixed set to catch regressions
Every change carries a risk: it might fix the case you were looking at and break something else. The academy course on shipping AI products puts it plainly: 'Tests make post-launch iteration safe.' The classification cookbook lists 'Regression detection: Ensuring prompt changes don't accidentally hurt accuracy' as a requirement for production evaluation, along with version control to track prompt performance over time. So each iteration reruns the same fixed cases, and the new version has to hold its score on all of them, not only on the case that prompted the change.
{
"skills": ["pdf-processing"],
"query": "Extract all text from this PDF file and save it to output.txt",
"files": ["test-files/document.pdf"],
"expected_behavior": [
"Successfully reads the PDF file using an appropriate PDF processing library or command-line tool",
"Extracts text content from all pages in the document without missing any pages",
"Saves the extracted text to a file named output.txt in a clear, readable format"
]
}An Anthropic write-up for startups describes a production version of this loop. Every change is made against a versioned set of instructions and tested against the records that failed. Then 'It runs the candidate change across a golden set plus random samples and surfaces any regressions before anything ships.' The team's first attempt went wrong in a telling way: 'Our first version overfitted.' It 'fixed' failures by encoding the specific case, so patches piled up and the agent got no better in general. They fixed this by requiring general principles and capping how many specifics a change could add. Their rule: 'fix the principle, not the example'.
Cost matters too, because this loop runs many times. The Building Evals cookbook notes that grading is 'a cost you will incur every time you re-run your eval, in perpetuity'. It recommends making evals quick and cheap to grade, since you will run them on every iteration.
A data science lead is designing an eval suite to compare prompt candidates before an A/B test, and a colleague suggests spending the team's limited time hand-labeling 80 carefully reviewed examples rather than building an automated grader over a larger set. Following the recommended approach to eval volume versus grading quality, how should the lead respond?
Correct answer: A — Favor a larger automatically graded set even with somewhat lower per-item signal, since more questions with automated grading generally outweigh fewer high-quality hand-graded ones
- A. Correct. The recommended practice is to prioritize volume with automated grading, since more questions with slightly lower-signal automated grading give a more reliable read on performance than a small number of high-quality hand-graded examples.
- B. Incorrect. Hand-grading quality does not automatically outweigh the statistical benefit of testing across a larger, automatically graded sample.
- C. Incorrect. Shrinking the eval set to enable full manual review sacrifices coverage and volume, working against reliable prompt comparisons.
- D. Incorrect. Relying on subjective impressions with no test set removes the empirical basis needed to compare prompt versions objectively.
3.Keep improving after launch
Launch is one more round of the loop, not the end of it. Anthropic's content moderation guide says to regularly assess the system using metrics such as precision and recall, and to 'Use this data to iteratively refine your moderation prompts, keywords, and assessment criteria.' The academy course ends with a concrete step: share the app with three real users and 'make one iteration based on what you learn, with tests that verify the change.'
The agent-evals post explains where post-launch signals come from. 'Production monitoring kicks in post-launch to detect distribution drift and unanticipated real-world failures.' Triage user feedback constantly, and sample transcripts to read every week. Each failure you find can become a new case in the fixed test set, so the next change is tested against it. The startup write-up recommends keeping several eval sets for key use cases and updating them regularly, so teams 'can prevent drift and evaluate future models.'
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.When a record fails, the quickest good fix is to add instructions that handle that exact record.Why is that wrong?
Encoding specific cases overfits: patches pile up and the system gets no better in general. Revise the general rule behind the failure, then check it against a golden set for regressions.
Covered in Test each change against a fixed set to catch regressions
2.Once a prompt passes its evaluation and ships, the iteration work is done.Why is that wrong?
Keep measuring performance after launch and use that data to refine prompts and criteria. Iteration continues in production.
Covered in Keep improving after launch
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/cookbook/misc-building-evalsSecondary source
“is an empirical science, and a process of continuous improvement”
↩︎ The loop: observe, refine, test again“Grading on the other hand is a cost you will incur every time you re-run your eval, in perpetuity”
↩︎ Test each change against a fixed set to catch regressions - 2.
“Continue this observe-refine-test cycle as you encounter new scenarios.”
↩︎ The loop: observe, refine, test again“Each iteration improves the Skill based on real agent behavior, not assumptions.”
↩︎ The loop: observe, refine, test again“using stronger language such as "MUST filter" instead of "always filter," or restructuring the workflow section”
↩︎ The loop: observe, refine, test again - 3.
“Tests make post-launch iteration safe.”
↩︎ Test each change against a fixed set to catch regressions“make one iteration based on what you learn, with tests that verify the change.”
↩︎ Keep improving after launch - 4.
“Regression detection: Ensuring prompt changes don't accidentally hurt accuracy”
↩︎ Test each change against a fixed set to catch regressions - 5.https://claude.com/blog/claude-code-guide-for-startupsSecondary source
“Every change is made against a versioned set of instructions and tested against the records that failed.”
↩︎ Test each change against a fixed set to catch regressions“It runs the candidate change across a golden set plus random samples and surfaces any regressions before anything ships.”
↩︎ Test each change against a fixed set to catch regressions“Our first version overfitted.”
↩︎ Test each change against a fixed set to catch regressions“update them regularly, so they can prevent drift and evaluate future models.”
↩︎ Keep improving after launch“fix the principle, not the example”
↩︎ Exam trap 1 - 6.
“Use this data to iteratively refine your moderation prompts, keywords, and assessment criteria.”
↩︎ Keep improving after launch“Use this data to iteratively refine your moderation prompts, keywords, and assessment criteria.”
↩︎ Exam trap 2 - 7.
“Production monitoring kicks in post-launch to detect distribution drift and unanticipated real-world failures.”
↩︎ Keep improving after launch