What you will be able to do
- Turn a vague quality goal into specific, measurable success criteria across several dimensions
- Name the parts of an eval (input, output, golden answer, score) and the agent-eval terms task, trial, transcript and outcome
- Build a test set that matches real-world traffic, includes edge cases and favours volume
- Use Claude to generate synthetic test inputs and draft golden answers
Key concept
Multidimensional evaluation — An eval scores a system against several separate success criteria (such as accuracy, safety, latency and cost), each with its own measurable target. Each criterion is then graded by whichever method can judge it reliably.
1.Start from success criteria, not from test cases
You can't design an evaluation dataset until you know what it is supposed to measure. Anthropic's guidance is to define success criteria first and then design evals against them. Good criteria are specific: "accurate sentiment classification" rather than "good performance". They are measurable, with quantitative metrics or well-defined qualitative scales. They are achievable, based on benchmarks or prior experiments. And they are relevant to what users actually need. Qualitative measures aren't banned. They are useful if you apply them consistently alongside quantitative ones, and even hazy topics like safety can be put into numbers, for example as a toxicity-flag rate over 10,000 trials.
One number is rarely enough. The guidance lists dimensions worth considering: task fidelity (including rare or hard inputs), consistency across similar inputs, relevance and coherence, tone and style, privacy preservation, use of context, latency and price. A real criterion combines several of them.
| Dimension | Measurable target on the held-out set |
|---|---|
| Task fidelity | F1 score of at least 0.85 |
| Safety | 99.5% of outputs are non-toxic |
| Error severity | 90% of errors would cause inconvenience, not egregious error |
| Latency | 95% response time < 200ms |
Sources1
2.The anatomy of an eval
Anthropic's evals cookbook says an eval has four parts. First, an input prompt, often a template filled with variable inputs. Second, the output the model produces. Third, a golden answer to compare it with. Fourth, a score from a grading method. The golden answer can be a required exact match, or an example of an ideal answer that gives a grader something to compare against. Which kind you write determines which graders you can use later.
tweets = [
{"text": "This movie was a total waste of time. 👎", "sentiment": "negative"},
{"text": "The new album is 🔥! Been on repeat all day.", "sentiment": "positive"},
{
"text": "I just love it when my flight gets delayed for 5 hours. #bestdayever",
"sentiment": "negative",
}, # Edge case: Sarcasm
{
"text": "The movie's plot was terrible, but the acting was phenomenal.",
"sentiment": "mixed",
}, # Edge case: Mixed sentiment
# ... 996 more tweets
]Agent evals need more vocabulary. A task is one test with defined inputs and success criteria. Each attempt at it is a trial, and you run several trials because outputs vary from run to run. The transcript is the full record of a trial: outputs, tool calls, reasoning and intermediate results. The outcome is the final state of the environment. A flight-booking agent can say the flight is booked, but the outcome is whether a reservation actually exists in the database. A collection of tasks with a shared goal, such as refunds, cancellations and escalations for a support agent, is an evaluation suite.
3.What goes into the test set
The first rule is to be task-specific: the test set should mirror the real distribution of inputs your system will see. Clean, typical inputs aren't enough. The documentation names edge cases to include on purpose: irrelevant or nonexistent input data, overly long input, poor, harmful or irrelevant user input in chat use cases, and ambiguous cases where even humans would find it hard to agree on an assessment. The sarcasm and mixed-sentiment rows above are there for this reason.
The shape of the dataset follows from the criterion. A consistency criterion asks whether a user who asks the same question twice gets semantically similar answers. That needs grouped data, not independent rows. The documentation's consistency example uses 50 groups, each with a few paraphrased versions of one question, and compares answers within each group by cosine similarity.
The second rule surprises people: prioritise volume over quality. Many test cases graded automatically, each with slightly lower signal, beat a few carefully hand-graded ones. This is also why the documentation says to write questions in forms that can be graded automatically, such as multiple choice, string match, code-graded or LLM-graded. How you write the test set now determines how cheaply you can grade it on every future run.
A team is defining the success criteria for a new claims-summarization feature before building the evaluation dataset. The current draft criterion reads: "the summaries should be concise and helpful." Which revised criterion best follows the SMART framework for evaluation design?
Correct answer: A — Require summaries to average at least 4 out of 5 on a completeness rubric across a 500-claim sample
- A. Correct. This criterion is specific (completeness rubric), measurable (numeric threshold), achievable (a 4/5 average rather than perfection), and relevant (tied directly to summary quality), and it defines a concrete sample size for measurement.
- B. Incorrect. Requiring a perfect score from every adjuster is not achievable in practice; SMART criteria should be realistic targets grounded in frontier model capability, not unattainable perfection.
- C. Incorrect. "Sound less robotic" and "read more naturally" are not measurable without a defined metric or scale, so this fails the Measurable requirement of SMART criteria.
- D. Incorrect. This describes a manual sign-off process, not a measurable success criterion for the model's output quality, so it does not give the evaluation team a quantifiable target.
Writing the test set has one more benefit. Two engineers can read the same spec and disagree on how edge cases should be handled. Once the expected behaviour is written down as test cases, that disagreement has to be settled.
4.Generating test inputs and golden answers
The cookbook describes two costs in evals: writing questions and golden answers, and grading. Writing is usually a one-time cost. When you have no real inputs, or aren't allowed to test on the ones you have for privacy reasons, Claude can generate them. The synthetic test data cookbook fills a prompt template's variables (for example DOCUMENTS and QUESTION for a support bot) with realistic values.
Before writing each test case, the generator plans: it analyses each variable, including who would write it, its tone, its format and its typical length. You can edit that plan to push the data towards reality. In the cookbook, the plan was changed to produce numbered policy lines, and later cases were asked to differ from earlier ones. For golden answers you can write them yourself, or have Claude draft one and then edit it.
The data may not match your real task distribution: queries that are too formal, documents that are too short, or no messy cases at all. The planning step (tone, length, format of each variable) is where you line generation up with real traffic. That is the same task-specific rule the documentation gives for any test set.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A small, carefully hand-graded test set gives a better eval than a large automatically graded one.Why is that wrong?
The documentation says the opposite: favour volume. Many automatically graded cases with slightly lower signal each beat a few hand-graded ones.
Covered in What goes into the test set
2.An evaluation dataset should hold clean, typical inputs so the scores reflect normal use.Why is that wrong?
The test set should mirror the real distribution of inputs, and that includes irrelevant, overly long, harmful and ambiguous inputs. Leave them out and your scores will overstate quality.
Covered in What goes into the test set
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Use quantitative metrics or well-defined qualitative scales.”
↩︎ Start from success criteria, not from test cases“Most use cases need multidimensional evaluation along several success criteria.”
↩︎ Start from success criteria, not from test cases“On a held-out test set of 10,000 diverse Twitter posts, the sentiment analysis model should achieve:”
↩︎ Start from success criteria, not from test cases“Example eval test cases: 1,000 tweets with human-labeled sentiments.”
↩︎ The anatomy of an eval“Ambiguous test cases where even humans would find it hard to reach an assessment consensus”
↩︎ What goes into the test set“Example eval test cases: 50 groups with a few paraphrased versions each.”
↩︎ What goes into the test set“Be task-specific: Design evals that mirror your real-world task distribution.”
↩︎ Generating test inputs and golden answers“Most use cases need multidimensional evaluation along several success criteria.”
↩︎ Key concept“Prioritize volume over quality: More questions with slightly lower signal automated grading is better than fewer questions with high-quality human hand-graded evals.”
↩︎ Exam trap 1“Be task-specific: Design evals that mirror your real-world task distribution. Don't forget to factor in edge cases!”
↩︎ Exam trap 2 - 2.https://platform.claude.com/cookbook/misc-building-evalsSecondary source
“The golden answer could be a mandatory exact match, or it could be an example of a perfect answer”
↩︎ The anatomy of an eval - 3.
“Because model outputs vary between runs, we run multiple trials to produce more consistent results.”
↩︎ The anatomy of an eval“The outcome is the final state in the environment at the end of the trial.”
↩︎ The anatomy of an eval“Two engineers reading the same initial spec could come away with different interpretations on how the AI should handle edge cases.”
↩︎ What goes into the test set - 4.https://platform.claude.com/cookbook/misc-generate-test-casesSecondary source
“maybe you aren't allowed to test on the ones you do have for privacy reasons”
↩︎ Generating test inputs and golden answers“To get golden answers, you can either write them yourself from scratch, or have Claude write an answer and then edit it to taste.”
↩︎ Generating test inputs and golden answers