CertSafari
    CCAR-P · Lessons

    Domain 4 · Lesson 21/38

    Designing LLM Evaluation Datasets: Success Criteria, Golden Answers and Edge Cases

    Design evaluation datasets and test frameworks using mixed methodologies

    8 min read
    2.67% of exam
    4 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Turn a vague quality goal into specific, measurable success criteria across several dimensions
    • Name the parts of an eval (input, output, golden answer, score) and the agent-eval terms task, trial, transcript and outcome
    • Build a test set that matches real-world traffic, includes edge cases and favours volume
    • Use Claude to generate synthetic test inputs and draft golden answers

    Key concept

    Multidimensional evaluation — An eval scores a system against several separate success criteria (such as accuracy, safety, latency and cost), each with its own measurable target. Each criterion is then graded by whichever method can judge it reliably.

    1.Start from success criteria, not from test cases

    You can't design an evaluation dataset until you know what it is supposed to measure. Anthropic's guidance is to define success criteria first and then design evals against them. Good criteria are specific: "accurate sentiment classification" rather than "good performance". They are measurable, with quantitative metrics or well-defined qualitative scales. They are achievable, based on benchmarks or prior experiments. And they are relevant to what users actually need. Qualitative measures aren't banned. They are useful if you apply them consistently alongside quantitative ones, and even hazy topics like safety can be put into numbers, for example as a toxicity-flag rate over 10,000 trials.

    One number is rarely enough. The guidance lists dimensions worth considering: task fidelity (including rare or hard inputs), consistency across similar inputs, relevance and coherence, tone and style, privacy preservation, use of context, latency and price. A real criterion combines several of them.

    One test set, several criteria: the documentation's worked sentiment-analysis example
    DimensionMeasurable target on the held-out set
    Task fidelityF1 score of at least 0.85
    Safety99.5% of outputs are non-toxic
    Error severity90% of errors would cause inconvenience, not egregious error
    Latency95% response time < 200ms

    Sources1

    2.The anatomy of an eval

    Anthropic's evals cookbook says an eval has four parts. First, an input prompt, often a template filled with variable inputs. Second, the output the model produces. Third, a golden answer to compare it with. Fourth, a score from a grading method. The golden answer can be a required exact match, or an example of an ideal answer that gives a grader something to compare against. Which kind you write determines which graders you can use later.

    A labelled dataset from the documentation. Each row pairs an input with its golden answer, and edge cases are marked on purposepython
    tweets = [
        {"text": "This movie was a total waste of time. 👎", "sentiment": "negative"},
        {"text": "The new album is 🔥! Been on repeat all day.", "sentiment": "positive"},
        {
            "text": "I just love it when my flight gets delayed for 5 hours. #bestdayever",
            "sentiment": "negative",
        },  # Edge case: Sarcasm
        {
            "text": "The movie's plot was terrible, but the acting was phenomenal.",
            "sentiment": "mixed",
        },  # Edge case: Mixed sentiment
        # ... 996 more tweets
    ]

    Agent evals need more vocabulary. A task is one test with defined inputs and success criteria. Each attempt at it is a trial, and you run several trials because outputs vary from run to run. The transcript is the full record of a trial: outputs, tool calls, reasoning and intermediate results. The outcome is the final state of the environment. A flight-booking agent can say the flight is booked, but the outcome is whether a reservation actually exists in the database. A collection of tasks with a shared goal, such as refunds, cancellations and escalations for a support agent, is an evaluation suite.

    Sources123

    3.What goes into the test set

    The first rule is to be task-specific: the test set should mirror the real distribution of inputs your system will see. Clean, typical inputs aren't enough. The documentation names edge cases to include on purpose: irrelevant or nonexistent input data, overly long input, poor, harmful or irrelevant user input in chat use cases, and ambiguous cases where even humans would find it hard to agree on an assessment. The sarcasm and mixed-sentiment rows above are there for this reason.

    The shape of the dataset follows from the criterion. A consistency criterion asks whether a user who asks the same question twice gets semantically similar answers. That needs grouped data, not independent rows. The documentation's consistency example uses 50 groups, each with a few paraphrased versions of one question, and compares answers within each group by cosine similarity.

    The second rule surprises people: prioritise volume over quality. Many test cases graded automatically, each with slightly lower signal, beat a few carefully hand-graded ones. This is also why the documentation says to write questions in forms that can be graded automatically, such as multiple choice, string match, code-graded or LLM-graded. How you write the test set now determines how cheaply you can grade it on every future run.

    A team is defining the success criteria for a new claims-summarization feature before building the evaluation dataset. The current draft criterion reads: "the summaries should be concise and helpful." Which revised criterion best follows the SMART framework for evaluation design?

    Writing the test set has one more benefit. Two engineers can read the same spec and disagree on how edge cases should be handled. Once the expected behaviour is written down as test cases, that disagreement has to be settled.

    Sources13

    4.Generating test inputs and golden answers

    The cookbook describes two costs in evals: writing questions and golden answers, and grading. Writing is usually a one-time cost. When you have no real inputs, or aren't allowed to test on the ones you have for privacy reasons, Claude can generate them. The synthetic test data cookbook fills a prompt template's variables (for example DOCUMENTS and QUESTION for a support bot) with realistic values.

    Before writing each test case, the generator plans: it analyses each variable, including who would write it, its tone, its format and its typical length. You can edit that plan to push the data towards reality. In the cookbook, the plan was changed to produce numbered policy lines, and later cases were asked to differ from earlier ones. For golden answers you can write them yourself, or have Claude draft one and then edit it.

    Sources14

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A small, carefully hand-graded test set gives a better eval than a large automatically graded one.Why is that wrong?

      The documentation says the opposite: favour volume. Many automatically graded cases with slightly lower signal each beat a few hand-graded ones.

      Covered in What goes into the test set

    2. 2.An evaluation dataset should hold clean, typical inputs so the scores reflect normal use.Why is that wrong?

      The test set should mirror the real distribution of inputs, and that includes irrelevant, overly long, harmful and ambiguous inputs. Leave them out and your scores will overstate quality.

      Covered in What goes into the test set

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Use quantitative metrics or well-defined qualitative scales.”
      ↩︎ Start from success criteria, not from test cases
      “Most use cases need multidimensional evaluation along several success criteria.”
      ↩︎ Start from success criteria, not from test cases
      “On a held-out test set of 10,000 diverse Twitter posts, the sentiment analysis model should achieve:”
      ↩︎ Start from success criteria, not from test cases
      “Example eval test cases: 1,000 tweets with human-labeled sentiments.”
      ↩︎ The anatomy of an eval
      “Ambiguous test cases where even humans would find it hard to reach an assessment consensus”
      ↩︎ What goes into the test set
      “Example eval test cases: 50 groups with a few paraphrased versions each.”
      ↩︎ What goes into the test set
      “Be task-specific: Design evals that mirror your real-world task distribution.”
      ↩︎ Generating test inputs and golden answers
      “Most use cases need multidimensional evaluation along several success criteria.”
      ↩︎ Key concept
      “Prioritize volume over quality: More questions with slightly lower signal automated grading is better than fewer questions with high-quality human hand-graded evals.”
      ↩︎ Exam trap 1
      “Be task-specific: Design evals that mirror your real-world task distribution. Don't forget to factor in edge cases!”
      ↩︎ Exam trap 2
    2. 2.
      “The golden answer could be a mandatory exact match, or it could be an example of a perfect answer”
      ↩︎ The anatomy of an eval
    3. 3.
      “Because model outputs vary between runs, we run multiple trials to produce more consistent results.”
      ↩︎ The anatomy of an eval
      “The outcome is the final state in the environment at the end of the trial.”
      ↩︎ The anatomy of an eval
      “Two engineers reading the same initial spec could come away with different interpretations on how the AI should handle edge cases.”
      ↩︎ What goes into the test set
    4. 4.
      “maybe you aren't allowed to test on the ones you do have for privacy reasons”
      ↩︎ Generating test inputs and golden answers
      “To get golden answers, you can either write them yourself from scratch, or have Claude write an answer and then edit it to taste.”
      ↩︎ Generating test inputs and golden answers

    Continue to page 2 of 2

    Mixed-Methodology Eval Frameworks: Code-Based, Model-Based and Human Grading