CertSafari
    CCAR-P · Lessons

    Domain 4 · Lesson 21/38

    Mixed-Methodology Eval Frameworks: Code-Based, Model-Based and Human Grading

    Design evaluation datasets and test frameworks using mixed methodologies

    9 min read
    2.67% of exam
    4 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Explain why grading cost, more than dataset cost, should shape how you design an eval
    • Pick a code-based, model-based or human grader for each criterion and state its trade-offs
    • Combine several graders on one task using weighted, binary or hybrid scoring, and calibrate model graders against humans
    • Tell capability suites apart from regression suites, and baseline tracking from one-off scores

    1.Grading is the cost you pay forever

    The evals cookbook separates two costs. Writing questions and golden answers is usually paid once. Grading is paid every time the eval runs, and a useful eval runs often. So the cookbook's advice is to design for fast, cheap grading from the start. The documentation says the same thing: structure questions so they can be graded automatically, by multiple choice, string match, code or an LLM.

    A mixed-methodology framework follows from this. Each success criterion gets the cheapest grader that can judge it reliably. Only the criteria that cheaper methods can't judge go to more expensive ones.

    Sources12

    2.Code-based, model-based and human graders

    Code-based grading uses ordinary code, mostly string matching and regular expressions. The cookbook calls it by far the best method when an eval allows it, because it is fast and reliable. Exact match is the standard case: the output must equal a predefined answer after whitespace and case are normalised, which suits categorical tasks like sentiment labels. The cookbook's other example asks for only the number of legs an animal has, so an exact-match grader can score it.

    A code-based grader from the documentation: normalise, compare with the golden answer, then aggregate accuracy over the datasetpython
    def evaluate_exact_match(model_output, correct_answer):
        return model_output.strip().lower() == correct_answer.lower()
    
    
    outputs = [
        get_completion(
            f"Classify this as 'positive', 'negative', 'neutral', or 'mixed': {tweet['text']}"
        )
        for tweet in tweets
    ]
    accuracy = sum(
        evaluate_exact_match(output, tweet["sentiment"])
        for output, tweet in zip(outputs, tweets)
    ) / len(tweets)
    print(f"Sentiment Analysis Accuracy: {accuracy * 100}%")

    Open-ended answers can't be graded with a string match. For these, the cookbook's golden answers are descriptions of a correct answer, not literal text. For example: a workout must have 50 or more reps of pulling leg exercises and must not count squats. Another: an email request must be declined because the assistant can't send email, though offering a draft is fine. A human can grade against descriptions like these, and so can Claude given a grader prompt. The cookbook notes that Claude is highly capable of grading tasks such as tone or free-form accuracy that used to need humans.

    The three grader families and their trade-offs
    GraderExample methodsStrengthsWeaknesses
    Code-basedString match (exact, regex, fuzzy), binary tests, static analysis, outcome verification, tool calls verification, transcript analysisFast, cheap, objective, reproducible, easy to debugBrittle to valid variations, lacking in nuance, limited for subjective tasks
    Model-basedRubric-based scoring, natural language assertions, pairwise comparison, reference-based evaluation, multi-judge consensusFlexible, scalable, captures nuance, handles open-ended and freeform outputNon-deterministic, more expensive than code, requires calibration with human graders
    HumanSME review, crowdsourced judgment, spot-check sampling, A/B testing, inter-annotator agreementGold standard quality, matches expert user judgment, calibrates model-based gradersExpensive, slow, often requires human experts at scale

    An evaluation team is assembling test cases for a clause-extraction feature used on uploaded contracts. The current dataset contains only well-formatted, standard commercial leases. Following evaluation design best practices, which addition would most improve the test set's ability to predict real-world failures?

    Sources123

    3.Combining and calibrating graders

    Mixed methodology works inside a single task, not only across the suite. A task can have several graders, each with several assertions, and each grader looks at part of the transcript or the outcome. For a coding agent, a code grader might run unit tests on the outcome, while a model grader checks instruction-following in the transcript. The task's score can be weighted (the combined scores must pass a threshold), binary (every grader must pass) or a hybrid of the two. The documentation also endorses mixing measure types: qualitative scales such as Likert ratings or expert rubrics are useful when applied consistently alongside quantitative metrics.

    Model graders need anchoring. The table lists calibration with human graders as a requirement, and human review is the gold standard those graders are measured against. Descript's video-editing team shows how this works in practice. They started with manual grading, then moved to LLM graders whose criteria were set by the product team, with periodic human calibration. In a separate harness experiment, an evaluator was calibrated with few-shot examples that included detailed score breakdowns, which reduced score drift across iterations.

    Sources134

    4.Suites that keep paying off

    A framework also has to work over time. Because the task bank is fixed, you can track latency, token usage, cost per task and error rates as baselines and regression checks at no extra cost. Comparisons against a baseline or an earlier version are the quantitative method the documentation calls A/B testing. The agent-evals post separates capability (quality) evals from regression evals. Capability evals should start at a low pass rate to give the team something to improve on. Regression evals protect what already works. Descript runs the two as separate suites. Evals also speed up model upgrades: a team with a suite can assess a new model in days instead of weeks.

    A team is designing a test framework for a Claude-based tool that answers factual questions about internal HR policies using short, well-defined answers (e.g., "What is the maximum number of paid sick days?"). To keep pace with frequent policy updates, which grading approach best supports automation while remaining accurate?

    Sources13

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.An LLM grader with a good rubric can be used as it is, since the model is capable of judging quality.Why is that wrong?

      Model-based graders are non-deterministic, and their judgments need to be checked against human graders to be trusted for accuracy.

      Covered in Combining and calibrating graders

    2. 2.Human grading is the gold standard, so it should be the default for a production eval suite.Why is that wrong?

      Humans are the most capable graders, but slow and expensive on every run. Their role is calibration and spot checks, and evals should be designed to avoid depending on human grading.

      Covered in Code-based, model-based and human graders

    3. 3.If the agent's final message says the task was done, the trial passed.Why is that wrong?

      Grade the outcome, meaning the actual final state of the environment, and not only what the transcript claims. "Your flight has been booked" passes only if the reservation exists.

      Covered in Combining and calibrating graders

    Practise it for real

    Build a small mixed-methodology sentiment eval: a labelled dataset with edge cases, graded by the documentation's exact-match function, with a review of what the code grader misses.

    1. 1.Write about 20 test cases in the documentation's tweets format ({"text": ..., "sentiment": ...}), including at least one sarcastic and one mixed-sentiment example.

      Why: The test set should mirror real inputs, including the edge cases.

      You should see: A Python list of dicts, each with a golden-answer label.

    2. 2.Run each case through the classification prompt and score it with evaluate_exact_match, then compute accuracy as in the documentation's example.

      Why: Exact match is the cheapest, most reproducible grader for categorical labels.

      You should see: A printed line such as 'Sentiment Analysis Accuracy: ...%'.

    3. 3.Read every failed case and label it either a model error or a grader error, meaning a valid answer the string check rejected.

      Why: Code graders are brittle to valid variations, and this step shows you how often that happens.

      You should see: A short list of failures, each labelled model error or grader error.

    4. 4.For any grader errors, either tighten the prompt's output format or add a model-graded check that compares the output with the golden label. Then spot-check a sample of those judgments yourself.

      Why: Model graders handle freeform output but need human calibration.

      You should see: A re-run where the remaining failures are real model errors.

    Stuck? Get a nudge

    If accuracy looks suspiciously low, print a few raw outputs before changing the prompt. The model may be answering correctly in a format your grader doesn't accept.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Automate when possible: Structure questions to allow for automated grading (for example, multiple-choice, string match, code-graded, LLM-graded).”
      ↩︎ Grading is the cost you pay forever
      “Exact match evals measure whether the model's output matches a predefined correct answer, typically after normalizing whitespace and case.”
      ↩︎ Code-based, model-based and human graders
      “Numbers provide clarity and scalability, but qualitative measures can be valuable if consistently applied along with quantitative measures.”
      ↩︎ Combining and calibrating graders
      “A/B testing: Compare performance against a baseline model or earlier version.”
      ↩︎ Suites that keep paying off
    2. 2.
      “Grading on the other hand is a cost you will incur every time you re-run your eval, in perpetuity”
      ↩︎ Grading is the cost you pay forever
      “building evals that can be quickly and cheaply graded should be at the center of your design choices.”
      ↩︎ Grading is the cost you pay forever
      “This is by far the best grading method if you can design an eval that allows for it”
      ↩︎ Code-based, model-based and human graders
      “It turns out that Claude is highly capable of grading itself”
      ↩︎ Code-based, model-based and human graders
      “You should mostly try to avoid designing evals that require human grading if you can help it.”
      ↩︎ Exam trap 2
    3. 3.
      “An essential component of effective evaluation design is to choose the right graders for the job.”
      ↩︎ Code-based, model-based and human graders
      “For each task, scoring can be weighted (combined grader scores must hit a threshold), binary (all graders must pass), or a hybrid.”
      ↩︎ Combining and calibrating graders
      “They evolved from manual grading to LLM graders with criteria defined by the product team and periodic human calibration”
      ↩︎ Combining and calibrating graders
      “latency, token usage, cost per task, and error rates can be tracked on a static bank of tasks”
      ↩︎ Suites that keep paying off
      “regularly run two separate suites for quality benchmarking and regression testing”
      ↩︎ Suites that keep paying off
      “Requires calibration with human graders for accuracy”
      ↩︎ Exam trap 1
      “Each grader evaluates some portion of either the transcript or the outcome.”
      ↩︎ Exam trap 3
    4. 4.
      “I calibrated the evaluator using few-shot examples with detailed score breakdowns.”
      ↩︎ Combining and calibrating graders