CertSafari
    CCAR-P · Lessons

    Domain 4 · Lesson 20/38

    Defining LLM Success Criteria: Accuracy, Safety and Security Metrics

    Define evaluation metrics

    8 min read
    2.67% of exam
    3 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Rewrite a vague quality goal as a specific, measurable, achievable and relevant success criterion
    • Choose an accuracy metric that fits the shape of the task, from exact match to F1 or a Likert scale
    • Express safety and sensitive-information handling as measurable rates over a defined number of trials

    Key concept

    Success criteria — Success criteria are the written, measurable targets an LLM application has to hit. You set them before you build any evaluation, and every metric you track exists to show whether one of them is met.

    1.From a vague goal to a success criterion

    An evaluation metric answers a question, so you have to write the question down first. Anthropic's testing guidance makes this the first step of the prompt-engineering cycle: define what success means, then design evaluations that measure performance against it. A metric that doesn't serve a stated criterion is just a number on a dashboard.

    The guidance names four properties of a good criterion. It is specific: it names the capability, such as "accurate sentiment classification" rather than "good performance". It is measurable: it uses a quantitative metric or a well-defined qualitative scale. Numbers give clarity and scale, and qualitative measures still help when you apply them consistently alongside the numbers. It is achievable: the target is based on industry benchmarks, prior experiments, AI research or expert knowledge, and doesn't ask for more than current frontier models can do. It is relevant: it fits the application's purpose. Citation accuracy can be critical in a medical app and matter much less in a casual chatbot.

    The same goal written as a bad and a good criterion (sentiment-analysis example from the testing guide)
    VersionCriterionWhat it gives you
    BadThe model should classify sentiments wellNothing to measure, no threshold, and no test set
    GoodF1 score of at least 0.85 on a held-out test set of 10,000 diverse Twitter posts, a 5% improvement over the current baselineA metric and threshold (measurable, specific), a representative data set (relevant), and a target anchored to today's baseline (achievable)

    Sources1

    2.Accuracy: pick the metric that fits the task

    "Accuracy" is a family of metrics, not one metric. The testing guide sorts quantitative metrics into task-specific and generic ones, and adds qualitative scales for outputs that can't be scored mechanically. Which one you use depends on what the output looks like.

    Accuracy-style metrics named in the testing guide, and the output each fits
    FamilyExamplesFits outputs that are…
    Generic quantitativeAccuracy, precision, recallLabels or decisions with a known correct answer
    Task-specific quantitativeF1 score, BLEU score, perplexityTasks with an established scoring convention
    Edge-case ratePercentage of edge cases handled without errorsRare or difficult inputs you've deliberately added to the test set
    Qualitative scaleLikert scale (coherence 1–5), expert rubricsOpen-ended text judged against defined criteria

    The simplest accuracy metric is exact match. The model's output is compared with a predefined correct answer, usually after normalizing whitespace and case. It works for categorical answers such as positive, negative or neutral sentiment. The guide's example runs a classification prompt over a set of human-labelled tweets and reports the share that match:

    Exact-match accuracy over a labelled sentiment setpython
    def evaluate_exact_match(model_output, correct_answer):
        return model_output.strip().lower() == correct_answer.lower()
    
    
    outputs = [
        get_completion(
            f"Classify this as 'positive', 'negative', 'neutral', or 'mixed': {tweet['text']}"
        )
        for tweet in tweets
    ]
    accuracy = sum(
        evaluate_exact_match(output, tweet["sentiment"])
        for output, tweet in zip(outputs, tweets)
    ) / len(tweets)
    print(f"Sentiment Analysis Accuracy: {accuracy * 100}%")

    Two further points shape an accuracy criterion. First, the threshold follows the stakes. Anthropic's customer-support guide sets a 100% target for the accuracy of quoted information, because a wrong quote breaks the interaction. The same guide asks for escalation accuracy of 95% or higher. Second, some criteria are about consistency rather than correctness: similar questions should get semantically similar answers. The guide measures this with cosine similarity between sentence embeddings, where values closer to 1 mean more similar answers.

    A fintech company built an automated support-ticket triage system that assigns each incoming ticket to exactly one of five fixed categories (billing, fraud, account access, technical issue, other). The team wants to measure how well the model performs this core categorization task using a labeled test set of 5,000 tickets, with automated, unambiguous grading. Which evaluation approach should the team use?

    Sources12

    3.Safety and security: making the "hazy" criteria measurable

    Teams often leave safety out of their metrics because it feels subjective. The testing guide says the opposite: ethics and safety can be quantified too. Its bad version of a safety criterion is simply "safe outputs". The good version sets a rate over a fixed number of trials, using a named detector: fewer than 0.1% of outputs out of 10,000 trials flagged for toxicity by the content filter.

    The guide's multidimensional example adds two more safety-style targets: 99.5% of outputs are non-toxic, and 90% of errors would cause inconvenience rather than egregious error. The second one measures how bad a mistake is, not only how often mistakes happen. The guide adds that you would also have to define what "inconvenience" and "egregious" mean before anyone can grade against them. For chat use cases, the test set should also include poor, harmful or irrelevant user input, so the safety rate reflects hostile inputs as well as friendly ones.

    The provided sources frame security-related criteria mainly as the handling of sensitive information. The testing guide lists this as a criterion to define: how does the model handle personal or sensitive information, and can it follow instructions not to use or share certain details? You can turn that into a rate, just like toxicity: the share of test cases in which protected details stay unshared. For agent evaluations, Anthropic lists static analysis (lint, type, security) among the code-based checks a grader can run on what an agent produces. Beyond these two, the sources don't define a dedicated security metric, so this lesson doesn't invent one.

    A documentation team built a tool that condenses long internal incident reports into short summaries. They have 200 incident reports paired with human-written reference summaries and want an automated metric that captures both the overlap of key information and the ordering of ideas between the model's summary and the reference. Which metric is best suited to this evaluation?

    Sources13

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1."The model should classify sentiments well" is a good enough success criterion to start evaluating against.Why is that wrong?

      A criterion has to name the specific capability and a measurable threshold on a defined test set. Otherwise nothing tells you whether a change helped.

      Covered in From a vague goal to a success criterion

    2. 2.Safety can't be quantified, so the success criterion should just say outputs must be safe.Why is that wrong?

      The guidance quantifies safety as a flag rate over a fixed number of trials, measured by a content filter.

      Covered in Safety and security: making the "hazy" criteria measurable

    3. 3.Every application should use the same accuracy target.Why is that wrong?

      Targets follow the application's purpose and stakes. A criterion that is critical in one domain can matter much less in another.

      Covered in Accuracy: pick the metric that fits the task

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Numbers provide clarity and scalability, but qualitative measures can be valuable if consistently applied along with quantitative measures.”
      ↩︎ From a vague goal to a success criterion
      “Your success metrics should not be unrealistic to current frontier model capabilities.”
      ↩︎ From a vague goal to a success criterion
      “should achieve an F1 score of at least 0.85 (Measurable, Specific) on a held-out test set”
      ↩︎ From a vague goal to a success criterion
      “perfect for tasks with clear-cut, categorical answers like sentiment analysis (positive, negative, neutral)”
      ↩︎ Accuracy: pick the metric that fits the task
      “Edge case analysis: Percentage of edge cases handled without errors.”
      ↩︎ Accuracy: pick the metric that fits the task
      “Values closer to 1 indicate higher similarity.”
      ↩︎ Accuracy: pick the metric that fits the task
      “90% of errors would cause inconvenience, not egregious error”
      ↩︎ Safety and security: making the "hazy" criteria measurable
      “Poor, harmful, or irrelevant user input”
      ↩︎ Safety and security: making the "hazy" criteria measurable
      “Can it follow instructions not to use or share certain details?”
      ↩︎ Safety and security: making the "hazy" criteria measurable
      “Building a successful LLM-based application starts with clearly defining your success criteria and then designing evaluations to measure performance against them.”
      ↩︎ Key concept
      “Specific: Clearly define what you want to achieve.”
      ↩︎ Exam trap 1
      “Less than 0.1% of outputs out of 10,000 trials flagged for toxicity by the content filter.”
      ↩︎ Exam trap 2
      “Strong citation accuracy might be critical for medical apps but less so for casual chatbots.”
      ↩︎ Exam trap 3
    2. 2.
      “Target 100% accuracy, as this is vital information for a successful customer interaction.”
      ↩︎ Accuracy: pick the metric that fits the task
      “Aim for an escalation accuracy of 95% or higher.”
      ↩︎ Accuracy: pick the metric that fits the task

    Continue to page 2 of 2

    Latency and Cost Metrics: Response Time, Cost per Task and Thresholds