CertSafari
    CCAR-P · Lessons

    Domain 4 · Lesson 22/38

    A/B Testing Prompts and Models Against a Baseline

    Conduct A/B testing and iterative improvements

    8 min read
    2.67% of exam
    4 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Write a success criterion that says how much better a new prompt or model version has to be than the current baseline
    • Run an offline A/B comparison of prompt variants on a shared test set and read what the results show
    • Decide when a change needs an online A/B test with real user traffic, and know what that test cannot tell you

    Key concept

    Baseline comparison (A/B testing) — An A/B test measures a candidate prompt, model or configuration against the version you already have, on the same measure. A score on its own tells you little. The difference from the baseline is the result.

    1.An A/B test needs a baseline and a target

    Anthropic's guidance on success criteria lists A/B testing as a quantitative way to measure performance. It defines A/B testing as comparing performance against a baseline model or an earlier version. So an A/B test always has two arms. The baseline is what you run today: the current prompt, model or pipeline. The variant is the change you think will help. Both arms are measured the same way, and the question is whether the variant is better by enough to matter.

    That means you decide on the measure and the target before you run the test. The documentation gives an example of a good criterion. A sentiment classifier should reach an F1 score of at least 0.85 on a held-out test set of 10,000 diverse Twitter posts, 'which is a 5% improvement over the current baseline'. The target is stated relative to the baseline. The guidance calls that part of the criterion 'Achievable' and says targets should rest on industry benchmarks, earlier experiments, AI research or expert knowledge. A variant that clears an absolute number but loses to the baseline has not won.

    Sources1

    2.Offline A/B: prompt variants on a shared test set

    The cheapest A/B test runs offline. Each variant goes through the same evaluation set, and no real users are involved. The success-criteria guidance says metrics should be measurable: quantitative metrics or well-defined qualitative scales. Offline comparison is where that pays off, because every variant gets a number on the same scale. The Building Evals cookbook describes this as knowing whether a change to your prompt improved a key metric.

    Anthropic's classification cookbook works through an example. It compares three prompt approaches for a ticket classifier: Simple, RAG, and RAG with chain-of-thought (CoT). The cookbook reports the results as a progression: 'Random baseline: ~10% → Simple: ~70% → RAG: 94% → RAG + CoT: 97%'. It says this data-driven approach showed which techniques added value and by how much. For production-scale testing it lists 'Multiple prompt variants: A/B testing different phrasings, structures, and approaches' and model comparisons across Claude models or temperature settings. It runs these in Promptfoo, an open-source evaluation toolkit. Each configuration ran on the same 68-example test set, to ensure fair comparison across all configurations, with each prediction checked automatically against ground-truth labels.

    Promptfoo accuracy (%) for three prompt variants across temperature settings, all run on the same 68-example test settext
    Haiku: T-0.0 Prompt: RAG w/ CoT    95.588235
    Haiku: T-0.2 Prompt: RAG w/ CoT    95.588235
    Haiku: T-0.8 Prompt: RAG w/ CoT    95.588235
    Haiku: T-0.0 Prompt: RAG           94.117647
    Haiku: T-0.4 Prompt: RAG           94.117647
    Haiku: T-0.6 Prompt: RAG           94.117647
    Haiku: T-0.4 Prompt: RAG w/ CoT    94.117647
    Haiku: T-0.6 Prompt: RAG w/ CoT    94.117647
    Haiku: T-0.2 Prompt: RAG           92.647059
    Haiku: T-0.8 Prompt: RAG           89.705882
    Haiku: T-0.6 Prompt: Simple        72.058824
    Haiku: T-0.0 Prompt: Simple        70.588235
    Haiku: T-0.2 Prompt: Simple        70.588235
    Haiku: T-0.8 Prompt: Simple        70.588235
    Haiku: T-0.4 Prompt: Simple        69.117647

    A support-ticket triage team maintains a categorical labeling eval (urgent/normal/low, with human-labeled ground truth) used to compare successive versions of their Claude-based classification prompt. Before promoting a revised prompt candidate, they want the fastest, most objective way to confirm it outperforms the current production prompt on this task. Which approach should they use?

    Sources123

    3.Online A/B: what live traffic tells you, and what it doesn't

    Offline results show how a variant scores on your test set. They don't show how users respond to it. For that, the success-criteria guidance points to user feedback, including 'Implicit measures like task completion rates'. Anthropic's engineering post on agent evals describes A/B testing as comparing variants with real user traffic. It lists this against the other ways to understand system performance, and each comes with a clear cost.

    Offline evals vs online A/B testing vs transcript review: what each gives you and what it costs
    MethodStrengthCost / limit
    Automated evalsFaster iteration, fully reproducible, no user impact, can run on every commitRequires more up-front investment to build and ongoing maintenance
    A/B testingMeasures actual user outcomes (retention, task completion); controls for confoundsSlow: days or weeks to reach significance, requires sufficient traffic, only tests changes you deploy
    Manual transcript reviewBuilds intuition for failure modes; catches subtle quality issues automated checks missTime-intensive, doesn't scale, typically only qualitative signal

    The post places each method at a stage. Automated evals run on each agent change and model upgrade as 'the first line of defense against quality problems'. 'A/B testing validates significant changes once you have sufficient traffic.' In practice, you screen variants offline first. Only the few that beat the baseline go to a live split, because each live test takes days or weeks. And when a metric moves in a live test, the test alone gives little signal on why. You find the reason by reading transcripts.

    A documentation team is iterating on a prompt that generates article summaries. They have 250 articles paired with human-written reference summaries and want an automated score that rewards a candidate summary for preserving the key information from the reference in a similar order, so they can compare prompt variants at scale during iteration. Which evaluation method best fits this need?

    Sources14

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Temperature is the main setting to tune when you A/B test a classification prompt.Why is that wrong?

      In Anthropic's Promptfoo comparison, the prompt approach (Simple vs RAG vs RAG w/ CoT) drove accuracy. The simple prompt stayed around 70% at every temperature, and CoT accuracy barely moved as temperature changed.

      Covered in Offline A/B: prompt variants on a shared test set

    2. 2.A live A/B test gives a quick answer, so run one for every small prompt tweak.Why is that wrong?

      Live A/B tests are slow and need enough traffic to reach significance. Use automated evals on every change and save live A/B tests for significant changes.

      Covered in Online A/B: what live traffic tells you, and what it doesn't

    3. 3.When a live A/B test shows a metric moved, the test also tells you why.Why is that wrong?

      A/B tests measure outcomes, not causes. Finding out why a metric moved means reading transcripts.

      Covered in Online A/B: what live traffic tells you, and what it doesn't

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “A/B testing: Compare performance against a baseline model or earlier version.”
      ↩︎ An A/B test needs a baseline and a target
      “which is a 5% improvement over the current baseline (Achievable).”
      ↩︎ An A/B test needs a baseline and a target
      “Base your targets on industry benchmarks, prior experiments, AI research, or expert knowledge.”
      ↩︎ An A/B test needs a baseline and a target
      “Measurable: Use quantitative metrics or well-defined qualitative scales.”
      ↩︎ Offline A/B: prompt variants on a shared test set
      “User feedback: Implicit measures like task completion rates.”
      ↩︎ Online A/B: what live traffic tells you, and what it doesn't
      “A/B testing: Compare performance against a baseline model or earlier version.”
      ↩︎ Key concept
    2. 2.
      “Whether you are trying to know if a change to your prompt made the model perform better on a key metric”
      ↩︎ Offline A/B: prompt variants on a shared test set
    3. 3.
      “Multiple prompt variants: A/B testing different phrasings, structures, and approaches”
      ↩︎ Offline A/B: prompt variants on a shared test set
      “This data-driven approach revealed exactly which techniques (RAG, chain-of-thought) added value and by how much.”
      ↩︎ Offline A/B: prompt variants on a shared test set
      “The same test set (68 examples): Ensuring fair comparison across all configurations”
      ↩︎ Offline A/B: prompt variants on a shared test set
      “Production recommendation: Use temperature=0.0 with RAG w/ CoT for maximum consistency and accuracy (95.59%)”
      ↩︎ Offline A/B: prompt variants on a shared test set
      “Temperature has minimal impact on CoT”
      ↩︎ Exam trap 1
    4. 4.
      “A/B testing validates significant changes once you have sufficient traffic.”
      ↩︎ Online A/B: what live traffic tells you, and what it doesn't
      “running on each agent change and model upgrade as the first line of defense against quality problems”
      ↩︎ Online A/B: what live traffic tells you, and what it doesn't
      “Slow; days or weeks to reach significance and requires sufficient traffic”
      ↩︎ Exam trap 2
      “without being able to thoroughly review the transcripts”
      ↩︎ Exam trap 3

    Continue to page 2 of 2

    Iterative Prompt Improvement: Observe, Refine, Retest