What you will be able to do
- Write a success criterion that says how much better a new prompt or model version has to be than the current baseline
- Run an offline A/B comparison of prompt variants on a shared test set and read what the results show
- Decide when a change needs an online A/B test with real user traffic, and know what that test cannot tell you
Key concept
Baseline comparison (A/B testing) — An A/B test measures a candidate prompt, model or configuration against the version you already have, on the same measure. A score on its own tells you little. The difference from the baseline is the result.
1.An A/B test needs a baseline and a target
Anthropic's guidance on success criteria lists A/B testing as a quantitative way to measure performance. It defines A/B testing as comparing performance against a baseline model or an earlier version. So an A/B test always has two arms. The baseline is what you run today: the current prompt, model or pipeline. The variant is the change you think will help. Both arms are measured the same way, and the question is whether the variant is better by enough to matter.
That means you decide on the measure and the target before you run the test. The documentation gives an example of a good criterion. A sentiment classifier should reach an F1 score of at least 0.85 on a held-out test set of 10,000 diverse Twitter posts, 'which is a 5% improvement over the current baseline'. The target is stated relative to the baseline. The guidance calls that part of the criterion 'Achievable' and says targets should rest on industry benchmarks, earlier experiments, AI research or expert knowledge. A variant that clears an absolute number but loses to the baseline has not won.
First, a measurable metric, such as accuracy, F1 or response time, or a qualitative scale applied consistently. Second, a target stated against the current version, such as 'a 5% improvement over the current baseline'. Without both, any difference you see can be argued either way.
Sources1
2.Offline A/B: prompt variants on a shared test set
The cheapest A/B test runs offline. Each variant goes through the same evaluation set, and no real users are involved. The success-criteria guidance says metrics should be measurable: quantitative metrics or well-defined qualitative scales. Offline comparison is where that pays off, because every variant gets a number on the same scale. The Building Evals cookbook describes this as knowing whether a change to your prompt improved a key metric.
Anthropic's classification cookbook works through an example. It compares three prompt approaches for a ticket classifier: Simple, RAG, and RAG with chain-of-thought (CoT). The cookbook reports the results as a progression: 'Random baseline: ~10% → Simple: ~70% → RAG: 94% → RAG + CoT: 97%'. It says this data-driven approach showed which techniques added value and by how much. For production-scale testing it lists 'Multiple prompt variants: A/B testing different phrasings, structures, and approaches' and model comparisons across Claude models or temperature settings. It runs these in Promptfoo, an open-source evaluation toolkit. Each configuration ran on the same 68-example test set, to ensure fair comparison across all configurations, with each prediction checked automatically against ground-truth labels.
Haiku: T-0.0 Prompt: RAG w/ CoT 95.588235
Haiku: T-0.2 Prompt: RAG w/ CoT 95.588235
Haiku: T-0.8 Prompt: RAG w/ CoT 95.588235
Haiku: T-0.0 Prompt: RAG 94.117647
Haiku: T-0.4 Prompt: RAG 94.117647
Haiku: T-0.6 Prompt: RAG 94.117647
Haiku: T-0.4 Prompt: RAG w/ CoT 94.117647
Haiku: T-0.6 Prompt: RAG w/ CoT 94.117647
Haiku: T-0.2 Prompt: RAG 92.647059
Haiku: T-0.8 Prompt: RAG 89.705882
Haiku: T-0.6 Prompt: Simple 72.058824
Haiku: T-0.0 Prompt: Simple 70.588235
Haiku: T-0.2 Prompt: Simple 70.588235
Haiku: T-0.8 Prompt: Simple 70.588235
Haiku: T-0.4 Prompt: Simple 69.117647Prompt approach. Simple stays around 70% at every temperature. RAG and RAG w/ CoT sit in the 89–96% range. The cookbook concludes that temperature has minimal impact on CoT and that simple prompts are temperature-agnostic. It recommends temperature=0.0 with RAG w/ CoT for maximum consistency and accuracy (95.59%).
A support-ticket triage team maintains a categorical labeling eval (urgent/normal/low, with human-labeled ground truth) used to compare successive versions of their Claude-based classification prompt. Before promoting a revised prompt candidate, they want the fastest, most objective way to confirm it outperforms the current production prompt on this task. Which approach should they use?
Correct answer: A — Run both prompt versions against the same held-out labeled set and compare exact-match accuracy against the ground-truth category labels
- A. Correct. Since ground-truth labels already exist for this categorical task, exact-match accuracy on a held-out set gives an objective, scalable, and directly comparable score between prompt versions before rollout.
- B. Incorrect. Manual agent voting is slow, inconsistent across raters, and does not scale the way automated grading against existing ground truth does.
- C. Incorrect. Self-reported confidence is not a validated accuracy measure and can be miscalibrated, so it does not substitute for grading against known correct labels.
- D. Incorrect. Shipping to all traffic without a controlled held-out comparison risks degrading production quality and conflates classification accuracy with an indirect proxy metric.
3.Online A/B: what live traffic tells you, and what it doesn't
Offline results show how a variant scores on your test set. They don't show how users respond to it. For that, the success-criteria guidance points to user feedback, including 'Implicit measures like task completion rates'. Anthropic's engineering post on agent evals describes A/B testing as comparing variants with real user traffic. It lists this against the other ways to understand system performance, and each comes with a clear cost.
| Method | Strength | Cost / limit |
|---|---|---|
| Automated evals | Faster iteration, fully reproducible, no user impact, can run on every commit | Requires more up-front investment to build and ongoing maintenance |
| A/B testing | Measures actual user outcomes (retention, task completion); controls for confounds | Slow: days or weeks to reach significance, requires sufficient traffic, only tests changes you deploy |
| Manual transcript review | Builds intuition for failure modes; catches subtle quality issues automated checks miss | Time-intensive, doesn't scale, typically only qualitative signal |
The post places each method at a stage. Automated evals run on each agent change and model upgrade as 'the first line of defense against quality problems'. 'A/B testing validates significant changes once you have sufficient traffic.' In practice, you screen variants offline first. Only the few that beat the baseline go to a live split, because each live test takes days or weeks. And when a metric moves in a live test, the test alone gives little signal on why. You find the reason by reading transcripts.
A documentation team is iterating on a prompt that generates article summaries. They have 250 articles paired with human-written reference summaries and want an automated score that rewards a candidate summary for preserving the key information from the reference in a similar order, so they can compare prompt variants at scale during iteration. Which evaluation method best fits this need?
Correct answer: A — Score each candidate summary against its reference using ROUGE-L and track the F1 score across prompt variants
- A. Correct. ROUGE-L measures the longest common subsequence between candidate and reference text, capturing whether key information appears in a similar order, and it is well suited to scoring summarization quality across many articles automatically.
- B. Incorrect. Exact string match is appropriate for short categorical answers, not free-text summaries, since two good summaries will rarely match a reference word-for-word.
- C. Incorrect. A binary PHI check evaluates a safety property, not whether the summary preserves the reference's key information and ordering.
- D. Incorrect. Latency measures response speed, not whether the generated content faithfully reflects the reference summary's content.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Temperature is the main setting to tune when you A/B test a classification prompt.Why is that wrong?
In Anthropic's Promptfoo comparison, the prompt approach (Simple vs RAG vs RAG w/ CoT) drove accuracy. The simple prompt stayed around 70% at every temperature, and CoT accuracy barely moved as temperature changed.
Covered in Offline A/B: prompt variants on a shared test set
2.A live A/B test gives a quick answer, so run one for every small prompt tweak.Why is that wrong?
Live A/B tests are slow and need enough traffic to reach significance. Use automated evals on every change and save live A/B tests for significant changes.
Covered in Online A/B: what live traffic tells you, and what it doesn't
3.When a live A/B test shows a metric moved, the test also tells you why.Why is that wrong?
A/B tests measure outcomes, not causes. Finding out why a metric moved means reading transcripts.
Covered in Online A/B: what live traffic tells you, and what it doesn't
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“A/B testing: Compare performance against a baseline model or earlier version.”
↩︎ An A/B test needs a baseline and a target“which is a 5% improvement over the current baseline (Achievable).”
↩︎ An A/B test needs a baseline and a target“Base your targets on industry benchmarks, prior experiments, AI research, or expert knowledge.”
↩︎ An A/B test needs a baseline and a target“Measurable: Use quantitative metrics or well-defined qualitative scales.”
↩︎ Offline A/B: prompt variants on a shared test set“User feedback: Implicit measures like task completion rates.”
↩︎ Online A/B: what live traffic tells you, and what it doesn't“A/B testing: Compare performance against a baseline model or earlier version.”
↩︎ Key concept - 2.https://platform.claude.com/cookbook/misc-building-evalsSecondary source
“Whether you are trying to know if a change to your prompt made the model perform better on a key metric”
↩︎ Offline A/B: prompt variants on a shared test set - 3.
“Multiple prompt variants: A/B testing different phrasings, structures, and approaches”
↩︎ Offline A/B: prompt variants on a shared test set“This data-driven approach revealed exactly which techniques (RAG, chain-of-thought) added value and by how much.”
↩︎ Offline A/B: prompt variants on a shared test set“The same test set (68 examples): Ensuring fair comparison across all configurations”
↩︎ Offline A/B: prompt variants on a shared test set“Production recommendation: Use temperature=0.0 with RAG w/ CoT for maximum consistency and accuracy (95.59%)”
↩︎ Offline A/B: prompt variants on a shared test set“Temperature has minimal impact on CoT”
↩︎ Exam trap 1 - 4.
“A/B testing validates significant changes once you have sufficient traffic.”
↩︎ Online A/B: what live traffic tells you, and what it doesn't“running on each agent change and model upgrade as the first line of defense against quality problems”
↩︎ Online A/B: what live traffic tells you, and what it doesn't“Slow; days or weeks to reach significance and requires sufficient traffic”
↩︎ Exam trap 2“without being able to thoroughly review the transcripts”
↩︎ Exam trap 3