What you will be able to do
- Tell baseline latency apart from time to first token, and state a response-time target as a percentage threshold
- Measure cost from API usage data and compare configurations on cost per completed task instead of price per token
- Combine accuracy, safety, latency and cost targets into one multidimensional set of thresholds
1.Latency: know which clock you are reading
Accuracy and safety describe what the model says. Latency and cost describe what it takes to get that answer, and the testing guide lists both as success criteria. The latency criterion asks what response time is acceptable, and the answer depends on the application's real-time requirements and what users expect. Anthropic defines latency as the time the model takes to process a prompt and generate an output. It depends on the size of the model, the complexity of the prompt, and the infrastructure between the model and the user.
Latency isn't one number. The latency guide separates two measurements:
| Measurement | What it measures | Use it when |
|---|---|---|
| Baseline latency | Time for the model to process the prompt and generate the response, without considering input and output tokens per second | You want a general idea of the model's speed |
| Time to first token (TTFT) | Time from sending the prompt until the first token of the response is generated | You stream responses and users judge responsiveness by the first visible output |
That answers the prediction: a streamed answer can finish on time and still start late, and TTFT is the measurement that catches it. The testing guide also lists operational metrics, response time in milliseconds and uptime as a percentage, and its worked criterion states latency as a percentage threshold instead of an average: 95% of responses under 200ms. Use-case guides turn latency into a business measure too. In ticket routing, one success metric tracks how quickly tickets are assigned after they are submitted.
2.Cost: measure per task, not per token
The cost criterion asks what your budget for running the model is. The testing guide names three factors: the cost of each API call, the size of the model, and how often it's used. To measure cost you need token counts from real calls. Every API response includes a usage block with input, output and cache token counts. The ticket-routing guide's evaluation reads these usage statistics on every call to calculate cost from the input and output tokens used. It records that cost next to accuracy and response time.
| Model | $/M input | $/M output |
|---|---|---|
| claude-fable-5 | 10.00 | 50.00 |
| claude-opus-5 | 5.00 | 25.00 |
| claude-sonnet-5 | 2.00 | 10.00 |
| claude-haiku-4-5 | 1.00 | 5.00 |
A price list is the wrong unit for comparing configurations. Anthropic's model-selection guidance says to compare models on cost per completed task, not per token. The cookbook gives the reason: a model with a higher sticker price can end up cheaper if it finishes the job in fewer turns. Its claims-adjuster example tracks cost and pass rate together, so no cost saving can quietly lower quality:
baseline (opus · effort=high)
✓ CLM-001 APPROVE (exp APPROVE) 4 turns $0.2844
✓ CLM-002 DENY (exp DENY) 3 turns $0.2355
✓ CLM-003 APPROVE (exp APPROVE) 4 turns $0.3117
✓ CLM-004 FRAUD (exp FRAUD) 3 turns $0.2444
✓ CLM-005 DENY (exp DENY) 3 turns $0.2451
✓ CLM-006 APPROVE (exp APPROVE) 4 turns $0.3001
✓ CLM-007 SUPERVISOR (exp SUPERVISOR) 4 turns $0.3305
✓ CLM-008 SUPERVISOR (exp SUPERVISOR) 5 turns $0.4194
✓ CLM-009 FRAUD (exp FRAUD) 3 turns $0.2643
✓ CLM-010 FRAUD (exp FRAUD) 3 turns $0.2709
→ 10/10 correct · 36 turns · $0.2906/task · $2.9063 totalIn insurance claims, a wrong label costs more than the inference it saves. The cost metric only counts as a win when the quality bar still holds. That's why both numbers sit on the same scoreboard.
3.Putting the metrics together: multidimensional thresholds
No single metric covers an LLM application. The testing guide says most use cases need evaluation along several success criteria. Its worked example sets four targets on one held-out set of 10,000 posts: an F1 score of at least 0.85, 99.5% of outputs non-toxic, 90% of errors causing only inconvenience, and 95% of responses under 200ms. Accuracy, safety and latency are each stated as numbers that pass or fail.
The ticket-routing guide adds a key point: an evaluation that reports accuracy, response time and cost still needs thresholds to judge them against. Its examples are 95% accuracy over 100 tests and a 50% average reduction in cost per classification compared with the current routing method. With explicit thresholds, comparing two prompts or two models becomes a check rather than a debate.
Thresholds also show the trade-offs. The ticket-routing guide says model choice depends on the trade-offs between cost, accuracy and response time. For routing, it names Claude Haiku 4.5 as the fastest and most cost-effective model in the Claude 4 family. Once an evaluation suite exists, latency, token usage, cost per task and error rates can all be tracked on a static bank of tasks, which gives you baselines and regression checks. One caution applies to every threshold: model outputs are nondeterministic, so the same configuration can land at a different pass rate and cost from one run to the next. Base your decisions on multiple trials, not a single run.
A voice-based customer service agent needs to feel responsive in real time. The team defines a success criterion stating that the vast majority of user turns must receive a response within a strict time budget, while still tolerating occasional slower responses during traffic spikes. Which latency metric definition best matches this success criterion?
Correct answer: A — The 95th-percentile response time across production queries, with a target such as 95% of responses under 200ms
- A. Correct. A percentile-based target (e.g., p95) directly captures 'the vast majority of turns are fast' while explicitly tolerating a small tail of slower responses, matching the stated criterion.
- B. A monthly average can hide a large share of slow responses behind many fast ones, so it doesn't guarantee that most individual turns meet the time budget.
- C. Output token count relates to response length and indirectly to latency, but it is not itself a measurement of response time against a budget.
- D. Requests-per-minute measures throughput and capacity, not how quickly any individual user turn is answered.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The model with the lowest price per token is the cheapest one to run.Why is that wrong?
A more expensive model can finish a task in fewer turns, so configurations should be compared on cost per completed task.
Covered in Cost: measure per task, not per token
2.One headline accuracy number is enough to decide whether a system is ready.Why is that wrong?
Most applications need separate targets for accuracy, safety, latency and cost, each with its own threshold.
Covered in Putting the metrics together: multidimensional thresholds
3.Total response time fully captures how responsive a streaming application feels.Why is that wrong?
Time to first token is a separate measurement. It tracks how long the user waits before any output appears, which matters most when streaming.
Covered in Latency: know which clock you are reading
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“What is the acceptable response time for the model?”
↩︎ Latency: know which clock you are reading“Operational: Response time (ms), uptime (%)”
↩︎ Latency: know which clock you are reading“Consider factors like the cost for each API call, the size of the model, and the frequency of usage.”
↩︎ Cost: measure per task, not per token“99.5% of outputs are non-toxic”
↩︎ Putting the metrics together: multidimensional thresholds“Most use cases need multidimensional evaluation along several success criteria.”
↩︎ Exam trap 2 - 2.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-latencyOfficial docs
“Latency refers to the time it takes for the model to process a prompt and generate an output.”
↩︎ Latency: know which clock you are reading“the time taken by the model to process the prompt and generate the response, without considering the input and output tokens per second.”
↩︎ Latency: know which clock you are reading“measures the time it takes for the model to generate the first token of the response”
↩︎ Exam trap 3 - 3.
“This metric tracks how quickly tickets are assigned after being submitted.”
↩︎ Latency: know which clock you are reading“The method extracts usage statistics for the API call to calculate cost based on input and output tokens used.”
↩︎ Cost: measure per task, not per token“A proper evaluation requires clear thresholds and benchmarks to determine what is a good result.”
↩︎ Putting the metrics together: multidimensional thresholds“Cost per classification: 50% reduction on average (across 100 tests) from current routing method”
↩︎ Putting the metrics together: multidimensional thresholds“The choice of model depends on the trade-offs between cost, accuracy, and response time.”
↩︎ Putting the metrics together: multidimensional thresholds“it is the fastest and most cost-effective model in the Claude 4 family while still delivering excellent results”
↩︎ Putting the metrics together: multidimensional thresholds - 4.
“carries a usage block with input, output, and cache token counts”
↩︎ Cost: measure per task, not per token“a model with a higher sticker price can be the cheaper option if it finishes the job in fewer turns”
↩︎ Cost: measure per task, not per token“mis-labels are more costly than inference savings”
↩︎ Cost: measure per task, not per token“the same configuration can land at a different pass rate and cost per task from one run to the next”
↩︎ Putting the metrics together: multidimensional thresholds - 5.
“latency, token usage, cost per task, and error rates can be tracked on a static bank of tasks”
↩︎ Putting the metrics together: multidimensional thresholds
Also cited
- https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligenceOfficial docs
“Compare on cost per completed task, not per token”
↩︎ Exam trap 1