CertSafari
    CCAR-P · Lessons

    Domain 4 · Lesson 20/38

    Latency and Cost Metrics: Response Time, Cost per Task and Thresholds

    Define evaluation metrics

    7 min read
    2.67% of exam
    6 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Tell baseline latency apart from time to first token, and state a response-time target as a percentage threshold
    • Measure cost from API usage data and compare configurations on cost per completed task instead of price per token
    • Combine accuracy, safety, latency and cost targets into one multidimensional set of thresholds

    1.Latency: know which clock you are reading

    Accuracy and safety describe what the model says. Latency and cost describe what it takes to get that answer, and the testing guide lists both as success criteria. The latency criterion asks what response time is acceptable, and the answer depends on the application's real-time requirements and what users expect. Anthropic defines latency as the time the model takes to process a prompt and generate an output. It depends on the size of the model, the complexity of the prompt, and the infrastructure between the model and the user.

    Latency isn't one number. The latency guide separates two measurements:

    Two latency measurements and what each tells you
    MeasurementWhat it measuresUse it when
    Baseline latencyTime for the model to process the prompt and generate the response, without considering input and output tokens per secondYou want a general idea of the model's speed
    Time to first token (TTFT)Time from sending the prompt until the first token of the response is generatedYou stream responses and users judge responsiveness by the first visible output

    That answers the prediction: a streamed answer can finish on time and still start late, and TTFT is the measurement that catches it. The testing guide also lists operational metrics, response time in milliseconds and uptime as a percentage, and its worked criterion states latency as a percentage threshold instead of an average: 95% of responses under 200ms. Use-case guides turn latency into a business measure too. In ticket routing, one success metric tracks how quickly tickets are assigned after they are submitted.

    Sources123

    2.Cost: measure per task, not per token

    The cost criterion asks what your budget for running the model is. The testing guide names three factors: the cost of each API call, the size of the model, and how often it's used. To measure cost you need token counts from real calls. Every API response includes a usage block with input, output and cache token counts. The ticket-routing guide's evaluation reads these usage statistics on every call to calculate cost from the input and output tokens used. It records that cost next to accuracy and response time.

    List prices per million tokens, as snapshotted in Anthropic's cost-optimization cookbook (it warns these rates change)
    Model$/M input$/M output
    claude-fable-510.0050.00
    claude-opus-55.0025.00
    claude-sonnet-52.0010.00
    claude-haiku-4-51.005.00

    A price list is the wrong unit for comparing configurations. Anthropic's model-selection guidance says to compare models on cost per completed task, not per token. The cookbook gives the reason: a model with a higher sticker price can end up cheaper if it finishes the job in fewer turns. Its claims-adjuster example tracks cost and pass rate together, so no cost saving can quietly lower quality:

    A baseline eval that records correctness, turns and dollars for each tasktext
    baseline (opus · effort=high)
      ✓ CLM-001  APPROVE    (exp APPROVE)    4 turns  $0.2844
      ✓ CLM-002  DENY       (exp DENY)       3 turns  $0.2355
      ✓ CLM-003  APPROVE    (exp APPROVE)    4 turns  $0.3117
      ✓ CLM-004  FRAUD      (exp FRAUD)      3 turns  $0.2444
      ✓ CLM-005  DENY       (exp DENY)       3 turns  $0.2451
      ✓ CLM-006  APPROVE    (exp APPROVE)    4 turns  $0.3001
      ✓ CLM-007  SUPERVISOR (exp SUPERVISOR) 4 turns  $0.3305
      ✓ CLM-008  SUPERVISOR (exp SUPERVISOR) 5 turns  $0.4194
      ✓ CLM-009  FRAUD      (exp FRAUD)      3 turns  $0.2643
      ✓ CLM-010  FRAUD      (exp FRAUD)      3 turns  $0.2709
      → 10/10 correct · 36 turns · $0.2906/task · $2.9063 total

    Sources134

    3.Putting the metrics together: multidimensional thresholds

    No single metric covers an LLM application. The testing guide says most use cases need evaluation along several success criteria. Its worked example sets four targets on one held-out set of 10,000 posts: an F1 score of at least 0.85, 99.5% of outputs non-toxic, 90% of errors causing only inconvenience, and 95% of responses under 200ms. Accuracy, safety and latency are each stated as numbers that pass or fail.

    The ticket-routing guide adds a key point: an evaluation that reports accuracy, response time and cost still needs thresholds to judge them against. Its examples are 95% accuracy over 100 tests and a 50% average reduction in cost per classification compared with the current routing method. With explicit thresholds, comparing two prompts or two models becomes a check rather than a debate.

    Thresholds also show the trade-offs. The ticket-routing guide says model choice depends on the trade-offs between cost, accuracy and response time. For routing, it names Claude Haiku 4.5 as the fastest and most cost-effective model in the Claude 4 family. Once an evaluation suite exists, latency, token usage, cost per task and error rates can all be tracked on a static bank of tasks, which gives you baselines and regression checks. One caution applies to every threshold: model outputs are nondeterministic, so the same configuration can land at a different pass rate and cost from one run to the next. Base your decisions on multiple trials, not a single run.

    A voice-based customer service agent needs to feel responsive in real time. The team defines a success criterion stating that the vast majority of user turns must receive a response within a strict time budget, while still tolerating occasional slower responses during traffic spikes. Which latency metric definition best matches this success criterion?

    Sources1354

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.The model with the lowest price per token is the cheapest one to run.Why is that wrong?

      A more expensive model can finish a task in fewer turns, so configurations should be compared on cost per completed task.

      Covered in Cost: measure per task, not per token

    2. 2.One headline accuracy number is enough to decide whether a system is ready.Why is that wrong?

      Most applications need separate targets for accuracy, safety, latency and cost, each with its own threshold.

      Covered in Putting the metrics together: multidimensional thresholds

    3. 3.Total response time fully captures how responsive a streaming application feels.Why is that wrong?

      Time to first token is a separate measurement. It tracks how long the user waits before any output appears, which matters most when streaming.

      Covered in Latency: know which clock you are reading

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “What is the acceptable response time for the model?”
      ↩︎ Latency: know which clock you are reading
      “Operational: Response time (ms), uptime (%)”
      ↩︎ Latency: know which clock you are reading
      “Consider factors like the cost for each API call, the size of the model, and the frequency of usage.”
      ↩︎ Cost: measure per task, not per token
      “Most use cases need multidimensional evaluation along several success criteria.”
      ↩︎ Exam trap 2
    2. 2.
      “Latency refers to the time it takes for the model to process a prompt and generate an output.”
      ↩︎ Latency: know which clock you are reading
      “the time taken by the model to process the prompt and generate the response, without considering the input and output tokens per second.”
      ↩︎ Latency: know which clock you are reading
      “measures the time it takes for the model to generate the first token of the response”
      ↩︎ Exam trap 3
    3. 3.
      “This metric tracks how quickly tickets are assigned after being submitted.”
      ↩︎ Latency: know which clock you are reading
      “The method extracts usage statistics for the API call to calculate cost based on input and output tokens used.”
      ↩︎ Cost: measure per task, not per token
      “A proper evaluation requires clear thresholds and benchmarks to determine what is a good result.”
      ↩︎ Putting the metrics together: multidimensional thresholds
      “Cost per classification: 50% reduction on average (across 100 tests) from current routing method”
      ↩︎ Putting the metrics together: multidimensional thresholds
      “The choice of model depends on the trade-offs between cost, accuracy, and response time.”
      ↩︎ Putting the metrics together: multidimensional thresholds
      “it is the fastest and most cost-effective model in the Claude 4 family while still delivering excellent results”
      ↩︎ Putting the metrics together: multidimensional thresholds
    4. 4.
      “carries a usage block with input, output, and cache token counts”
      ↩︎ Cost: measure per task, not per token
      “a model with a higher sticker price can be the cheaper option if it finishes the job in fewer turns”
      ↩︎ Cost: measure per task, not per token
      “mis-labels are more costly than inference savings”
      ↩︎ Cost: measure per task, not per token
      “the same configuration can land at a different pass rate and cost per task from one run to the next”
      ↩︎ Putting the metrics together: multidimensional thresholds
    5. 5.
      “latency, token usage, cost per task, and error rates can be tracked on a static bank of tasks”
      ↩︎ Putting the metrics together: multidimensional thresholds

    Also cited

    Ready to test yourself?

    Practise the 12 questions on this subdomain.