CertSafari
    CCAR-P · Lessons

    Domain 3 · Lesson 14/38

    Accuracy-Latency Trade-offs: Model Choice and Effort

    Evaluate accuracy-latency trade-offs and justify configuration decisions

    9 min read
    2.38% of exam
    6 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Name the latency metric (baseline latency or time to first token) a configuration change is meant to improve.
    • Tell apart levers that trade intelligence for speed from levers that only remove waste.
    • Justify a speed-first or capability-first starting model for a given workload.
    • Choose an effort level and defend it, allowing for model-specific defaults and behaviour.

    Key concept

    Effort as the intelligence-for-latency dial — Most latency decisions buy speed by spending less reasoning. The effort parameter makes that trade explicit inside one model, so model tier and effort are the two dials you have to justify with evals.

    1.Say which latency you are optimising

    Before you can justify a configuration, you have to say which latency you mean. Anthropic's latency guidance defines latency as the time the model takes to process a prompt and generate an output. Three things shape it: the size of the model, how complex the prompt is, and the infrastructure between the model and the user.

    The guidance names two measurements. Baseline latency is the overall time the model takes to process the prompt and produce a response. It gives you a rough sense of the model's speed. Time to first token (TTFT) is the time from sending the prompt to receiving the first token of the response. TTFT is what users notice when output is streamed, because it decides how long they look at a blank screen.

    The distinction matters because each lever moves a different metric. Some levers shorten total generation time. Others only change when the first visible text appears. A justification that says "this makes it faster" without naming the metric is incomplete, because a lever can improve one metric and leave the other untouched.

    Sources1

    2.Two kinds of lever: trades and free wins

    Anthropic's cost-and-intelligence guide pictures configuration as a frontier. Some levers exchange capability for savings. Others remove waste and leave quality alone. The guide frames this in terms of cost, but the same split works for latency. Its list of real trade-offs is model choice, effort, output caps and task budgets, an elapsed-time clock, and multi-model architectures. Each of these makes a response cheaper or faster by making it less thorough. Its free wins, such as prompt caching and token hygiene, cut spend without touching quality.

    That gives you a working rule. When you pull a trade-off lever, you need a measurement showing that quality held. When you pull a free lever, you don't. A scenario that asks you to cut latency "without affecting accuracy" is pointing you at the free group. A scenario that accepts some loss of capability is inviting the trade-offs, and the right answer usually names how you would check the loss.

    Sources2

    3.Model choice: speed-first or capability-first

    The latency guidance calls model choice one of the most direct levers. For speed-critical applications it recommends Claude Haiku 4.5 as the fastest option that still keeps high intelligence.

    Routing a time-sensitive summarisation task to Claude Haiku 4.5python
    client = anthropic.Anthropic()
    
    # For time-sensitive applications, use Claude Haiku 4.5
    message = client.messages.create(
        model="claude-haiku-4-5",
        max_tokens=100,
        messages=[
            {
                "role": "user",
                "content": "Summarize this customer feedback in 2 sentences: [feedback text]",
            }
        ],
    )
    print(message.content[0].text)

    The model-selection guide gives two ways to start. Speed-first means you begin with Claude Haiku 4.5, test thoroughly, and upgrade only when you find a specific capability gap. It suits prototyping, tight latency requirements, cost-sensitive deployments and high-volume straightforward tasks. Capability-first means you implement with Claude Opus 5.5, optimise your prompts for it, and then gain efficiency over time by lowering effort or moving to a smaller model. If evals at xhigh or max effort still fall short on demanding reasoning or long agentic work, you step up to Claude Fable 5.1. Capability-first suits complex reasoning, nuanced understanding and applications where accuracy matters more than cost. The guide also notes that most workloads start with Claude Opus 5.5.

    Starting points by requirement, from the model-selection guide
    When you need...Start withExample use cases
    The highest available capabilityClaude Fable 5.1Agent sessions that run for hours, multistep deep research
    Complex agentic coding and enterprise workClaude Opus 5.5Multihour autonomous coding agents, large-scale refactoring
    Speed and capability for everyday workloadsClaude Sonnet 5Code generation, data analysis, agentic tool use
    The lowest latency and price, with extended thinkingClaude Haiku 4.5Real-time applications, high-volume processing, sub-agent tasks

    A team is building a real-time customer support chat feature that must respond within one to two seconds, handles a very high volume of simple intent-classification queries, and has a tight cost budget. Which model should they start with to satisfy the latency and cost requirements while keeping accuracy at an acceptable level for this task?

    Sources13

    4.Effort: the fine-grained dial inside one model

    Model choice is the coarse dial and effort is the fine one. Effort sets how many tokens Claude spends on a response. That covers the text, tool calls and their arguments, and thinking when it is active. Lower effort also produces fewer and terser tool calls, which matters in agent loops where every call is a round trip. The model-selection guide says outright that tuning effort is often a better lever than switching models.

    Setting effort on a single request with output_configpython
    client = anthropic.Anthropic()
    
    response = client.messages.create(
        model="claude-opus-5-5",
        max_tokens=4096,
        messages=[
            {
                "role": "user",
                "content": "Analyze the trade-offs between microservices and monolithic architectures",
            }
        ],
        output_config={"effort": "medium"},
    )
    Effort levels and their intended use
    LevelDescriptionTypical use case
    maxAbsolute maximum capability with no constraints on token spendingTasks requiring the deepest possible reasoning
    xhighExtended capability for long-horizon workLong-running agentic and coding tasks (over 30 minutes)
    highSpends as many tokens as the task needs; the default except on Claude Opus 5.5Complex reasoning, difficult coding problems, agentic tasks
    mediumBalanced approach with moderate token savings; the default on Claude Opus 5.5Agentic tasks balancing speed, cost, and performance
    lowMost efficient; significant token savings with some capability reductionSimpler tasks needing the best speed and lowest costs, such as subagents

    Effort also steers thinking. Thinking is adaptive: Claude decides for each request whether to reason and how deeply, and effort is the main control over that decision. Thinking improves results on maths, coding and long agentic work, but it has a cost. Reasoning tokens are billed as output tokens and count toward max_tokens, so high effort adds real latency. If you want Claude to think less often, lower effort before you try steering it with the prompt. On Claude Opus 5.5, adaptive thinking is always on and cannot be disabled, which makes effort the main lever there.

    Three cautions come up when you defend an effort setting. First, defaults differ between models. Most models default to high, but Claude Opus 5.5 defaults to medium, so a request that omits effort runs one level lower than it did on Claude Opus 5. Second, don't carry settings over from an earlier model. Run a fresh effort sweep on your own evals. Third, effort is not a length control. On Claude Opus 5 it changes how much the model thinks, not how long the visible response is, so ask for brevity in the prompt instead. At the top of the range, the Claude Opus 4.7 guidance warns that max often costs far more for small gains.

    An engineering team runs many parallel subagents performing simple lookups inside a larger agentic workflow. They want to minimize latency and token cost per subagent call without materially degrading the quality needed for those simple lookups. Which effort configuration decision best matches this need?

    Sources456

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.If a workload is too slow, the first fix is to move to a smaller model.Why is that wrong?

      Anthropic says adjusting effort on your current model is often a better lever than changing models. Try an effort sweep before you drop a tier.

      Covered in Model choice: speed-first or capability-first

    2. 2.A request that leaves out effort runs at high on every model.Why is that wrong?

      Claude Opus 5.5 defaults to medium, so leaving effort out runs one level lower than it did on Claude Opus 5.

      Covered in Effort: the fine-grained dial inside one model

    3. 3.Lowering effort is a reliable way to get shorter visible answers.Why is that wrong?

      On Claude Opus 5, effort changes how much the model thinks, not how long the visible response is. Ask for brevity in the prompt.

      Covered in Effort: the fine-grained dial inside one model

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Time to first token (TTFT): This metric measures the time it takes for the model to generate the first token of the response”
      ↩︎ Say which latency you are optimising
      “Latency can be influenced by various factors, such as the size of the model, the complexity of the prompt”
      ↩︎ Say which latency you are optimising
      “One of the most direct ways to reduce latency is to select the appropriate model for your use case.”
      ↩︎ Model choice: speed-first or capability-first
    2. 2.
      “Managing cost well means understanding how each cost lever affects output quality, because some levers trade against quality and some don't.”
      ↩︎ Two kinds of lever: trades and free wins
      “Tradeoffs exchange cost for intelligence: model choice, effort, output caps and task budgets, an elapsed-time clock, and multi-model architectures.”
      ↩︎ Two kinds of lever: trades and free wins
    3. 3.
      “Upgrade only if necessary for specific capability gaps.”
      ↩︎ Model choice: speed-first or capability-first
      “Applications where accuracy outweighs cost considerations”
      ↩︎ Model choice: speed-first or capability-first
      “Several Claude models support an effort parameter that trades intelligence for latency and cost within a single model.”
      ↩︎ Key concept
      “Tuning effort is often a better lever than switching models.”
      ↩︎ Exam trap 1
    4. 4.
      “Lower effort also means fewer and terser tool calls.”
      ↩︎ Effort: the fine-grained dial inside one model
      “Run an effort sweep on your own evals rather than carrying settings over from an earlier model”
      ↩︎ Effort: the fine-grained dial inside one model
      “On most workloads max adds significant cost for relatively small quality gains”
      ↩︎ Effort: the fine-grained dial inside one model
      “Most Claude models default to high effort, spending as many tokens as needed for excellent results; Claude Opus 5.5 defaults to medium.”
      ↩︎ Exam trap 2
      “Effort controls thinking volume, not visible response length”
      ↩︎ Exam trap 3
    5. 5.
      “If you want Claude to think less often, lower the effort level before reaching for prompt-based steering.”
      ↩︎ Effort: the fine-grained dial inside one model
    6. 6.
      “the tokens Claude spends reasoning are billed as output tokens, even when the thinking text isn't returned to you”
      ↩︎ Effort: the fine-grained dial inside one model

    Continue to page 2 of 2

    Cutting Latency Without Losing Accuracy: Tokens, Streaming, Fast Mode