CertSafari
    CCAR-P · Lessons

    Domain 3 · Lesson 14/38

    Cutting Latency Without Losing Accuracy: Tokens, Streaming, Fast Mode

    Evaluate accuracy-latency trade-offs and justify configuration decisions

    7 min read
    2.38% of exam
    5 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Trim input and output tokens without letting max_tokens cut off reasoning.
    • Improve perceived responsiveness with streaming and omitted thinking display, and say what each one does and does not change.
    • Decide when fast mode is worth its premium, and list its constraints.
    • Back a configuration decision with benchmark evidence and measured trade-off patterns.

    1.Token hygiene, and where max_tokens stops being free

    The cheapest speed-up removes tokens nobody needed. Anthropic's latency guidance is simple: the fewer tokens the model has to process and generate, the faster it responds. It gives four tactics. Be clear but concise, and don't cut so far that instructions become ambiguous, because Claude has no context on your use case. Ask Claude directly for shorter responses. Set max_tokens as a hard cap on output length. Experiment with temperature, since lower values such as 0.2 can sometimes give more focused, shorter output. The guidance adds that finding the balance between clarity, quality and token count may take experimentation.

    The max_tokens cap is where a free lever quietly becomes a trade. When thinking is active, thinking tokens count toward max_tokens. A cap sized for the visible answer can therefore cut the reasoning short, and the answer gets worse. The cost guide's rule: if attempts end with stop_reason: max_tokens, raise the cap. In its measurements, 64,000 covered all but 2 of 14,000 turns at the default effort. The effort documentation likewise tells you to set a large max_tokens at the higher effort levels.

    Sources12

    2.Perceived latency: streaming and thinking display

    Some latency problems are really about how long the user waits to see something. Streaming lets the model start sending its response before the whole output is complete. You can then update the interface or do other work in parallel as tokens arrive. The total work stays the same, but the wait before the user sees output gets much shorter.

    Thinking adds a wrinkle, because on thinking models the reasoning comes before the answer. The display field on the thinking configuration decides what you get back. "summarized" returns a readable summary of the reasoning. "omitted" returns thinking blocks with an empty thinking field and the encrypted signature. "updates" (beta) returns short progress updates between tool calls. If your app never shows thinking to users, set display to "omitted". When streaming, the server then skips the thinking tokens and sends only the signature, so the answer text starts arriving sooner.

    Omitted display costs no accuracy, because the reasoning still happens and the block is passed back the same way in multi-turn conversations. It is already the default on Claude Opus 5.5, Claude Fable 5.1, Claude Sonnet 5 and several other current models. The setting you need to justify is therefore "summarized", since that is the one that adds time before the first answer text.

    Sources13

    3.Fast mode: paying for speed instead of trading accuracy

    When users need more output throughput from an Opus-class model, fast mode gives up to 2.5x higher output tokens per second on Claude Opus 5.5, Claude Opus 5 and Claude Opus 4.8. It runs the same model weights with a faster inference configuration, so you trade money for speed, not accuracy. On Claude Opus 5.5 it costs $8 per million input tokens and $40 per million output tokens. You opt in with speed="fast" and a beta header.

    Opting a request into fast modepython
    client = anthropic.Anthropic()
    
    response = client.beta.messages.create(
        model="claude-opus-5-5",
        max_tokens=4096,
        speed="fast",
        betas=["fast-mode-2026-02-01"],
        messages=[
            {"role": "user", "content": "Refactor this module to use dependency injection"}
        ],
    )

    The documentation's fallback example handles a RateLimitError on a fast request by deleting the speed parameter and retrying at standard speed. That keeps requests flowing, but check the constraints below before you rely on it. A retry at the other speed does not share the cached prefix.

    Fast mode constraints that affect a configuration decision
    AreaBehaviour
    Prompt cachingSwitching between fast and standard speed invalidates the prompt cache
    TTFTBenefits are focused on output tokens per second (OTPS), not time to first token
    Batch APIFast mode is not available with the Batch API
    Priority TierFast mode is not available with a Priority Tier commitment
    Claude Platform on AWSFast mode is not currently available

    Sources4

    4.Justifying the configuration with measurements

    A configuration decision is justified when measurements on your own workload back it. The model-selection guide lists the steps: build benchmark tests specific to your use case, test with your real prompts and data, and compare accuracy, response quality and edge-case handling across models before weighing performance against cost. The cost guide adds that you should compare models on cost per completed task, not per token. It also stresses that its own numbers are directional, so you have to measure your own workload.

    Measured patterns from Anthropic's cost-and-intelligence guide that move speed, cost or quality
    SituationDo thisMeasured effect or condition
    You want agent runs to finish soonerTell the model that time matters, and show it the elapsed time33% to 69% less time, 28% to 54% lower cost per task, scores up to 1.9 points lower
    Costs are too high; quality is fineSweep effort down on your current modelTrade-off lever: confirm quality on evals
    Quality isn't good enoughRestore effort if you lowered it; otherwise try the next tier up at low effortMoves back up the frontier
    You can check outputs (tests, a verifier)Run everything at low effort and re-run failures at highPass rate held at about half the cost on the coding benchmark
    A lower-cost model stalls only on hard decisionsAdd a frontier advisorPays off when priced well above the executor and actually consulted

    The elapsed-time row is the clearest accuracy-latency trade in the guidance. Runs finish much sooner, and the accuracy cost is stated: scores up to 1.9 points lower. Justifying it means arguing that the drop is acceptable for your use case. The re-run-failures row shows the opposite approach. When a verifier can spot wrong answers, low effort becomes the fast default and only failures pay for high effort, so the pass rate holds. Multi-model setups use the same idea at architecture level: an executor passes hard decisions to an advisor, or an orchestrator hands bulk work to cheaper workers.

    A financial analytics application uses extended thinking to solve complex multi-step valuation problems. Response latency has become unacceptable for end users, but outputs must remain highly accurate given the complexity of the problems. Which adjustment best balances these constraints?

    Sources52

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Fast mode is the fix when users complain about the delay before the first token.Why is that wrong?

      Fast mode raises output tokens per second. For TTFT, look at streaming and omitted thinking display instead.

      Covered in Fast mode: paying for speed instead of trading accuracy

    2. 2.Setting thinking display to "omitted" lowers the bill because less thinking text is returned.Why is that wrong?

      The full thinking tokens are still billed. Omitting speeds up the first answer text but does not reduce cost.

      Covered in Perceived latency: streaming and thinking display

    3. 3.A tight max_tokens is a safe latency cap that only limits answer length.Why is that wrong?

      Thinking tokens count toward max_tokens, so a tight cap can cut the reasoning short as well as the answer.

      Covered in Token hygiene, and where max_tokens stops being free

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “The fewer tokens the model has to process and generate, the faster the response will be.”
      ↩︎ Token hygiene, and where max_tokens stops being free
      “Use the max_tokens parameter to set a hard limit on the maximum length of the generated response.”
      ↩︎ Token hygiene, and where max_tokens stops being free
      “This can significantly improve the perceived responsiveness of your application, as users can see the model's output in real time.”
      ↩︎ Perceived latency: streaming and thinking display
    2. 2.
      “Raise max_tokens; 64,000 covered all but 2 of 14,000 turns measured at the default effort”
      ↩︎ Token hygiene, and where max_tokens stops being free
      “Compare on cost per completed task, not per token”
      ↩︎ Justifying the configuration with measurements
      “runs took 33% to 69% less time at a 28% to 54% lower cost per task, with scores up to 1.9 points lower”
      ↩︎ Justifying the configuration with measurements
    3. 3.
      “The primary benefit is faster time-to-first-text-token when streaming”
      ↩︎ Perceived latency: streaming and thinking display
      “You're still charged for the full thinking tokens. Omitting reduces latency, not cost.”
      ↩︎ Exam trap 2
      “Thinking tokens count toward max_tokens, so set it high enough to leave room for both the thinking and the response text.”
      ↩︎ Exam trap 3
    4. 4.
      “Fast mode runs the same model with a faster inference configuration. There is no change to intelligence or capabilities.”
      ↩︎ Fast mode: paying for speed instead of trading accuracy
      “Switching between fast and standard speed invalidates the prompt cache.”
      ↩︎ Fast mode: paying for speed instead of trading accuracy
      “Speed benefits are focused on output tokens per second (OTPS), not time to first token (TTFT)”
      ↩︎ Exam trap 1
    5. 5.
      “Create benchmark tests specific to your use case - having a good evaluation set is the most important step in the process.”
      ↩︎ Justifying the configuration with measurements