CertSafari
    CCAR-P · Lessons

    Domain 4 · Lesson 24/38

    Optimizing token usage, latency, and cost-performance trade-offs

    Optimize token usage, latency, and cost-performance trade-offs

    18 min read
    2.67% of exam
    6 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Separate cost levers that are free from levers that trade quality away
    • Turn on prompt caching first and judge a loop by its cache-read share
    • Choose between the 5-minute and 1-hour cache duration from the pauses in real traffic
    • Cut tokens and latency with prompt hygiene, max_tokens, model choice, and streaming
    • Pick a model, an effort level, and a multi-model architecture on cost per completed task

    Key concept

    Cost-to-intelligence frontier — The curve on which paying more buys more capability. Some levers move a workload toward the frontier by cutting spend at the same quality; only the rest move it along the frontier, buying cost with intelligence or the reverse.

    1.Two kinds of levers: free wins and trade-offs

    Once a workload leaves the prototype stage, cost becomes a design constraint alongside quality. The useful way to organise the optimisation work is not by technique but by what each technique costs you in output quality. One group of levers is free: prompt caching, token hygiene, auditing a prompt against the model you are actually running, batch processing, and workspace spend limits as a backstop. The other group buys cost with intelligence: model choice, effort, output caps and task budgets, an elapsed-time clock, and multi-model architectures.

    Work the free group to exhaustion before touching the second. A team that opens the optimisation effort by downgrading its model has skipped the levers that would have paid more and cost nothing. In Anthropic's own measurements the single largest lever was caching, not model choice.

    Sources1

    2.Prompt caching is the first lever, and usually the biggest

    An agent loop resends its whole conversation on every turn: system prompt, tool definitions, and every prior turn. That is why a 40-turn task sends its first turn 40 times and task cost grows with roughly the square of the turn count. Caching does not stop the resending. It makes each resend cost about a tenth as much, because the prefix is billed at the cache-read rate while only the new tokens pay the 1.25x cache-write rate.

    The measured effect is large: caching cut agent-loop cost by a factor of 2.7 to 5.3 on the guide's benchmarks, and cut a small triage agent's bill by 83%, or 88% once input trimming was added. Across measured runs, cache reads are routinely the largest single component of task cost.

    Cache pricing per million tokens: writes cost a premium, reads cost a fraction of input
    ModelInput5m writes1h writesHits and refreshes
    Claude Fable 5.1$10 / MTok$12.50 / MTok$20 / MTok$0.25 / MTok
    Claude Opus 5.5$4 / MTok$5 / MTok$8 / MTok$0.20 / MTok
    Claude Sonnet 5$2 / MTok$2.50 / MTok$4 / MTok$0.20 / MTok
    Claude Haiku 4.5$1 / MTok$1.25 / MTok$2 / MTok$0.10 / MTok

    Most models read at the standard 0.1x multiplier, but not all: hits on Claude Fable 5.1 and Claude Mythos 5.1 are 0.025x of base input, and hits on Claude Opus 5.5 are 0.05x. Those multipliers change which caching strategy is cheapest, which is the subject of the next section.

    Caching also gives you a health metric. Over a full day of real traffic, agent loops read a median 84% of their input from the cache, and the top 10% of harnesses read 94% or more. Deep in a task a well-built loop pays full price on under 1% of its input, so a cache-read share below roughly 80% is a signal that something is breaking the cache rather than a fact of life.

    A support-bot application sends the same 30,000-token product manual as part of the system prompt on every request, followed by a short unique user question. Requests arrive roughly every 30-90 seconds throughout the day. Which change most directly reduces both cost and latency for this workload without changing the model or the manual content?

    Sources12

    3.Picking the cache duration from real pause data

    The cache lifetime defaults to 5 minutes, and a read refreshes the entry at no extra cost. Critically, the clock starts when the request begins, not when its response ends: a response that streams for 4 minutes leaves only about a minute for the follow-up request to start and still hit.

    A 1-hour duration is available at a higher write price — 2x base input instead of 1.25x. Because a miss on either duration bills the whole prefix at the write price, the decision is purely about the distribution of gaps between consecutive requests in a conversation. Count those gaps rather than guessing.

    Requesting the 1-hour TTL on the automatic cache breakpointjson
    { "cache_control": { "type": "ephemeral", "ttl": "1h" } }

    Turns arriving seconds apart: stay on the default — with nothing paused it cost 15% less than the 1-hour setting on Claude Sonnet 5, and about 15% to 18% less on Claude Opus 5.5. Pauses between 5 minutes and an hour on more than about 1 turn in 20, with long gaps rare: buy the 1-hour duration. Gaps over an hour common: stay on the default, since such a gap expires both durations and the 1-hour setting then re-writes at its higher price. The published threshold is that the 1-hour duration pays only when at least about 40% of long pauses end within the hour.

    Two model-specific wrinkles. On Claude Fable 5.1, whose cache reads cost 0.025x of input while its writes keep the standard multipliers, keep-alive requests that re-read the prefix are cheap: keeping the 5-minute cache warm cost 13% to 20% less per session than the 1-hour cache whenever pauses ran for minutes, and only near 45-minute pauses did the 1-hour cache win. On Claude Sonnet 5, by contrast, keep-alives saved about 8% at 1 paused turn in 20 and nothing by 2 in 20, so the 1-hour duration is the recommendation there.

    Sources21

    4.Token hygiene, measured before you send

    Trimming tokens is the other free win, and it compounds with caching: the triage agent went from 83% to 88% cheaper once input trimming was added on top of the cache. Trimming needs a measurement, though, and guessing at token counts from character length is how prompt budgets drift. The token counting endpoint takes the same structured input as a Messages request — system prompts, tools, images, PDFs — and returns the input token total before you spend anything.

    Counting input tokens for a request before sending itpython
    response = client.messages.count_tokens(
        model="claude-opus-5-5",
        system="You are a scientist",
        messages=[{"role": "user", "content": "Hello, Claude"}],
    )

    Counting serves three jobs: managing rate limits and costs ahead of time, routing requests to a model, and optimising a prompt to a specific length. It is free, subject to a requests-per-minute limit by usage tier, and supported on all active models. Note the gap in coverage: server tools such as web search and code execution, the MCP connector, and image or document blocks given as a URL return an invalid_request_error, so send images and PDFs as base64 to count them.

    Counts are per-tokenizer, not universal. Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, and Claude Mythos 5 share the tokenizer introduced with Claude Opus 4.7, and a prompt counts roughly 30 percent higher there than on earlier models. The endpoint counts under the tokenizer of the model you pass, so a migration budget comes from counting the same request twice, once per model, and comparing input_tokens.

    Sources3

    5.Latency: what to measure and what moves it

    Latency is the time to process a prompt and generate an output, and it has two distinct measurements. Baseline latency is the model's overall speed at processing a prompt and producing a response, ignoring tokens per second. Time to first token is the delay before the first token appears, which is what a user actually perceives in a streaming interface. Optimise for the one your experience is judged on.

    Latency levers and what each one changes
    LeverWhat it does
    Choose claude-haiku-4-5Fastest response times while maintaining high intelligence, for speed-critical applications
    Be clear but conciseFewer input tokens to process, without dropping the context Claude needs
    Ask for shorter responsesCurbs unwanted length at the source, so fewer output tokens are generated
    Set max_tokensA hard limit on generated response length, preventing overly long outputs
    Lower temperature (for example, 0.2)Can lead to more focused and shorter responses
    Enable streamingImproves perceived responsiveness; output arrives before generation completes

    The token levers pull in the same direction as the cost levers — the fewer tokens the model processes and generates, the faster the response — but they are not free in the way caching is. Cutting a prompt too hard removes context Claude lacks about your use case, so it may not make the intended leaps of logic. Finding the balance among prompt clarity, output quality, and token count takes experimentation.

    Streaming is the one lever that changes nothing about the work done. It only changes when the user sees it: the response starts arriving before the full output is complete, so your interface can update in real time while generation continues. It improves perceived responsiveness without reducing baseline latency at all.

    Latency can also be bought directly. Claude Opus 5.5, Claude Opus 5, and Claude Opus 4.8 support fast mode, a research preview delivering up to 2.5x higher output speed — at premium pricing. It is a latency lever that moves cost the wrong way, so it belongs in the trade-off group, not the free-win group.

    A team runs a nightly job that scores 50,000 archived support tickets against a rubric. The job has no interactive user waiting on results and currently completes in about six hours using standard synchronous API calls, at a cost the finance team has flagged as too high. What should the team do to reduce cost for this specific workload?

    Sources4

    6.Trading cost against intelligence deliberately

    Now the levers that do cost quality. The first is not model choice: several models expose an effort parameter that trades intelligence for latency and cost inside a single model, and tuning effort is often a better lever than switching models. Defaults differ — start at the default (high) on Claude Fable 5.1 and Claude Opus 5, at medium on Claude Opus 5.5 — and move from there on your evals.

    Model choice then has two sensible entry points. Start cheap with Claude Haiku 4.5, test thoroughly, and upgrade only for specific capability gaps, which suits prototyping, tight latency requirements, cost-sensitive work, and high-volume straightforward tasks. Or start capability-first with Claude Opus 5.5, optimise prompts for it, then increase efficiency by lowering effort or downgrading later — the right route for complex reasoning and high-autonomy agentic work. Either way, compare candidates on cost per completed task, not cost per token, because a cheap model that fails and retries is not cheap.

    Matching a situation to its lever
    Your situationDo this
    Any workload, any modelTurn on prompt caching and trim unneeded tokens; both are free
    Costs are too high; quality is fineSweep effort down on your current model
    Quality isn't good enoughIf you lowered effort, restore it; otherwise try the next tier up at low effort
    Attempts end with stop_reason: max_tokensRaise max_tokens
    You can check outputs (tests, a verifier)Run everything at low effort and re-run failures at high
    A lower-cost model stalls only on hard decisionsAdd a frontier advisor
    The work exceeds one context windowDelegate partitions to cheaper workers

    Two rows in that table are worth dwelling on. If you can verify outputs automatically, run everything at low effort and re-run only the failures at high; on the coding benchmark measured, the pass rate held at about half the cost. And raising max_tokens is nearly free insurance: 64,000 covered all but 2 of 14,000 measured turns, and 128,000 cost nothing extra per solved task, because you only pay for tokens generated.

    Architecture is the last trade-off. Multi-model strategies pair a lower-cost model with a frontier one so most tokens bill at the lower rate, in two measured shapes: an executor that escalates hard decisions to an advisor, and an orchestrator that delegates bulk work to cheaper workers. The advisor shape only pays when the advisor is priced well above the executor and is actually consulted, so price the advisor's model alone at low effort and measure the consult rate first.

    For work that can wait, batch processing is the blunt instrument: submit requests asynchronously, with most batches finishing in under an hour and costs reduced by 50%. Results are accessible when all messages complete or after 24 hours, whichever comes first, and batches expire if they do not complete in 24 hours. That is the latency you are selling for the discount.

    Finally, treat every number in this lesson as a starting hypothesis. The published results are Anthropic-internal and directional rather than guarantees, so the last step of any optimisation is measuring the lever on your own workload.

    An agentic coding assistant runs long sessions that accumulate hundreds of tool calls, and the conversation regularly approaches the model's context window limit before the task is finished. Engineers want strategies that let the session continue productively without losing critical state. Which approaches address this goal?(Select 3)

    Sources516

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.The biggest cost saving in an agent loop comes from moving to a cheaper model.Why is that wrong?

      Caching was the largest lever by a wide margin in Anthropic's measurements, and cache reads are routinely the largest single component of task cost — so it outweighs most model-choice decisions.

      Covered in Prompt caching is the first lever, and usually the biggest

    2. 2.If users pause between turns, the 1-hour cache duration is the cheaper setting.Why is that wrong?

      Only when pauses land inside the hour. A gap over an hour expires both durations, and the 1-hour setting then re-writes the prefix at its higher write price, losing on every such gap.

      Covered in Picking the cache duration from real pause data

    3. 3.Batch processing and context editing are free wins, since neither changes the model or the prompt.Why is that wrong?

      Batch processing trades latency for its discount, and context editing cost more than it saved in the run measured — so both carry a caveat despite sitting near the free-win group.

      Covered in Two kinds of levers: free wins and trade-offs

    4. 4.Within one model the cost-quality point is fixed, so reducing cost means switching models.Why is that wrong?

      The effort parameter trades intelligence for latency and cost inside a single model, and is often the better lever to reach for first.

      Covered in Trading cost against intelligence deliberately

    5. 5.Fast mode is a way to cut cost as well as latency, since faster runs are cheaper runs.Why is that wrong?

      Fast mode buys output speed at premium pricing. It is a latency lever that raises cost, so it belongs with the trade-offs rather than the free wins.

      Covered in Latency: what to measure and what moves it

    6. 6.A token count measured on your current model carries over to the model you are migrating to.Why is that wrong?

      Counts depend on the tokenizer of the model you pass. The tokenizer introduced with Claude Opus 4.7 counts a prompt roughly 30 percent higher than earlier models, so budgets must be re-measured per model.

      Covered in Token hygiene, measured before you send

    7. 7.A 5-minute cache entry survives 5 minutes after the previous response finishes.Why is that wrong?

      The lifetime is measured from the start of the request that writes or reads the entry, so a long streaming response eats into the window before the next request must arrive.

      Covered in Picking the cache duration from real pause data

    Practise it for real

    Measure the token cost of one real prompt, then size the migration cost of moving it to a newer tokenizer — before spending anything on generation.

    1. 1.Call client.messages.count_tokens with your production system prompt and one representative user message, passing the model you run today.

      Why: Counting is free and returns the input total, so the prompt budget becomes a measurement instead of an estimate.

      You should see: A JSON response of the form {"input_tokens": 14} with your own total.

    2. 2.Add your tool definitions to the same call and count again.

      Why: Tool schemas are resent on every turn and are easy to forget in a budget; the documented weather-tool example counts 403 input tokens.

      You should see: A visibly higher input_tokens, showing what the tool block costs on each request.

    3. 3.Repeat the call with model set to a model on the newer tokenizer, such as claude-fable-5-1, and compare the two input_tokens values.

      Why: The endpoint counts under the tokenizer of the model you pass, and the Claude Opus 4.7 tokenizer counts roughly 30 percent higher.

      You should see: A higher count for the same request, which is the migration premium for your specific content.

    4. 4.Add cache_control at the top level of a Messages request so the growing prefix is cached automatically, then inspect the usage object in the response.

      Why: Automatic caching moves the breakpoint to the last cacheable block as the conversation grows, so you do not maintain markers by hand.

      You should see: usage fields showing cache-write tokens on the first call and cache-read tokens on the next call with the same prefix.

    5. 5.Run several turns seconds apart and compute the share of input tokens read from cache.

      Why: A median agent loop reads 84% of its input from cache; below about 80% something is breaking the prefix.

      You should see: A cache-read share you can compare against the 84% median and the 94% top-decile figure.

    Stuck? Get a nudge

    If the count endpoint rejects your request, check for a server tool, an MCP connector, or an image given by URL — send images and PDFs as base64 instead.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Free wins cut spend without touching quality: prompt caching, token hygiene, a prompt audit against the model you are running”
      ↩︎ Two kinds of levers: free wins and trade-offs
      “Tradeoffs exchange cost for intelligence: model choice, effort, output caps and task budgets, an elapsed-time clock, and multi-model architectures.”
      ↩︎ Two kinds of levers: free wins and trade-offs
      “batch processing trades latency for its discount”
      ↩︎ Two kinds of levers: free wins and trade-offs
      “Turn on prompt caching before any other lever, because every turn of an agentic task resends the entire growing conversation”
      ↩︎ Prompt caching is the first lever, and usually the biggest
      “A 40-turn task sends its first turn 40 times, so task cost grows with roughly the square of turn count.”
      ↩︎ Prompt caching is the first lever, and usually the biggest
      “agent loops read a median 84% of their input from the cache”
      ↩︎ Prompt caching is the first lever, and usually the biggest
      “the 1-hour duration pays only when at least about 40% of long pauses end within the hour”
      ↩︎ Picking the cache duration from real pause data
      “Keeping the 5-minute cache warm cost 13% to 20% less per session than the 1-hour cache whenever pauses ran for minutes”
      ↩︎ Picking the cache duration from real pause data
      “Compare on cost per completed task, not per token”
      ↩︎ Trading cost against intelligence deliberately
      “Run everything at low effort and re-run failures at high”
      ↩︎ Trading cost against intelligence deliberately
      “directional, not guarantees, so measure on your own workload”
      ↩︎ Trading cost against intelligence deliberately
      “The first group of levers on this page moves a workload toward that frontier by cutting cost without touching quality”
      ↩︎ Key concept
      “cache reads are routinely the largest single component of task cost, making caching worth more than most model-choice decisions”
      ↩︎ Exam trap 1
      “A gap over an hour expires both durations, and the 1-hour setting then re-writes the prefix at its higher write price”
      ↩︎ Exam trap 2
      “context editing, a token-hygiene lever, cost more than it saved in the run measured in this section”
      ↩︎ Exam trap 3
    2. 2.
      “Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.”
      ↩︎ Prompt caching is the first lever, and usually the biggest
      “By default, the cache has a 5-minute lifetime. The cache is refreshed for no additional cost each time the cached content is used.”
      ↩︎ Picking the cache duration from real pause data
      “The lifetime is measured from the start of the request that writes or reads the cache entry, not from the end of its response.”
      ↩︎ Exam trap 7
    3. 3.
      “Optimize prompts to a specific length”
      ↩︎ Token hygiene, measured before you send
      “Send images and PDFs as base64 to count them.”
      ↩︎ Token hygiene, measured before you send
      “The token counting endpoint counts under the tokenizer of the model you pass.”
      ↩︎ Token hygiene, measured before you send
      “roughly 30 percent higher than on models before Claude Opus 4.7”
      ↩︎ Exam trap 6
    4. 4.
      “Time to first token (TTFT): This metric measures the time it takes for the model to generate the first token of the response”
      ↩︎ Latency: what to measure and what moves it
      “The fewer tokens the model has to process and generate, the faster the response will be.”
      ↩︎ Latency: what to measure and what moves it
      “Streaming is a feature that allows the model to start sending back its response before the full output is complete.”
      ↩︎ Latency: what to measure and what moves it
    5. 5.
      “Several Claude models support an effort parameter that trades intelligence for latency and cost within a single model.”
      ↩︎ Trading cost against intelligence deliberately
      “Multi-model strategies pair a lower-cost model with a frontier model so that most tokens are billed at the lower rate.”
      ↩︎ Trading cost against intelligence deliberately
      “Tuning effort is often a better lever than switching models.”
      ↩︎ Exam trap 4
      “delivers up to 2.5x higher output speed at premium pricing”
      ↩︎ Exam trap 5
    6. 6.
      “most batches finishing in less than 1 hour while reducing costs by 50% and increasing throughput”
      ↩︎ Trading cost against intelligence deliberately

    Ready to test yourself?

    Practise the 12 questions on this subdomain.