What you will be able to do
- Name the latency metric (baseline latency or time to first token) a configuration change is meant to improve.
- Tell apart levers that trade intelligence for speed from levers that only remove waste.
- Justify a speed-first or capability-first starting model for a given workload.
- Choose an effort level and defend it, allowing for model-specific defaults and behaviour.
Key concept
Effort as the intelligence-for-latency dial — Most latency decisions buy speed by spending less reasoning. The effort parameter makes that trade explicit inside one model, so model tier and effort are the two dials you have to justify with evals.
1.Say which latency you are optimising
Before you can justify a configuration, you have to say which latency you mean. Anthropic's latency guidance defines latency as the time the model takes to process a prompt and generate an output. Three things shape it: the size of the model, how complex the prompt is, and the infrastructure between the model and the user.
The guidance names two measurements. Baseline latency is the overall time the model takes to process the prompt and produce a response. It gives you a rough sense of the model's speed. Time to first token (TTFT) is the time from sending the prompt to receiving the first token of the response. TTFT is what users notice when output is streamed, because it decides how long they look at a blank screen.
The distinction matters because each lever moves a different metric. Some levers shorten total generation time. Others only change when the first visible text appears. A justification that says "this makes it faster" without naming the metric is incomplete, because a lever can improve one metric and leave the other untouched.
Sources1
2.Two kinds of lever: trades and free wins
Anthropic's cost-and-intelligence guide pictures configuration as a frontier. Some levers exchange capability for savings. Others remove waste and leave quality alone. The guide frames this in terms of cost, but the same split works for latency. Its list of real trade-offs is model choice, effort, output caps and task budgets, an elapsed-time clock, and multi-model architectures. Each of these makes a response cheaper or faster by making it less thorough. Its free wins, such as prompt caching and token hygiene, cut spend without touching quality.
That gives you a working rule. When you pull a trade-off lever, you need a measurement showing that quality held. When you pull a free lever, you don't. A scenario that asks you to cut latency "without affecting accuracy" is pointing you at the free group. A scenario that accepts some loss of capability is inviting the trade-offs, and the right answer usually names how you would check the loss.
Sources2
3.Model choice: speed-first or capability-first
The latency guidance calls model choice one of the most direct levers. For speed-critical applications it recommends Claude Haiku 4.5 as the fastest option that still keeps high intelligence.
client = anthropic.Anthropic()
# For time-sensitive applications, use Claude Haiku 4.5
message = client.messages.create(
model="claude-haiku-4-5",
max_tokens=100,
messages=[
{
"role": "user",
"content": "Summarize this customer feedback in 2 sentences: [feedback text]",
}
],
)
print(message.content[0].text)The model-selection guide gives two ways to start. Speed-first means you begin with Claude Haiku 4.5, test thoroughly, and upgrade only when you find a specific capability gap. It suits prototyping, tight latency requirements, cost-sensitive deployments and high-volume straightforward tasks. Capability-first means you implement with Claude Opus 5.5, optimise your prompts for it, and then gain efficiency over time by lowering effort or moving to a smaller model. If evals at xhigh or max effort still fall short on demanding reasoning or long agentic work, you step up to Claude Fable 5.1. Capability-first suits complex reasoning, nuanced understanding and applications where accuracy matters more than cost. The guide also notes that most workloads start with Claude Opus 5.5.
| When you need... | Start with | Example use cases |
|---|---|---|
| The highest available capability | Claude Fable 5.1 | Agent sessions that run for hours, multistep deep research |
| Complex agentic coding and enterprise work | Claude Opus 5.5 | Multihour autonomous coding agents, large-scale refactoring |
| Speed and capability for everyday workloads | Claude Sonnet 5 | Code generation, data analysis, agentic tool use |
| The lowest latency and price, with extended thinking | Claude Haiku 4.5 | Real-time applications, high-volume processing, sub-agent tasks |
A team is building a real-time customer support chat feature that must respond within one to two seconds, handles a very high volume of simple intent-classification queries, and has a tight cost budget. Which model should they start with to satisfy the latency and cost requirements while keeping accuracy at an acceptable level for this task?
Correct answer: B — Claude Haiku 4.5
- A. Incorrect. Claude Opus 4.8 targets complex agentic coding and enterprise work with moderate comparative latency, which is more capability than a simple, high-volume classification task needs and adds unnecessary cost and delay.
- B. Correct. Claude Haiku 4.5 is described as the fastest model with near-frontier intelligence at the most economical price point, explicitly suited to real-time applications and high-volume, cost-sensitive workloads like this one.
- C. Incorrect. Claude Fable 5 is positioned as the most capable model for long-running agents with slower comparative latency and premium pricing, which does not match a latency-sensitive, cost-constrained classification workload.
- D. Incorrect. Claude Sonnet 5 offers a strong balance of speed and intelligence for frontier coding and agentic work, but it is not the most cost-effective or fastest option when the task itself is simple and volume is very high.
4.Effort: the fine-grained dial inside one model
Model choice is the coarse dial and effort is the fine one. Effort sets how many tokens Claude spends on a response. That covers the text, tool calls and their arguments, and thinking when it is active. Lower effort also produces fewer and terser tool calls, which matters in agent loops where every call is a round trip. The model-selection guide says outright that tuning effort is often a better lever than switching models.
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=4096,
messages=[
{
"role": "user",
"content": "Analyze the trade-offs between microservices and monolithic architectures",
}
],
output_config={"effort": "medium"},
)| Level | Description | Typical use case |
|---|---|---|
| max | Absolute maximum capability with no constraints on token spending | Tasks requiring the deepest possible reasoning |
| xhigh | Extended capability for long-horizon work | Long-running agentic and coding tasks (over 30 minutes) |
| high | Spends as many tokens as the task needs; the default except on Claude Opus 5.5 | Complex reasoning, difficult coding problems, agentic tasks |
| medium | Balanced approach with moderate token savings; the default on Claude Opus 5.5 | Agentic tasks balancing speed, cost, and performance |
| low | Most efficient; significant token savings with some capability reduction | Simpler tasks needing the best speed and lowest costs, such as subagents |
Effort also steers thinking. Thinking is adaptive: Claude decides for each request whether to reason and how deeply, and effort is the main control over that decision. Thinking improves results on maths, coding and long agentic work, but it has a cost. Reasoning tokens are billed as output tokens and count toward max_tokens, so high effort adds real latency. If you want Claude to think less often, lower effort before you try steering it with the prompt. On Claude Opus 5.5, adaptive thinking is always on and cannot be disabled, which makes effort the main lever there.
Three cautions come up when you defend an effort setting. First, defaults differ between models. Most models default to high, but Claude Opus 5.5 defaults to medium, so a request that omits effort runs one level lower than it did on Claude Opus 5. Second, don't carry settings over from an earlier model. Run a fresh effort sweep on your own evals. Third, effort is not a length control. On Claude Opus 5 it changes how much the model thinks, not how long the visible response is, so ask for brevity in the prompt instead. At the top of the range, the Claude Opus 4.7 guidance warns that max often costs far more for small gains.
An engineering team runs many parallel subagents performing simple lookups inside a larger agentic workflow. They want to minimize latency and token cost per subagent call without materially degrading the quality needed for those simple lookups. Which effort configuration decision best matches this need?
Correct answer: A — Set effort to low for the subagent calls, since low effort is designed for simple, high-volume tasks like subagent lookups where speed and cost outweigh marginal quality gains.
- A. Correct. Documentation describes low effort as the most efficient setting, recommended for simpler tasks needing the best speed and lowest cost, explicitly including subagent tasks.
- B. Incorrect. Max effort removes constraints on token spending for the deepest possible reasoning, which is intended for frontier-difficulty problems, not narrowly scoped simple lookups, and would add unnecessary latency and cost.
- C. Incorrect. Xhigh is intended for long-horizon agentic and coding tasks with token budgets in the millions, which is a mismatch for a narrowly scoped, simple subagent lookup.
- D. Incorrect. Leaving effort at the high default does not address the stated goal of minimizing latency and token cost; high effort spends more tokens than necessary for simple lookups.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If a workload is too slow, the first fix is to move to a smaller model.Why is that wrong?
Anthropic says adjusting effort on your current model is often a better lever than changing models. Try an effort sweep before you drop a tier.
2.A request that leaves out effort runs at high on every model.Why is that wrong?
Claude Opus 5.5 defaults to medium, so leaving effort out runs one level lower than it did on Claude Opus 5.
3.Lowering effort is a reliable way to get shorter visible answers.Why is that wrong?
On Claude Opus 5, effort changes how much the model thinks, not how long the visible response is. Ask for brevity in the prompt.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-latencyOfficial docs
“Time to first token (TTFT): This metric measures the time it takes for the model to generate the first token of the response”
↩︎ Say which latency you are optimising“Latency can be influenced by various factors, such as the size of the model, the complexity of the prompt”
↩︎ Say which latency you are optimising“One of the most direct ways to reduce latency is to select the appropriate model for your use case.”
↩︎ Model choice: speed-first or capability-first - 2.https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligenceOfficial docs
“Managing cost well means understanding how each cost lever affects output quality, because some levers trade against quality and some don't.”
↩︎ Two kinds of lever: trades and free wins“Tradeoffs exchange cost for intelligence: model choice, effort, output caps and task budgets, an elapsed-time clock, and multi-model architectures.”
↩︎ Two kinds of lever: trades and free wins - 3.
“Upgrade only if necessary for specific capability gaps.”
↩︎ Model choice: speed-first or capability-first“Applications where accuracy outweighs cost considerations”
↩︎ Model choice: speed-first or capability-first“Several Claude models support an effort parameter that trades intelligence for latency and cost within a single model.”
↩︎ Key concept“Tuning effort is often a better lever than switching models.”
↩︎ Exam trap 1 - 4.
“Lower effort also means fewer and terser tool calls.”
↩︎ Effort: the fine-grained dial inside one model“Run an effort sweep on your own evals rather than carrying settings over from an earlier model”
↩︎ Effort: the fine-grained dial inside one model“On most workloads max adds significant cost for relatively small quality gains”
↩︎ Effort: the fine-grained dial inside one model“Most Claude models default to high effort, spending as many tokens as needed for excellent results; Claude Opus 5.5 defaults to medium.”
↩︎ Exam trap 2“Effort controls thinking volume, not visible response length”
↩︎ Exam trap 3 - 5.
“If you want Claude to think less often, lower the effort level before reaching for prompt-based steering.”
↩︎ Effort: the fine-grained dial inside one model - 6.
“the tokens Claude spends reasoning are billed as output tokens, even when the thinking text isn't returned to you”
↩︎ Effort: the fine-grained dial inside one model