What you will be able to do
- Trim input and output tokens without letting max_tokens cut off reasoning.
- Improve perceived responsiveness with streaming and omitted thinking display, and say what each one does and does not change.
- Decide when fast mode is worth its premium, and list its constraints.
- Back a configuration decision with benchmark evidence and measured trade-off patterns.
1.Token hygiene, and where max_tokens stops being free
The cheapest speed-up removes tokens nobody needed. Anthropic's latency guidance is simple: the fewer tokens the model has to process and generate, the faster it responds. It gives four tactics. Be clear but concise, and don't cut so far that instructions become ambiguous, because Claude has no context on your use case. Ask Claude directly for shorter responses. Set max_tokens as a hard cap on output length. Experiment with temperature, since lower values such as 0.2 can sometimes give more focused, shorter output. The guidance adds that finding the balance between clarity, quality and token count may take experimentation.
The max_tokens cap is where a free lever quietly becomes a trade. When thinking is active, thinking tokens count toward max_tokens. A cap sized for the visible answer can therefore cut the reasoning short, and the answer gets worse. The cost guide's rule: if attempts end with stop_reason: max_tokens, raise the cap. In its measurements, 64,000 covered all but 2 of 14,000 turns at the default effort. The effort documentation likewise tells you to set a large max_tokens at the higher effort levels.
The 300-token cap covers thinking and response text together, so the model runs out of room to reason. Look for stop_reason: max_tokens in the responses and raise the cap. To keep replies short, ask for concise answers in the prompt.
2.Perceived latency: streaming and thinking display
Some latency problems are really about how long the user waits to see something. Streaming lets the model start sending its response before the whole output is complete. You can then update the interface or do other work in parallel as tokens arrive. The total work stays the same, but the wait before the user sees output gets much shorter.
Thinking adds a wrinkle, because on thinking models the reasoning comes before the answer. The display field on the thinking configuration decides what you get back. "summarized" returns a readable summary of the reasoning. "omitted" returns thinking blocks with an empty thinking field and the encrypted signature. "updates" (beta) returns short progress updates between tool calls. If your app never shows thinking to users, set display to "omitted". When streaming, the server then skips the thinking tokens and sends only the signature, so the answer text starts arriving sooner.
Omitted display costs no accuracy, because the reasoning still happens and the block is passed back the same way in multi-turn conversations. It is already the default on Claude Opus 5.5, Claude Fable 5.1, Claude Sonnet 5 and several other current models. The setting you need to justify is therefore "summarized", since that is the one that adds time before the first answer text.
3.Fast mode: paying for speed instead of trading accuracy
When users need more output throughput from an Opus-class model, fast mode gives up to 2.5x higher output tokens per second on Claude Opus 5.5, Claude Opus 5 and Claude Opus 4.8. It runs the same model weights with a faster inference configuration, so you trade money for speed, not accuracy. On Claude Opus 5.5 it costs $8 per million input tokens and $40 per million output tokens. You opt in with speed="fast" and a beta header.
client = anthropic.Anthropic()
response = client.beta.messages.create(
model="claude-opus-5-5",
max_tokens=4096,
speed="fast",
betas=["fast-mode-2026-02-01"],
messages=[
{"role": "user", "content": "Refactor this module to use dependency injection"}
],
)The documentation's fallback example handles a RateLimitError on a fast request by deleting the speed parameter and retrying at standard speed. That keeps requests flowing, but check the constraints below before you rely on it. A retry at the other speed does not share the cached prefix.
| Area | Behaviour |
|---|---|
| Prompt caching | Switching between fast and standard speed invalidates the prompt cache |
| TTFT | Benefits are focused on output tokens per second (OTPS), not time to first token |
| Batch API | Fast mode is not available with the Batch API |
| Priority Tier | Fast mode is not available with a Priority Tier commitment |
| Claude Platform on AWS | Fast mode is not currently available |
Sources4
4.Justifying the configuration with measurements
A configuration decision is justified when measurements on your own workload back it. The model-selection guide lists the steps: build benchmark tests specific to your use case, test with your real prompts and data, and compare accuracy, response quality and edge-case handling across models before weighing performance against cost. The cost guide adds that you should compare models on cost per completed task, not per token. It also stresses that its own numbers are directional, so you have to measure your own workload.
| Situation | Do this | Measured effect or condition |
|---|---|---|
| You want agent runs to finish sooner | Tell the model that time matters, and show it the elapsed time | 33% to 69% less time, 28% to 54% lower cost per task, scores up to 1.9 points lower |
| Costs are too high; quality is fine | Sweep effort down on your current model | Trade-off lever: confirm quality on evals |
| Quality isn't good enough | Restore effort if you lowered it; otherwise try the next tier up at low effort | Moves back up the frontier |
| You can check outputs (tests, a verifier) | Run everything at low effort and re-run failures at high | Pass rate held at about half the cost on the coding benchmark |
| A lower-cost model stalls only on hard decisions | Add a frontier advisor | Pays off when priced well above the executor and actually consulted |
The elapsed-time row is the clearest accuracy-latency trade in the guidance. Runs finish much sooner, and the accuracy cost is stated: scores up to 1.9 points lower. Justifying it means arguing that the drop is acceptable for your use case. The re-run-failures row shows the opposite approach. When a verifier can spot wrong answers, low effort becomes the fast default and only failures pay for high effort, so the pass rate holds. Multi-model setups use the same idea at architecture level: an executor passes hard decisions to an advisor, or an orchestrator hands bulk work to cheaper workers.
A financial analytics application uses extended thinking to solve complex multi-step valuation problems. Response latency has become unacceptable for end users, but outputs must remain highly accurate given the complexity of the problems. Which adjustment best balances these constraints?
Correct answer: A — Keep the existing thinking budget but set display to "omitted" so text streaming begins immediately, since full thinking tokens still execute but their streaming overhead is removed rather than the reasoning itself being cut.
- A. Correct. Setting display to "omitted" eliminates thinking-token streaming overhead and lets text streaming begin immediately, improving perceived latency while preserving the full reasoning depth needed for accurate valuations.
- B. Incorrect. Disabling thinking entirely removes the step-by-step reasoning that complex multi-step valuation problems rely on for accuracy, trading away the quality the team explicitly needs to preserve.
- C. Incorrect. Forcing a uniformly small thinking budget across all requests would degrade accuracy on genuinely complex valuation problems that need more thorough reasoning, rather than only trimming the requests that can tolerate it.
- D. Incorrect. A larger thinking budget increases the tokens spent on reasoning and does not by itself reduce latency; it works against the team's goal of reducing unacceptable response times.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Fast mode is the fix when users complain about the delay before the first token.Why is that wrong?
Fast mode raises output tokens per second. For TTFT, look at streaming and omitted thinking display instead.
Covered in Fast mode: paying for speed instead of trading accuracy
2.Setting thinking display to "omitted" lowers the bill because less thinking text is returned.Why is that wrong?
The full thinking tokens are still billed. Omitting speeds up the first answer text but does not reduce cost.
Covered in Perceived latency: streaming and thinking display
3.A tight max_tokens is a safe latency cap that only limits answer length.Why is that wrong?
Thinking tokens count toward max_tokens, so a tight cap can cut the reasoning short as well as the answer.
Covered in Token hygiene, and where max_tokens stops being free
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-latencyOfficial docs
“The fewer tokens the model has to process and generate, the faster the response will be.”
↩︎ Token hygiene, and where max_tokens stops being free“Use the max_tokens parameter to set a hard limit on the maximum length of the generated response.”
↩︎ Token hygiene, and where max_tokens stops being free“This can significantly improve the perceived responsiveness of your application, as users can see the model's output in real time.”
↩︎ Perceived latency: streaming and thinking display - 2.https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligenceOfficial docs
“Raise max_tokens; 64,000 covered all but 2 of 14,000 turns measured at the default effort”
↩︎ Token hygiene, and where max_tokens stops being free“Compare on cost per completed task, not per token”
↩︎ Justifying the configuration with measurements“runs took 33% to 69% less time at a 28% to 54% lower cost per task, with scores up to 1.9 points lower”
↩︎ Justifying the configuration with measurements - 3.
“The primary benefit is faster time-to-first-text-token when streaming”
↩︎ Perceived latency: streaming and thinking display“You're still charged for the full thinking tokens. Omitting reduces latency, not cost.”
↩︎ Exam trap 2“Thinking tokens count toward max_tokens, so set it high enough to leave room for both the thinking and the response text.”
↩︎ Exam trap 3 - 4.
“Fast mode runs the same model with a faster inference configuration. There is no change to intelligence or capabilities.”
↩︎ Fast mode: paying for speed instead of trading accuracy“Switching between fast and standard speed invalidates the prompt cache.”
↩︎ Fast mode: paying for speed instead of trading accuracy“Speed benefits are focused on output tokens per second (OTPS), not time to first token (TTFT)”
↩︎ Exam trap 1 - 5.
“Create benchmark tests specific to your use case - having a good evaluation set is the most important step in the process.”
↩︎ Justifying the configuration with measurements