What you will be able to do
- Explain what thinking adds to a response and how thinking tokens are billed and counted
- Choose between adaptive thinking and manual extended thinking (budget_tokens), and pick the right display setting
- Choose an effort level for a workload and know each model's default
- Explain what fast mode speeds up, what it costs, and where it is not available
1.Thinking: scratch work before the answer
A model that answers in a single pass has to get everything right first time, with no scratch work and no chance to check itself. Thinking removes that constraint. When thinking is active, Claude works through the problem in its own words before answering. It restates the question, tries approaches, checks intermediate results and drops paths that don't hold up. This is why thinking helps most on math, coding, analysis and long-running agentic work, where the quality of the answer depends on the intermediate steps.
The reasoning arrives in thinking content blocks before the text blocks. What you see is never the raw chain of thought. It is a summary, and depending on the display setting it may be empty. Each block carries a signature, an encrypted copy of the full reasoning, which you must pass back unchanged in multi-turn and tool-use conversations.
Thinking costs tokens. Thinking tokens are billed as output tokens even when the text isn't returned to you, and they count toward max_tokens alongside the response. Set max_tokens high enough to leave room for both. Thinking also takes up context. On Claude Opus 4.5 and later Opus models, and on Claude Sonnet 4.6 and later Sonnet models, the API keeps earlier thinking blocks, which then count and bill as input on later turns. On earlier Opus and Sonnet models and on all Haiku models, the API strips them automatically.
2.Adaptive thinking versus manual extended thinking
There are two ways to configure thinking. With adaptive thinking, thinking: {type: "adaptive"}, Claude decides when to think and how deeply based on the request, so thinking usage varies from one request to the next. With manual extended thinking you set a budget: thinking: {type: "enabled", budget_tokens: N}. Manual mode is still useful when you need predictable latency or precise control over thinking costs.
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=16000,
thinking={"type": "enabled", "budget_tokens": 10000},
messages=[
{
"role": "user",
"content": "Are there an infinite number of prime numbers such that n mod 4 == 3?",
}
],
)budget_tokens must be at least 1,024 and less than max_tokens. The one exception is interleaved thinking, where the budget spans all the thinking blocks in a turn. The budget is a target, not a cap: Claude may stop well short of it, and max_tokens remains the hard ceiling. For simple tasks, start near the minimum. For complex ones, start at 16,000 or more. Above 32k, use batch processing to avoid timeouts. To see what the budget actually costs, check usage.output_tokens_details.thinking_tokens.
| Model | Thinking by default | Can you send thinking: {type: "disabled"}? |
|---|---|---|
| Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5, Claude Mythos Preview | On | No. The request is rejected. |
| Claude Opus 5 | On | Only at effort high or below. At xhigh or max it returns a 400 error. |
| Claude Sonnet 5 | On | Yes |
| Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Sonnet 4.6 | Off until you set thinking: {type: "adaptive"} | Thinking is off by default |
The display field controls what comes back. "summarized" returns readable summary text. "omitted", the default on Claude Opus 5.5, Opus 4.8 and other newer models, returns thinking blocks with an empty thinking field. Use omitted when your app doesn't show thinking to users. The main benefit is a faster time to the first text token when streaming, because the server skips streaming the thinking. It saves no money: the block is billed the same either way.
A team using Claude Sonnet 5 wants to guarantee the lowest possible latency on simple lookup queries by turning reasoning off completely for those requests. Simply omitting the thinking configuration does not achieve this. What should they do instead?
Correct answer: C — Set thinking type to disabled explicitly in the request
- A. On Claude Sonnet 5, adaptive thinking is on by default, so omitting the thinking parameter still leaves thinking active rather than turning it off.
- B. Manual thinking with type enabled and budget_tokens is rejected with a 400 error on Claude Sonnet 5, since it only supports adaptive thinking.
- C. Claude Sonnet 5 requires passing thinking type disabled explicitly to turn off adaptive thinking, since it is on by default otherwise.
- D. Lowering effort only reduces how much Claude thinks; it does not disable thinking outright, since thinking remains active by default on this model.
3.Effort levels: one dial for tokens, speed and cost
The effort parameter controls how many tokens Claude spends on a response. It trades thoroughness against efficiency on the same model. You set it with output_config.effort. It applies to every output token, including text, tool calls and thinking, so it works whether or not thinking is on. Lower effort also means fewer and terser tool calls.
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=4096,
messages=[
{
"role": "user",
"content": "Analyze the trade-offs between microservices and monolithic architectures",
}
],
output_config={"effort": "medium"},
)| Level | What it does | Typical use |
|---|---|---|
| max | Maximum capability, no limit on token spending | The deepest reasoning and most thorough analysis |
| xhigh | Extended capability for long-horizon work (not every model that supports max supports it) | Agentic and coding tasks running over 30 minutes |
| high | As many tokens as the task needs. The default on every supported model except Claude Opus 5.5 | Complex reasoning, hard coding problems, agentic tasks |
| medium | A balance with moderate token savings. The default on Claude Opus 5.5 | Agentic tasks that balance speed, cost and performance |
| low | The most efficient setting, with some loss of capability | Simple, speed-sensitive tasks such as subagents |
On Claude Opus 5.5, adaptive thinking is always on, so effort is the main control over how much the model reasons and what a request costs. A request that omits effort runs one level lower than it did on Claude Opus 5. At higher levels, set a large max_tokens, because it is a hard limit on thinking plus response text. Don't assume effort controls length: on Claude Opus 5 it governs thinking volume, not how long the visible response is, so ask for the length you want in the prompt. Anthropic's guidance for Opus 4.7 adds that max adds significant cost for small gains on most workloads, and that shallow reasoning should be fixed by raising effort.
Not reliably. On Claude Opus 5, effort controls how much the model thinks, not how long the visible response is. To get shorter answers, ask for a length in the prompt.
A prompt engineer is building an extraction task that must follow a specific output format consistently, including on tricky edge cases. They want the most reliable way to steer Claude's format and structure using a small number of well-crafted demonstrations. What should they do?
Correct answer: A — Provide three to five diverse examples that mirror the real use case and cover edge cases, wrapped in example tags
- A. Multishot prompting with three to five relevant, diverse examples wrapped in example tags is the recommended approach for steering format, tone, and structure reliably.
- B. A single untagged example gives Claude less structure to distinguish instructions from demonstrations and does not cover the diversity edge cases require.
- C. Examples are one of the most reliable levers for output format; replacing them with prose instructions alone forgoes that benefit for a format-sensitive task.
- D. A large set of near-duplicate examples risks Claude picking up unintended narrow patterns instead of generalizing to the diverse cases the task actually requires.
Sources4
4.Fast mode: the same model, faster output, higher price
Fast mode delivers up to 2.5x higher output tokens per second on Claude Opus 5.5, Claude Opus 5 and Claude Opus 4.8. You opt in with speed: "fast" and the fast-mode-2026-02-01 beta header. It is not a smaller model: the weights and behaviour are the same, with no change to intelligence. The speed-up is in output tokens per second (OTPS), not time to first token (TTFT), so the gain is most visible when streaming.
| Model | Input | Output |
|---|---|---|
| Claude Opus 5.5 | $8 USD / MTok | $40 USD / MTok |
| Claude Opus 5 / Claude Opus 4.8 | $10 USD / MTok | $50 USD / MTok |
The multiplier applies across the full context window. Prompt caching and data residency multipliers apply on top of it. There are several limits. Switching between fast and standard speed invalidates the prompt cache. Fast mode is not available with the Batch API or a Priority Tier commitment, and not currently on Claude Platform on AWS. The documented fallback pattern catches a rate-limit error and retries the request without speed.
Sources5
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Fast mode swaps in a smaller, less capable model to get its speed.Why is that wrong?
Fast mode runs the same model with a faster inference configuration. Capability doesn't change, only speed and price.
Covered in Fast mode: the same model, faster output, higher price
2.Setting display to "omitted" reduces the cost of thinking.Why is that wrong?
Omitted only hides the summary text and speeds up the first streamed text token. The thinking is billed exactly the same.
Covered in Adaptive thinking versus manual extended thinking
3.budget_tokens is a hard cap that Claude always spends in full.Why is that wrong?
The budget is a target. Claude may stop well before it, and max_tokens is the hard ceiling.
Covered in Adaptive thinking versus manual extended thinking
4.Every Claude model defaults to high effort, so omitting effort on Claude Opus 5.5 behaves the same as on Claude Opus 5.Why is that wrong?
Claude Opus 5.5 defaults to medium, one level lower than Claude Opus 5.
Covered in Effort levels: one dial for tokens, speed and cost
Practise it for real
See manual extended thinking, the thinking token count and the display setting in real API responses
1.Send the extended-thinking example request on claude-sonnet-4-6 with max_tokens=16000 and thinking={"type": "enabled", "budget_tokens": 10000}.
Why: This is manual mode: you set the thinking budget yourself.
You should see: The response contains thinking blocks followed by text blocks.
2.Read usage.output_tokens_details.thinking_tokens from the response.
Why: This field shows how many billed output tokens went on reasoning.
You should see: A count that can sit well below 10000, because the budget is a target, not a quota.
3.Resend with budget_tokens set to 500.
Why: This tests the minimum budget constraint.
You should see: The API rejects the request, because the minimum is 1,024 tokens.
4.Restore budget_tokens to 10000, add "display": "omitted" to the thinking object, and resend.
Why: This shows what omitted changes and what it leaves alone.
You should see: The thinking blocks come back with an empty thinking field but still carry a signature. The text answer is unchanged.
Stuck? Get a nudge
If a request fails, check that budget_tokens is at least 1,024 and less than max_tokens.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“When thinking is active, Claude works through the problem in its own words before answering”
↩︎ Thinking: scratch work before the answer“thinking improves performance on complex tasks like math, coding, analysis, and long-running agentic work”
↩︎ Thinking: scratch work before the answer“the tokens Claude spends reasoning are billed as output tokens, even when the thinking text isn't returned to you”
↩︎ Thinking: scratch work before the answer“which lets Claude decide when and how deeply to think based on the request”
↩︎ Adaptive thinking versus manual extended thinking“Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Opus 5.5, and Claude Mythos Preview reject thinking: {type: "disabled"}.”
↩︎ Adaptive thinking versus manual extended thinking“The primary benefit is faster time-to-first-text-token when streaming”
↩︎ Adaptive thinking versus manual extended thinking“Either way the block is billed the same and passed back the same in multi-turn conversations.”
↩︎ Exam trap 2 - 2.
“the API automatically strips previous thinking blocks from the conversation history when you pass them back”
↩︎ Thinking: scratch work before the answer - 3.
“Manual mode remains useful when your workload requires predictable latency or precise control over thinking costs.”
↩︎ Adaptive thinking versus manual extended thinking“Minimum of 1,024 tokens. The API rejects smaller values.”
↩︎ Adaptive thinking versus manual extended thinking“For thinking budgets above 32k, use batch processing to avoid networking issues.”
↩︎ Adaptive thinking versus manual extended thinking“The budget is a target rather than a strict cap.”
↩︎ Exam trap 3 - 4.
“The effort parameter lets you control how many tokens Claude spends when responding to requests.”
↩︎ Effort levels: one dial for tokens, speed and cost“Because effort applies to every output token, it works whether or not thinking is enabled.”
↩︎ Effort levels: one dial for tokens, speed and cost“Adaptive thinking is always on and can't be turned off, so effort is the primary control for how much the model reasons”
↩︎ Effort levels: one dial for tokens, speed and cost“Effort controls thinking volume, not visible response length”
↩︎ Effort levels: one dial for tokens, speed and cost“On most workloads max adds significant cost for relatively small quality gains”
↩︎ Effort levels: one dial for tokens, speed and cost“Most Claude models default to high effort, spending as many tokens as needed for excellent results; Claude Opus 5.5 defaults to medium.”
↩︎ Exam trap 4 - 5.
“up to 2.5x higher output tokens per second”
↩︎ Fast mode: the same model, faster output, higher price“Speed benefits are focused on output tokens per second (OTPS), not time to first token (TTFT)”
↩︎ Fast mode: the same model, faster output, higher price“Switching between fast and standard speed invalidates the prompt cache.”
↩︎ Fast mode: the same model, faster output, higher price“Fast mode is not available with the Batch API.”
↩︎ Fast mode: the same model, faster output, higher price“Fast mode runs the same model with a faster inference configuration. There is no change to intelligence or capabilities.”
↩︎ Exam trap 1