What you will be able to do
- Write basic, guided and structured chain-of-thought prompts
- Explain how extended thinking relates to manual chain of thought, including its token cost
- Adjust reasoning depth on Claude Sonnet 5 with adaptive thinking and the effort parameter, then measure the effect
1.Chain of thought: ask for the reasoning before the answer
Chain-of-thought (CoT) prompting asks Claude to reason step by step before giving its answer. It helps because a model answering in one pass gets no scratch work. On a complex task, the quality of the answer depends on intermediate steps. Without room to reason, those steps get squeezed into the answer or skipped. CoT gives them space in the output.
The simplest version adds one sentence to the prompt. Claude's blog shows it on a donor-email task:
Draft personalized emails to donors asking for contributions to this year's Care for Kids program.
Program information:
<program>
{{PROGRAM_DETAILS}}
</program>
Donor information:
<donor>
{{DONOR_DETAILS}}
</donor>
Think step-by-step before you write the email."Think step-by-step" leaves it to Claude to decide what to think about. Guided CoT names the stages: first work out which messaging would appeal to this donor given their history, then which parts of the program would resonate, and only then write the email. The blog suggests CoT when the task needs several analytical steps, when you need reasoning you can review, when you want specific factors considered, or when extended thinking isn't available.
| Variant | What you add to the prompt | What you gain |
|---|---|---|
| Basic | "Think step-by-step" | Reasoning happens, but Claude chooses its shape |
| Guided | Named reasoning stages in order | Claude considers the specific factors you list |
| Structured | Tags separating reasoning from the final answer | The answer can be separated from the reasoning |
2.Keep the reasoning apart from the answer
With basic and guided CoT, the reasoning and the answer come back as one block of text. If code consumes the result, or a user should see only the email, that is a problem. Structured CoT fixes it with tags: reasoning goes in one tag and the deliverable goes in another. This uses the same XML-tag habit the prompting guide recommends for any complex prompt, since tags help Claude parse content unambiguously.
Think before you write the email in <thinking> tags. First, analyze what messaging would appeal to this donor. Then, identify relevant program aspects. Finally, write the personalized donor email in <email> tags, using your analysis.While reviewing a few-shot prompt for a sentiment classifier, an engineer notices that all five included examples are short, positive-sentiment product reviews written in the same sentence structure. Model outputs on negative or mixed reviews are now unreliable. What is the most likely cause and fix?
Correct answer: A — The examples are not diverse enough and Claude has picked up unintended patterns from the repeated structure; add negative, mixed, and structurally varied examples
- A. Correct. Examples must be diverse and cover edge cases so Claude does not pick up unintended patterns; a set that is uniform in sentiment and structure fails to represent the full task and biases outputs toward that pattern.
- B. Incorrect. Missing <thinking> tags is unrelated to a diversity problem; sentiment classification of this kind does not require an intermediate reasoning trace to fix a biased example set.
- C. Incorrect. Simply adding more examples without addressing their lack of diversity would likely reinforce the same unintended pattern rather than correct it.
- D. Incorrect. Example placement relative to instructions affects clarity but does not resolve a root diversity problem in the content of the examples themselves.
3.Extended thinking: built-in reasoning
Manual CoT puts the reasoning in the visible answer. Extended thinking moves it into its own part of the response. When thinking is active, Claude works through the problem before answering. It restates the question, tries approaches, checks intermediate results, and drops paths that fail. That reasoning comes back in thinking content blocks ahead of the text blocks.
{
"content": [
{
"type": "thinking",
"thinking": "Let me break this down. The question has two parts, so I'll start with the simpler one and use its result to constrain the second...",
"signature": "WaUjzkypQ2mUEVM36O2Txu...."
},
{
"type": "text",
"text": "Based on my analysis..."
}
]
}Two details come up on exams. First, the thinking text you get back is a summary, not the raw chain of thought. On many models the default is to return thinking blocks with an empty thinking field. Second, thinking is not free. Reasoning tokens are billed as output tokens even when you don't see the text, and they count toward max_tokens along with the answer.
The blog says that where extended thinking is available, it is generally better than manual CoT. It also says explicit CoT prompting can still help on complex tasks even with thinking on, because the two approaches are complementary. Manual CoT is still the right tool when you need reasoning in the answer that you can review yourself.
4.Tuning reasoning depth on Claude Sonnet 5
On Claude Sonnet 5, whether and how much Claude thinks depends on two settings. Adaptive thinking is on by default: a request with no thinking field runs with adaptive thinking. To turn it off, pass thinking: {type: "disabled"}. The older manual form, thinking: {type: "enabled", budget_tokens: N}, has been removed and returns a 400 error. The replacement is adaptive thinking plus the effort parameter, which defaults to high. xhigh is recommended for the hardest coding and agentic work. low is for short, scoped, latency-sensitive tasks that don't need much intelligence.
The guide says which to change first. If reasoning on complex problems is shallow, raise effort to high or xhigh instead of trying to fix it with prompt wording. If latency forces you to stay on low effort, add a targeted CoT-style line: "This task involves multistep reasoning. Think carefully through the problem before responding." Thinking can also be steered down. If large system prompts cause more thinking blocks than you want, a line saying to think only when it will meaningfully improve the answer, and otherwise respond directly, can reduce them.
Sonnet 5 turns adaptive thinking on by default, and max_tokens is a hard limit on thinking plus response text. max_tokens needs to be raised for workloads that used to run without thinking.
Sources4
5.Measure whether the reasoning helps
None of these techniques should be applied on faith. Anthropic's classification guide measured every step on an insurance-ticket classifier. A simple classifier reached about 70% accuracy. Retrieval raised that to 94%. Adding chain-of-thought reasoning raised it to 97% and resolved most of the remaining edge cases. The guide credits the gain to explicit reasoning about each decision, which helped Claude tell ambiguous cases apart.
The measurements are what made that result useful: they showed which techniques added value and how much. The Sonnet 5 guide gives the same instruction for any prompting change: measure its effect on performance. Reasoning costs output tokens and latency. It is worth adding only when your evaluation shows it pays for itself.
An engineer is building a Claude prompt that analyzes several long internal documents and answers a question about them. Responses sometimes ignore details buried in the middle of the documents. Which of the following changes are consistent with documented long-context prompting guidance? (Select all that apply.)(Select 3)
Correct answers: A, B, D — Place the long documents near the top of the prompt, above the query and instructions; Wrap each document in <document> tags with source and content subtags to structure the input clearly; Ask Claude to quote the relevant passages inside <quotes> tags before producing the final analysis
- A. Correct. Long-form documents should be placed near the top of the prompt, above the query and instructions, which improves performance across models.
- B. Correct. Wrapping each document in structured tags such as <document>, <source>, and <document_content> helps Claude parse multi-document input unambiguously.
- C. Incorrect. Placing the query at the very top, before the documents, contradicts documented guidance; queries placed at the end can improve response quality, especially with complex multi-document inputs.
- D. Correct. Asking Claude to quote relevant passages before analysis grounds the response and helps it cut through noise in long documents.
- E. Incorrect. Long-context guidance explicitly addresses multi-document analysis via structured tags rather than requiring the input to be reduced to a single document.
- F. Incorrect. Removing XML structure works against the documented recommendation to use tags for clarity when handling multiple documents and metadata.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.On Claude Sonnet 5, you can control reasoning depth by setting thinking: {type: "enabled", budget_tokens: N}.Why is that wrong?
Manual extended thinking has been removed from Sonnet 5 and the request fails. Use adaptive thinking with the effort parameter instead.
Covered in Tuning reasoning depth on Claude Sonnet 5
2.If the thinking text isn't returned, the reasoning costs nothing.Why is that wrong?
Reasoning tokens are billed as output tokens whether or not the thinking text is shown, and they count toward max_tokens.
Covered in Extended thinking: built-in reasoning
3.When Claude reasons too shallowly on hard problems, the first fix is to add more "think carefully" instructions.Why is that wrong?
The Sonnet 5 guidance is to raise effort first. Prompting for deeper thinking is the fallback when effort must stay low for latency.
Covered in Tuning reasoning depth on Claude Sonnet 5
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“the quality of the answer depends on intermediate work that would otherwise be compressed into the response itself or skipped”
↩︎ Chain of thought: ask for the reasoning before the answer“When thinking is active, Claude works through the problem in its own words before answering”
↩︎ Extended thinking: built-in reasoning“what you see is never the raw chain of thought: the text in a thinking block is a summary of Claude's reasoning.”
↩︎ Extended thinking: built-in reasoning“the tokens Claude spends reasoning are billed as output tokens, even when the thinking text isn't returned to you”
↩︎ Exam trap 2 - 2.https://claude.com/blog/best-practices-for-prompt-engineeringSecondary source
“Chain of thought (CoT) prompting involves requesting step-by-step reasoning before answering.”
↩︎ Chain of thought: ask for the reasoning before the answer“Use tags to separate reasoning from the final answer.”
↩︎ Keep the reasoning apart from the answer“When available, extended thinking is generally preferable to manual chain of thought prompting.”
↩︎ Extended thinking: built-in reasoning“The two approaches are complementary, not mutually exclusive.”
↩︎ Extended thinking: built-in reasoning - 3.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practicesOfficial docs
“XML tags help Claude parse complex prompts unambiguously, especially when your prompt mixes instructions, context, examples, and variable inputs.”
↩︎ Keep the reasoning apart from the answer - 4.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5Official docs
“On Claude Sonnet 5, adaptive thinking is on by default.”
↩︎ Tuning reasoning depth on Claude Sonnet 5“This task involves multistep reasoning. Think carefully through the problem before responding.”
↩︎ Tuning reasoning depth on Claude Sonnet 5“As always, measure the effect of any prompting changes on performance.”
↩︎ Measure whether the reasoning helps“Manual extended thinking (thinking: {type: "enabled", budget_tokens: N}) is not supported on Claude Sonnet 5 and returns a 400 error.”
↩︎ Exam trap 1“If you observe shallow reasoning on complex problems, raise effort to high or xhigh rather than prompting around it.”
↩︎ Exam trap 3 - 5.
“Chain-of-thought reasoning pushed our accuracy to 97%, resolving most of the edge cases that challenged the previous approaches.”
↩︎ Measure whether the reasoning helps“By explicitly reasoning through each classification decision, Claude better distinguishes between ambiguous cases.”
↩︎ Measure whether the reasoning helps