What you will be able to do
- Route a response by its stop_reason before parsing its content
- Detect truncated, empty and incomplete responses and recover from each correctly
- Improve format consistency with prompting when you are not using constrained decoding
- Apply verification techniques that treat confident output as unconfirmed until it is checked
1.Read stop_reason before you read the content
Defensive parsing starts before any JSON gets parsed. Every successful Messages API response includes stop_reason, which says why generation stopped. It is not an error. The request succeeded, and this field tells you whether the output is complete and whether you should use it as it is.
| stop_reason | When it occurs | What to do |
|---|---|---|
| end_turn | Claude finished its response naturally. | Use the response. |
| max_tokens | The response reached your max_tokens limit. | Raise max_tokens or continue the response. |
| stop_sequence | Claude emitted one of your stop_sequences. | Read stop_sequence to see which one fired. |
| tool_use | Claude is calling a tool. | Run the tool and return the result. |
| pause_turn | A server-tool loop reached its iteration limit. | Send the assistant content back to continue. |
| refusal | Claude declined to respond. | Read stop_details and retry on a fallback model. |
| model_context_window_exceeded | The response filled the model's context window. | Treat the response as truncated. |
Sources1
2.Truncated, empty and unfinished responses
Truncation. When stop_reason is max_tokens, the output was cut off. If the last content block is a tool_use block, that tool call is incomplete, so don't run it. Retry with a higher max_tokens instead:
# Check if response was truncated during tool use
if response.stop_reason == "max_tokens":
# Check if the last content block is an incomplete tool_use
last_block = response.content[-1]
if last_block.type == "tool_use":
# Send the request with higher max_tokens
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=4096, # Increased limit
messages=messages,
tools=tools,
)**Empty end_turn.** A response can end with end_turn and still contain no content, only 2–3 tokens. This usually happens after tool results. A common cause is adding a text block right after a tool_result, which teaches Claude to expect the user to always add text after a tool result, so it ends its turn. Send tool results without extra text. If a response still comes back empty, don't just resend it, because Claude has already decided it is done. As a last resort, add a continuation prompt in a new user message.
Unfinished server tool calls. A tool_use response can include a server_tool_use block that has no matching result block yet. This typically happens when Claude calls a server tool and one of your client tools in parallel: the API returns so you can run the client tools first. Nothing else marks this state. You have to check each server_tool_use or mcp_tool_use id for a matching result block, and a call without one completes in a later response.
A summarization service parses the response text as JSON. For very long transcripts, the JSON string is occasionally cut off mid-object and the parser raises a decode error. Investigating, the engineer finds stop_reason is "max_tokens" on every one of the failing calls. What should the defensive parsing logic do?
Correct answer: A — Catch the decode error, check stop_reason for max_tokens, and treat the response as incomplete
- A. Correct. stop_reason "max_tokens" is the documented signal that the response was cut off before completion; defensive code should branch on it and handle the response as incomplete rather than as malformed JSON.
- B. Incorrect. A lenient parser might mask trailing-comma typos, but it cannot recover data that was never generated because the response was truncated mid-object.
- C. Incorrect. Lowering max_tokens makes truncation more likely, not less, worsening the exact failure being investigated.
- D. Incorrect. The failures are explained by output length, not network transience; blindly resending the same request risks the same truncation again.
Sources1
3.Format consistency when you rely on prompting
If a request does not use constrained decoding, the format depends on how clearly you ask for it. The consistency guide suggests several techniques. Spell out the output format precisely as JSON, XML or a custom template, naming the keys and their allowed values. For example, "sentiment" (positive/negative/neutral). Prefill the assistant turn with the start of the structure you want. Give an example format for Claude to follow, as the market-intelligence prompt does with a full <competitor> template before it says "Now, analyze AcmeGiant and AcmeDataCo using this format."
The prompting best-practices guide backs up the examples technique: make examples relevant to your use case and varied enough to cover edge cases, and wrap them in <example> tags so Claude can tell them apart from instructions. None of these techniques gives the guarantee constrained decoding does, so their output still has to be parsed and validated before anything downstream relies on it.
4.Valid is not the same as true
Schema enforcement guarantees structure, not accuracy. A response that is well-formed, correctly typed and complete can still contain a hallucinated fact. The hallucination guide says outright that even the most advanced models can produce text that is factually wrong or inconsistent with the context they were given. A confident tone and a clean schema tell you nothing about correctness. Output that matters needs a check that looks at its content.
| Technique | How it works |
|---|---|
| Allow "I don't know" | Explicitly give Claude permission to admit uncertainty, for example by saying it lacks enough information |
| Direct quotes first | For long documents (>20k tokens), extract word-for-word quotes, then base the analysis only on them |
| Verify with citations | Have each claim backed by a supporting quote; a claim without one is retracted |
| Chain-of-thought verification | Ask for step-by-step reasoning before the final answer to expose faulty logic |
| Best-of-N verification | Run the same prompt several times; inconsistencies between outputs can indicate hallucinations |
| Iterative refinement | Feed outputs into follow-up prompts that verify or expand on earlier statements |
| External knowledge restriction | Instruct Claude to use only the provided documents, not its general knowledge |
These techniques also produce things a parser can check. If the prompt permits "I don't have enough information to confidently assess this", that exact string can be detected and routed. Cited claims can be checked against their quotes. Empty [] brackets where the model removed an unsupported claim can be counted. Disagreement across N runs is a measurable signal. So being skeptical can be built into the pipeline as checks, not left to a reviewer's judgement.
A weather tool occasionally fails because an upstream API times out. The integration currently returns a bare tool_result of "error" with is_error: true, and Claude keeps repeating the same failing call in a loop. What change to the tool_result content would most likely stop the repeated failing calls?
Correct answer: A — Replace the generic "error" text with a specific message describing the failure and a retry delay
- A. Correct. Instructive error messages that name the failure and suggest a concrete next step give Claude the context needed to adapt, such as waiting before retrying, instead of repeating the identical call.
- B. Incorrect. Hiding the error would cause Claude to treat failed data as valid, corrupting downstream output rather than stopping the loop.
- C. Incorrect. Omitting a required tool_result block after a tool_use block violates the API's formatting requirements and produces a request error, not a resolved loop.
- D. Incorrect. Converting a custom client tool into a server tool isn't a configuration option available to arbitrary user-defined tools, and it wouldn't address the message-quality problem.
Sources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A response with stop_reason end_turn is always complete and safe to parse.Why is that wrong?
Claude can return an empty end_turn response with no content, usually after tool results. Check for content as well as the stop reason.
Covered in Truncated, empty and unfinished responses
2.The fix for an empty response is to send the same conversation, including the empty turn, again.Why is that wrong?
Claude has already decided it is done, so resending changes nothing. Fix the message structure, and as a last resort add a continuation prompt in a new user message.
Covered in Truncated, empty and unfinished responses
3.Once structured outputs guarantee a valid schema, the content can be trusted without further checks.Why is that wrong?
Schema compliance covers structure only. Claude can still produce factually incorrect or inconsistent text, so important claims need grounding techniques such as citations or best-of-N comparison.
Covered in Valid is not the same as true
Practise it for real
See stop_reason signal a truncated response, then fix it by changing the request.
1.Call client.messages.create with model "claude-opus-5-5", max_tokens=10 and the user message "Explain quantum physics".
Why: A 10-token limit is too small for this question, so the response will almost certainly be cut off.
You should see: A successful response with no error raised, containing a few words of text.
2.Print response.stop_reason, and branch on it before using response.content.
Why: Truncation shows up only in this field. The text itself gives no sign that it is incomplete.
You should see: stop_reason is "max_tokens".
3.Send the same request with max_tokens=1024 and check stop_reason again.
Why: The documented fix for max_tokens is to raise the limit or continue the response.
You should see: stop_reason is "end_turn" and the explanation is complete.
Stuck? Get a nudge
Only an "end_turn" response should reach your parser. Route every other value to its own handling path.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Check this field to decide whether to use the response as-is, continue the conversation, retry, or fall back to another model.”
↩︎ Read stop_reason before you read the content“Unlike errors, which indicate failures in processing your request, stop_reason tells you why Claude completed its response generation.”
↩︎ Read stop_reason before you read the content“you'll need to retry the request with a higher max_tokens value to get the full tool use.”
↩︎ Truncated, empty and unfinished responses“Never add text blocks immediately after tool results: This teaches Claude to expect user input after every tool use.”
↩︎ Truncated, empty and unfinished responses“There is no other marker for the state; detect it by checking each server_tool_use or mcp_tool_use block's id for a matching result block.”
↩︎ Truncated, empty and unfinished responses“Sometimes Claude returns an empty response (exactly 2–3 tokens with no content) with stop_reason: "end_turn".”
↩︎ Exam trap 1“Don't retry empty responses without modification: Sending the empty response back won't help.”
↩︎ Exam trap 2 - 2.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/increase-consistencyOfficial docs
“Precisely define your desired output format using JSON, XML, or custom templates so that Claude follows every output formatting element you require.”
↩︎ Format consistency when you rely on prompting“Prefill the Assistant turn with your desired format. This trick bypasses Claude's friendly preamble and enforces your structure.”
↩︎ Format consistency when you rely on prompting - 3.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practicesOfficial docs
“Examples are one of the most reliable ways to steer Claude's output format, tone, and structure.”
↩︎ Format consistency when you rely on prompting - 4.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-hallucinationsOfficial docs
“Explicitly give Claude permission to admit uncertainty. This simple technique can drastically reduce false information.”
↩︎ Valid is not the same as true“If it can't find a quote, it must retract the claim.”
↩︎ Valid is not the same as true“Inconsistencies across outputs could indicate hallucinations.”
↩︎ Valid is not the same as true“Even the most advanced language models, like Claude, can sometimes generate text that is factually incorrect or inconsistent with the given context.”
↩︎ Exam trap 3