What you will be able to do
- Capture and use request IDs to trace a specific Claude API call
- Recognise failures that arrive inside a successful 200 response by reading stop_reason
- Diagnose tool-use and message-structure errors caused by the integration layer
- Decide whether a failure comes from your integration code or from the model's output
1.Request IDs: tracing a single call
Trace analysis needs a way to follow one call from start to finish. Every Claude API response, success or failure, carries a unique request-id header such as req_018EeWyXxfu5pfWkrYcMdjWG. The same value appears as request_id in error bodies. Log it next to your own trace data. When a 500 keeps coming back and you contact support, this ID is what they need.
client = anthropic.Anthropic()
message = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Claude"}],
)
print(f"Request ID: {message._request_id}")Python and TypeScript expose _request_id directly on top-level response objects. C#, Go, Java and PHP expose it through their raw-response accessors, and Ruby through middleware. On Claude Platform on AWS you get two IDs: the AWS x-amzn-requestid, which you use for CloudTrail lookups, and the Anthropic request-id, which you use for Anthropic support tickets.
Sources1
2.When the failure is inside a 200 response
Request IDs help you trace calls that failed. The harder failures to debug are the ones where no exception is raised at all. Every successful Messages API response includes stop_reason, and it tells you something different from an error does. An error means the request could not be processed. stop_reason says why Claude stopped generating, and it can report a truncated or declined answer even though the HTTP call succeeded.
| stop_reason | When it occurs | What to do |
|---|---|---|
| end_turn | Claude finished naturally | Use the response |
| max_tokens | Response reached your max_tokens limit | Raise max_tokens or continue the response |
| stop_sequence | Claude emitted one of your stop_sequences | Read stop_sequence to see which fired |
| tool_use | Claude is calling a tool | Run the tool and return the result |
| pause_turn | A server-tool loop reached its iteration limit | Send the assistant content back to continue |
| refusal | Claude declined to respond | Read stop_details and retry on a fallback model |
| model_context_window_exceeded | Response filled the model's context window | Treat the response as truncated |
A summary that stops mid-sentence with status 200 is therefore not an SDK bug. Check stop_reason for max_tokens. Truncation matters most in the middle of a tool call. If the last content block is an unfinished tool_use, you have to resend the request with a higher limit.
# Check if response was truncated during tool use
if response.stop_reason == "max_tokens":
# Check if the last content block is an incomplete tool_use
last_block = response.content[-1]
if last_block.type == "tool_use":
# Send the request with higher max_tokens
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=4096, # Increased limit
messages=messages,
tools=tools,
)An empty end_turn reply of just a few tokens often points back at your own message history. A common cause is adding text blocks right after tool results, which teaches Claude to end its turn and wait for more user input. Sending the empty response back unchanged won't help. Fix the message structure first, and add a continuation prompt in a new user message only as a last resort. Streaming has one more trap: over server-sent events, an error can arrive after the API has already returned 200, and it doesn't follow the standard error mechanisms.
An agent sends a tool_use request for a `query_database` tool. The client-side tool executor throws a connection exception, and the developer's code returns a tool_result block with `is_error: true` and the exception message as content, then sends this back to Claude in the next turn. Claude's subsequent response apologizes and suggests checking the database connection instead of returning query results. Where does the root cause of this failed turn lie?
Correct answer: A — In the integration layer, because the tool executor failed to reach the database; Claude's reply correctly reflects the is_error tool_result it was given
- A. Correct. The failure occurred in the client-side tool executor (the integration layer) before any model reasoning happened; Claude only saw the is_error tool_result the application constructed and responded appropriately to that signal, so its output is not the source of the defect.
- B. Claude does not autonomously retry tool calls on its own initiative when given an error result; the application is responsible for deciding whether and how to retry the failed tool invocation.
- C. is_error is a general-purpose flag for signaling any tool execution failure back to Claude, including connection failures, and is not restricted to cases where the tool name was misspelled.
- D. Acknowledging a reported tool failure is the expected and desired model behavior given an is_error tool_result; no prompt tuning is meant to suppress this signal, since suppressing it would hide real integration failures from the user.
3.400s that your message plumbing causes
The integration layer is at fault. Several documented invalid_request_error messages come from how your code rebuilds the conversation, not from anything the model did. If you edit, reorder, filter or reconstruct thinking or redacted_thinking blocks in the latest assistant message, you get a 400 that names the position of the offending block. The fix is to send the whole assistant message back unchanged and then append your tool_result. A missing tool_result for any tool_use id also triggers a 400, and so does a tool_result that isn't the first content block. Two more come from the model configuration rather than the loop. Claude 4.6 and later models reject a prefilled final assistant message. Newer models reject thinking.type: "enabled", and the message tells you to switch to adaptive thinking.
A tool that fails while running, such as a downstream API returning 500, is different. It is not an API error at all. Return the failure to Claude as a tool_result with is_error: true, and Claude will work the error into its reply.
{
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": "toolu_01A09q90qw90lq917835lq9",
"content": "ConnectionError: the weather service API is not available (HTTP 500)",
"is_error": true
}
]
}4.Isolating the origin: integration layer or model output
Taken together, these cases give you a sequence of checks. First, did the call return an HTTP error? If so, the status and type tell you whether the fix is your request, your account, or waiting out a transient condition. Second, on a 200, what is stop_reason? Truncation and refusal are signals the model reports on purpose, not bugs. Third, if the output is complete but wrong, is the model misbehaving, or is it reacting to what you sent? The tool-use troubleshooting guide traces many apparent model failures back to the definitions and messages the integration supplied.
| Symptom | Likely cause | Fix |
|---|---|---|
| Claude calls tool A when you wanted tool B | Description ambiguity | Differentiate tools by when to use them, not only what they do |
| Parameter that doesn't exist in your schema | Model over-generation without strict mode | Add strict: true if the schema is in the supported subset |
| Claude refuses to act on a tool result | Your own instructions are delivered inside the tool_result content | Move instructions to a user turn after the tool_result |
| String comparison on tool inputs fails with newer models | Unicode and forward-slash escaping differs between model versions | Parse with json.loads() or JSON.parse() |
| Every request is a cache miss | tool_choice, thinking configuration or effort varying between requests | Hold them constant or place the cache_control breakpoint before the variation |
The same reasoning applies when Claude calls a tool with missing required parameters. The handle-tool-calls guide says this usually means Claude didn't have enough information to use the tool correctly, and the recommended fix during development is to write more detailed tool descriptions. You can also return the error as a tool_result, and Claude will retry 2-3 times with corrections.
A team's usage sits well under their published requests-per-minute limit, yet they begin seeing 429 errors after a marketing campaign causes API traffic to triple within a few minutes. Historical usage was steady before the spike. What is the most likely explanation, and what should the team change to avoid recurrence?
Correct answer: A — The spike likely triggered acceleration limits designed to catch sharp usage increases; the team should ramp traffic up gradually and keep usage patterns more consistent
- A. Correct. The documentation notes that a sharp increase in usage can trigger 429 errors from acceleration limits even while under the standard published limits, and the recommended mitigation is to ramp traffic up gradually and maintain consistent usage patterns.
- B. Published RPM limits apply per model to the Messages API itself; the Message Batches API has its own separate rate limits, so this statement misattributes which endpoint the published limit governs.
- C. A 429 rate_limit_error does not indicate key revocation; a revoked or invalid key produces a 401 authentication_error instead, so this diagnosis points to the wrong error type.
- D. Exceeding the maximum request payload size produces a 413 request_too_large error, not a 429, and is unrelated to the number of requests sent per minute.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.An HTTP 200 with no SDK exception means the response is complete and usable.Why is that wrong?
stop_reason can report truncation (max_tokens, model_context_window_exceeded) or a refusal on a successful response. It is separate from the error mechanism.
Covered in When the failure is inside a 200 response
2.An empty end_turn response can be fixed by resending the same conversation.Why is that wrong?
Claude has already decided its turn is done. Fix the message structure, for example by removing text after tool results, and use a new continuation prompt only as a last resort.
Covered in When the failure is inside a 200 response
3.A 400 about thinking blocks after a tool call means the model produced malformed output.Why is that wrong?
The integration changed the thinking blocks before sending them back. The fix is to resend the assistant message unchanged.
Covered in 400s that your message plumbing causes
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/api/errorsOfficial docs
“Every API response includes a unique request-id header.”
↩︎ Request IDs: tracing a single call“The Python and TypeScript SDKs expose the request ID as a _request_id property on top-level response objects.”
↩︎ Request IDs: tracing a single call“When contacting support about a specific request, include this ID to help quickly resolve your issue.”
↩︎ Request IDs: tracing a single call“an error can occur after the API returns a 200 response”
↩︎ When the failure is inside a 200 response“Claude 4.6 and later models and Claude Mythos Preview do not support prefilling assistant messages.”
↩︎ 400s that your message plumbing causes - 2.
“Unlike errors, which indicate failures in processing your request, stop_reason tells you why Claude completed its response generation.”
↩︎ When the failure is inside a 200 response“you'll need to retry the request with a higher max_tokens value to get the full tool use”
↩︎ When the failure is inside a 200 response“Unlike errors, which indicate failures in processing your request, stop_reason tells you why Claude completed its response generation.”
↩︎ Exam trap 1“Don't retry empty responses without modification: Sending the empty response back won't help.”
↩︎ Exam trap 2 - 3.
“Return one tool_result for every tool_use block in the assistant response.”
↩︎ 400s that your message plumbing causes“your application is altering the assistant's thinking blocks before sending them back.”
↩︎ 400s that your message plumbing causes“Differentiate tools by WHEN to use them, not only WHAT they do.”
↩︎ Isolating the origin: integration layer or model output“Parse with json.loads() or JSON.parse(). Never do raw string matching on serialized input.”
↩︎ Isolating the origin: integration layer or model output“your application is altering the assistant's thinking blocks before sending them back.”
↩︎ Exam trap 3 - 4.
“This happens because the model you requested has removed extended thinking”
↩︎ 400s that your message plumbing causes - 5.
“it usually means that there wasn't enough information for Claude to use the tool correctly”
↩︎ Isolating the origin: integration layer or model output