CertSafari
    CLAUDE-CERTIFIED-DEVELOPER-FOUNDATIONS-CCDV-F · Lessons

    Domain 4 · Lesson 11/25

    Tracing Claude Failures: Integration Layer or Model Output

    Debugging and Error Handling

    9 min read
    2.6% of exam
    5 sources
    Published 29 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Capture and use request IDs to trace a specific Claude API call
    • Recognise failures that arrive inside a successful 200 response by reading stop_reason
    • Diagnose tool-use and message-structure errors caused by the integration layer
    • Decide whether a failure comes from your integration code or from the model's output

    1.Request IDs: tracing a single call

    Trace analysis needs a way to follow one call from start to finish. Every Claude API response, success or failure, carries a unique request-id header such as req_018EeWyXxfu5pfWkrYcMdjWG. The same value appears as request_id in error bodies. Log it next to your own trace data. When a 500 keeps coming back and you contact support, this ID is what they need.

    Reading the request ID from a successful response in the Python SDKpython
    client = anthropic.Anthropic()
    
    message = client.messages.create(
        model="claude-sonnet-5",
        max_tokens=1024,
        messages=[{"role": "user", "content": "Hello, Claude"}],
    )
    print(f"Request ID: {message._request_id}")

    Python and TypeScript expose _request_id directly on top-level response objects. C#, Go, Java and PHP expose it through their raw-response accessors, and Ruby through middleware. On Claude Platform on AWS you get two IDs: the AWS x-amzn-requestid, which you use for CloudTrail lookups, and the Anthropic request-id, which you use for Anthropic support tickets.

    Sources1

    2.When the failure is inside a 200 response

    Request IDs help you trace calls that failed. The harder failures to debug are the ones where no exception is raised at all. Every successful Messages API response includes stop_reason, and it tells you something different from an error does. An error means the request could not be processed. stop_reason says why Claude stopped generating, and it can report a truncated or declined answer even though the HTTP call succeeded.

    stop_reason values and the action each calls for
    stop_reasonWhen it occursWhat to do
    end_turnClaude finished naturallyUse the response
    max_tokensResponse reached your max_tokens limitRaise max_tokens or continue the response
    stop_sequenceClaude emitted one of your stop_sequencesRead stop_sequence to see which fired
    tool_useClaude is calling a toolRun the tool and return the result
    pause_turnA server-tool loop reached its iteration limitSend the assistant content back to continue
    refusalClaude declined to respondRead stop_details and retry on a fallback model
    model_context_window_exceededResponse filled the model's context windowTreat the response as truncated

    A summary that stops mid-sentence with status 200 is therefore not an SDK bug. Check stop_reason for max_tokens. Truncation matters most in the middle of a tool call. If the last content block is an unfinished tool_use, you have to resend the request with a higher limit.

    Detecting a tool call cut off by max_tokens and retrying with a higher limitpython
    # Check if response was truncated during tool use
    if response.stop_reason == "max_tokens":
        # Check if the last content block is an incomplete tool_use
        last_block = response.content[-1]
        if last_block.type == "tool_use":
            # Send the request with higher max_tokens
            response = client.messages.create(
                model="claude-opus-5-5",
                max_tokens=4096,  # Increased limit
                messages=messages,
                tools=tools,
            )

    An empty end_turn reply of just a few tokens often points back at your own message history. A common cause is adding text blocks right after tool results, which teaches Claude to end its turn and wait for more user input. Sending the empty response back unchanged won't help. Fix the message structure first, and add a continuation prompt in a new user message only as a last resort. Streaming has one more trap: over server-sent events, an error can arrive after the API has already returned 200, and it doesn't follow the standard error mechanisms.

    An agent sends a tool_use request for a `query_database` tool. The client-side tool executor throws a connection exception, and the developer's code returns a tool_result block with `is_error: true` and the exception message as content, then sends this back to Claude in the next turn. Claude's subsequent response apologizes and suggests checking the database connection instead of returning query results. Where does the root cause of this failed turn lie?

    Sources21

    3.400s that your message plumbing causes

    The integration layer is at fault. Several documented invalid_request_error messages come from how your code rebuilds the conversation, not from anything the model did. If you edit, reorder, filter or reconstruct thinking or redacted_thinking blocks in the latest assistant message, you get a 400 that names the position of the offending block. The fix is to send the whole assistant message back unchanged and then append your tool_result. A missing tool_result for any tool_use id also triggers a 400, and so does a tool_result that isn't the first content block. Two more come from the model configuration rather than the loop. Claude 4.6 and later models reject a prefilled final assistant message. Newer models reject thinking.type: "enabled", and the message tells you to switch to adaptive thinking.

    A tool that fails while running, such as a downstream API returning 500, is different. It is not an API error at all. Return the failure to Claude as a tool_result with is_error: true, and Claude will work the error into its reply.

    Reporting a tool execution failure back to Claudejson
    {
      "role": "user",
      "content": [
        {
          "type": "tool_result",
          "tool_use_id": "toolu_01A09q90qw90lq917835lq9",
          "content": "ConnectionError: the weather service API is not available (HTTP 500)",
          "is_error": true
        }
      ]
    }

    Sources314

    4.Isolating the origin: integration layer or model output

    Taken together, these cases give you a sequence of checks. First, did the call return an HTTP error? If so, the status and type tell you whether the fix is your request, your account, or waiting out a transient condition. Second, on a 200, what is stop_reason? Truncation and refusal are signals the model reports on purpose, not bugs. Third, if the output is complete but wrong, is the model misbehaving, or is it reacting to what you sent? The tool-use troubleshooting guide traces many apparent model failures back to the definitions and messages the integration supplied.

    Apparent model misbehaviour traced to its likely cause
    SymptomLikely causeFix
    Claude calls tool A when you wanted tool BDescription ambiguityDifferentiate tools by when to use them, not only what they do
    Parameter that doesn't exist in your schemaModel over-generation without strict modeAdd strict: true if the schema is in the supported subset
    Claude refuses to act on a tool resultYour own instructions are delivered inside the tool_result contentMove instructions to a user turn after the tool_result
    String comparison on tool inputs fails with newer modelsUnicode and forward-slash escaping differs between model versionsParse with json.loads() or JSON.parse()
    Every request is a cache misstool_choice, thinking configuration or effort varying between requestsHold them constant or place the cache_control breakpoint before the variation

    The same reasoning applies when Claude calls a tool with missing required parameters. The handle-tool-calls guide says this usually means Claude didn't have enough information to use the tool correctly, and the recommended fix during development is to write more detailed tool descriptions. You can also return the error as a tool_result, and Claude will retry 2-3 times with corrections.

    A team's usage sits well under their published requests-per-minute limit, yet they begin seeing 429 errors after a marketing campaign causes API traffic to triple within a few minutes. Historical usage was steady before the spike. What is the most likely explanation, and what should the team change to avoid recurrence?

    Sources35

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.An HTTP 200 with no SDK exception means the response is complete and usable.Why is that wrong?

      stop_reason can report truncation (max_tokens, model_context_window_exceeded) or a refusal on a successful response. It is separate from the error mechanism.

      Covered in When the failure is inside a 200 response

    2. 2.An empty end_turn response can be fixed by resending the same conversation.Why is that wrong?

      Claude has already decided its turn is done. Fix the message structure, for example by removing text after tool results, and use a new continuation prompt only as a last resort.

      Covered in When the failure is inside a 200 response

    3. 3.A 400 about thinking blocks after a tool call means the model produced malformed output.Why is that wrong?

      The integration changed the thinking blocks before sending them back. The fix is to resend the assistant message unchanged.

      Covered in 400s that your message plumbing causes

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Every API response includes a unique request-id header.”
      ↩︎ Request IDs: tracing a single call
      “The Python and TypeScript SDKs expose the request ID as a _request_id property on top-level response objects.”
      ↩︎ Request IDs: tracing a single call
      “When contacting support about a specific request, include this ID to help quickly resolve your issue.”
      ↩︎ Request IDs: tracing a single call
      “an error can occur after the API returns a 200 response”
      ↩︎ When the failure is inside a 200 response
      “Claude 4.6 and later models and Claude Mythos Preview do not support prefilling assistant messages.”
      ↩︎ 400s that your message plumbing causes
    2. 2.
      “Unlike errors, which indicate failures in processing your request, stop_reason tells you why Claude completed its response generation.”
      ↩︎ When the failure is inside a 200 response
      “you'll need to retry the request with a higher max_tokens value to get the full tool use”
      ↩︎ When the failure is inside a 200 response
      “Unlike errors, which indicate failures in processing your request, stop_reason tells you why Claude completed its response generation.”
      ↩︎ Exam trap 1
      “Don't retry empty responses without modification: Sending the empty response back won't help.”
      ↩︎ Exam trap 2
    3. 3.
      “Return one tool_result for every tool_use block in the assistant response.”
      ↩︎ 400s that your message plumbing causes
      “your application is altering the assistant's thinking blocks before sending them back.”
      ↩︎ 400s that your message plumbing causes
      “Differentiate tools by WHEN to use them, not only WHAT they do.”
      ↩︎ Isolating the origin: integration layer or model output
      “Parse with json.loads() or JSON.parse(). Never do raw string matching on serialized input.”
      ↩︎ Isolating the origin: integration layer or model output
      “your application is altering the assistant's thinking blocks before sending them back.”
      ↩︎ Exam trap 3
    4. 4.
      “This happens because the model you requested has removed extended thinking”
      ↩︎ 400s that your message plumbing causes
    5. 5.
      “it usually means that there wasn't enough information for Claude to use the tool correctly”
      ↩︎ Isolating the origin: integration layer or model output

    Ready to test yourself?

    Practise the 20 questions on this subdomain.