What you will be able to do
- Match an exact error message to its category (server, usage limits, authentication, network, request, tool) before choosing a fix
- Predict when Claude Code retries a failed request by itself and when it gives the failure back to you
- Tell a temporary 429 throttle apart from a spend-cap 429, and read the Claude API HTTP error types
- Diagnose tool-use protocol errors, and agent runs that work locally but fail once deployed
Key concept
Retry-safe versus retry-unsafe failures — Claude Code retries a failed request on its own when nothing has been completed yet, or when resending with a change can succeed. It stops and reports the failure when a retry would repeat side effects such as tool calls, or would fail in exactly the same way.
1.Start from the exact message, not a guess
Operational debugging starts with the exact text of the failure. Claude Code's error reference works as a lookup table: you find the message, and it tells you which section explains it. This matters because similar-looking symptoms belong to very different categories. A 429 can mean you have hit a usage limit. An MCP server asking you to sign in again is an authentication problem, not a broken tool. A corporate proxy's certificate shows up as a network error. If you classify the message first, you avoid fixes that cannot work, such as reinstalling the CLI to cure an expired OAuth token.
| Message | Section |
|---|---|
| API Error: Repeated 529 Overloaded errors | Server errors |
| Request rejected (429) | Usage limits |
| Credit balance is too low | Usage limits |
| Invalid API key | Authentication |
| MCP server "<name>" requires re-authorization (token expired) | Authentication |
| unable to get local issuer certificate | Network |
| Prompt is too long / Input is too long for requested model | Request errors |
| API Error: 400 orphaned tool_result in conversation history | Request errors |
| File is covered by a Read deny rule in your permission settings | Tool errors |
It is listed under Authentication. The credential script on their own machine is the suspect, not the API's availability.
Sources1
2.What Claude Code retries for you, and what it hands back
Many transient failures never need you to act. Claude Code retries server errors, overloaded responses and timeouts, as long as they arrive before any of Claude's response has streamed. It also retries dropped connections that happen before Claude has completed any part of its response. A stalled stream is aborted and re-sent at most once, and that retry sits outside the 10-attempt budget. Some retries change the request instead of resending it unchanged. If the input plus max_tokens exceeds the context limit, Claude Code shrinks max_tokens, and it compacts the conversation when no reduction can fit. When an apiKeyHelper script supplies the credential and the API returns a 401 or 403, Claude Code runs the script again and uses its fresh output.
Other failures come straight back to you on purpose. A TLS certificate validation failure, for example from a TLS-inspecting proxy or a missing NODE_EXTRA_CA_CERTS bundle, is reported on the first attempt, because resending would not fix it. If an error arrives after Claude has finished a block of text or a tool call, Claude Code does not resend the request. Resending could run the same tool calls twice. Instead it keeps what Claude completed, runs any tool calls Claude finished, and continues the turn from their results.
| Failure | Claude Code's behaviour |
|---|---|
| 5xx, overloaded or timeout before any response streamed | Retried automatically |
| Temporary 429 throttle | Retried automatically |
| Gateway spend-limit 429 | Not treated as a throttle |
| Input plus max_tokens exceeds the context limit | Retries with a reduced max_tokens, or compacts if nothing fits |
| AWS or Google Cloud credentials fail to load | Discards cached credentials, retries up to two times, then reports |
| TLS certificate validation failure | Reported on the first attempt |
| Error after a completed text block or tool call | Not re-run; keeps the completed output and continues the turn |
During a long debugging session, Claude Code raises "Autocompact is thrashing: the context refilled to the limit..." immediately after automatic compaction completes, and it stops retrying. The team was asking Claude to read an entire 40,000-line production log file to find the root cause of an outage. What is the most effective way to recover and continue the investigation?
Correct answer: B — Ask Claude to read the log file in smaller chunks, such as a specific line range, or delegate the analysis to a subagent instead.
- A. Incorrect. CLAUDE_CODE_MAX_RETRIES governs retries of failed API requests, not the compaction-thrashing loop. Thrashing happens because the file itself refills the context after each compaction, so retrying compaction more times does not solve the underlying problem.
- B. Correct. Auto-compaction succeeded but the oversized file output immediately refilled the context window several times in a row, so Claude Code stopped retrying to avoid wasted API calls. Reading the file in smaller chunks or moving the work to a subagent's separate context window is the documented recovery path.
- C. Incorrect. Disabling auto-compaction removes the safeguard that prevents the context window from overflowing, so it would make the thrashing problem worse, not better, since the full log would still need to fit in context.
- D. Incorrect. This is more drastic than necessary and discards useful investigation context. The documented recovery steps are to chunk the file, use a targeted /compact, or delegate to a subagent, not to abandon the conversation entirely.
Sources1
3.Reading Claude API HTTP errors
If you run agents directly on the Claude API, the HTTP status and error type tell you where the problem lies. Two details are worth remembering. First, not every 429 is a throttle. The API returns a 429 when an organisation hits a rate limit, reaches its usage tier's monthly spend cap, or reaches a spend limit on the Claude Code workspace. A tier spend-cap 429 has no retry-after header and keeps failing until access resumes, so backing off will not help. Second, a 400 can also mean you reached an organisation or workspace spend limit you set yourself.
| Status | Error type | What to do |
|---|---|---|
| 400 | invalid_request_error | Fix the request format or content; also check spend limits you set |
| 401 | authentication_error | Check whether the API key is malformed, revoked or expired |
| 403 | permission_error | Check organisation access and workspace settings in the Claude Console |
| 413 | request_too_large | Reduce the request size |
| 429 | rate_limit_error | Back off for a throttle; a spend-cap 429 has no retry-after |
| 500 | api_error | Retry with exponential backoff; if it persists, contact support with the request ID |
| 504 | timeout_error | Consider the streaming Messages API for long requests |
| 529 | overloaded_error | The API is temporarily overloaded |
The official SDKs already retry transient failures (connection errors, rate limits, 5xx) with exponential backoff, twice by default, and honour retry-after when it is present. You can change or disable this with max_retries. Watch for one blind spot in logs: on a streaming response, an error can arrive after the API has already returned 200. Treating a 200 status as proof of success will hide these mid-stream failures.
Sources2
4.Diagnosing tool-use failures in your own agent
When an agent built on the Messages API misbehaves around tools, the fault is usually in how your application builds the conversation, not in the model. Protocol errors are specific and fixable. Every tool_use block needs a matching tool_result, and the tool_result blocks must come before any text in the user message. If thinking blocks are altered before being sent back, the request fails with a 400. The fix is to send the whole assistant message back unchanged and then append your tool_result.
| Symptom | Fix |
|---|---|
| tool_use ids were found without tool_result blocks immediately after | Return one tool_result per tool_use; put tool_result blocks before any text |
| Claude calls tool A when you wanted tool B | Sharpen descriptions: say WHEN to use each tool, not only WHAT it does |
| Parameter that doesn't exist in your schema | Add strict: true if the schema is in the supported subset |
| Claude refuses to act on a tool result | Move your instructions out of the tool_result into a user turn |
| String comparison on tool inputs fails with newer models | Parse with json.loads() or JSON.parse() |
The refusal case is easy to misdiagnose as the model being difficult. Claude treats instructions found inside tool results as possibly untrusted third-party content. So if your own orchestration instructions travel inside a tool result, Claude may ask for confirmation instead of acting on them. Keep tool results to data only.
A team's project relies on an internal MCP server for ticket lookups. Running /mcp shows the server status as connected, but Claude reports it has no tools available from that server and cannot look up tickets. What is the correct next step to diagnose the failure?
Correct answer: A — Select Reconnect for the server from /mcp, and if the tool count stays at zero, run claude --debug mcp to see the server's stderr output.
- A. Correct. A server that shows connected but lists zero tools has started but isn't returning a tool list. The documented step is to select Reconnect from /mcp, and if the count stays at zero, run claude --debug mcp to inspect the server's stderr output.
- B. Incorrect. A relative path in command or args typically causes the server to fail to start entirely, which shows as failed in /mcp, not connected with zero tools. That is not the situation described here.
- C. Incorrect. Project-scoped approval does not silently expire; once approved, a server stays enabled. This does not explain a connected server returning no tools.
- D. Incorrect. Safe mode disables all MCP servers for the session, so the server would not show as connected at all, let alone connected with zero tools. This is not a diagnostic step for this symptom.
Sources3
5.When an agent works locally but fails once deployed
The Agent SDK runs the Claude Code CLI as a subprocess, so many production-only failures are really about starting that process. If the SDK cannot find the CLI, check that claude --version works in the same environment the application runs in. A service manager or IDE often runs with a different PATH from your shell. If the binary exists but will not launch inside a container, the usual cause is a bundled binary built for a different architecture or libc, or one that lost its execute permission during the image build. Reinstall the SDK during the image build, or rebuild the image for the target architecture. On Windows, the SDK refuses to run npm's claude.cmd shim at all. Point cli_path at a native claude.exe instead.
Sources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Every 429 is a temporary throttle, so exponential backoff will always get the request through eventually.Why is that wrong?
A 429 can also mean the organisation reached its usage tier's monthly spend cap. That 429 carries no retry-after header and keeps failing until access resumes, so backing off does not help.
Covered in Reading Claude API HTTP errors
2.Claude Code keeps retrying TLS certificate errors, so a misconfigured corporate proxy will eventually work if you wait.Why is that wrong?
Certificate validation failures are reported on the first attempt so you can fix the certificate setup, for example with NODE_EXTRA_CA_CERTS. Only transient TLS conditions such as a handshake timeout are retried.
Covered in What Claude Code retries for you, and what it hands back
3.Putting follow-up instructions inside the tool_result content is the most reliable way to steer the agent's next step.Why is that wrong?
Claude treats instructions inside tool results as possibly untrusted, so it may refuse or ask for confirmation. Send instructions in a user turn after the tool_result and keep the result to data only.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://code.claude.com/docs/en/errorsOfficial docs
“Request rejected (429)”
↩︎ Start from the exact message, not a guess“Re-sending it unchanged would fail the same way, so Claude Code retries with a reduced max_tokens”
↩︎ What Claude Code retries for you, and what it hands back“Claude Code re-runs the script and retries with its fresh output, within the full retry budget.”
↩︎ What Claude Code retries for you, and what it hands back“Temporary 429 throttles, but not a gateway’s spend-limit 429, which isn’t a throttle”
↩︎ What Claude Code retries for you, and what it hands back“Claude Code doesn’t re-run the request, because that could execute the same tool calls twice.”
↩︎ Key concept“Claude Code reports the error on the first attempt, so you can fix the certificate setup right away”
↩︎ Exam trap 2 - 2.https://platform.claude.com/docs/en/api/errorsOfficial docs
“The official SDKs automatically retry transient failures (such as connection errors, rate limits, and 5xx server errors) with exponential backoff, twice by default”
↩︎ Reading Claude API HTTP errors“an error can occur after the API returns a 200 response”
↩︎ Reading Claude API HTTP errors“A tier spend-cap 429 has no retry-after header and keeps failing until access resumes”
↩︎ Exam trap 1 - 3.
“Return one tool_result for every tool_use block in the assistant response. Put tool_result blocks before any text.”
↩︎ Diagnosing tool-use failures in your own agent“Send the entire assistant message back unchanged, then append your tool_result.”
↩︎ Diagnosing tool-use failures in your own agent“Claude is trained to treat instructions inside tool results as potentially untrusted third-party content.”
↩︎ Exam trap 3 - 4.
“Processes you launch outside your shell, such as from an IDE or a service manager, often run with a different PATH.”
↩︎ When an agent works locally but fails once deployed“a binary that doesn’t match the container’s architecture or libc”
↩︎ When an agent works locally but fails once deployed