What you will be able to do
- Identify a Claude API failure from its HTTP status, error type and SDK exception class
- Decide whether an error should be retried with backoff or needs a change to the request
- Configure and reason about the SDKs' automatic retry behaviour
- Tell a transient 429 rate limit apart from a spend-cap 429 and a spend-limit 400
Key concept
Transient versus request-change errors — Some Claude API errors come from a temporary condition, such as rate limits, overload or internal failures. These clear up if you wait and retry. Others mean the request, key, account or state is wrong, and sending the same request again gives the same result, so you have to fix the cause first.
1.What a Claude API error looks like
To debug a failed call, first read what the API actually returned. Every error arrives as JSON with the same shape: a top-level error object holding a machine-readable type and a human-readable message, plus a request_id that identifies this particular call. The type is the field to branch on. The message is there for people to read.
{
"type": "error",
"error": {
"type": "not_found_error",
"message": "The requested resource could not be found."
},
"request_id": "req_011CSHoEeqs5C35K2UUqR7Fy"
}If you use an official SDK, you don't handle this raw JSON. The SDK raises a typed exception instead, and the class names differ by language. In Python, for example, a 404 surfaces as anthropic.NotFoundError. The Go SDK is the exception: it has a single *anthropic.Error type for every status, and you branch on StatusCode. The documentation also warns that the set of type values can grow over time, so code that string-matches error messages is fragile twice over. Catch the typed classes, and put the most specific ones first.
Sources1
2.Sorting status codes: fix the request, or wait and retry
Once you know the error type, you need a recovery strategy, and the status code mostly tells you which one. Most 4xx errors describe something about your request, key, account or resource state, so they need a change on your side. The 5xx errors, the 504 timeout and 429 rate limits describe conditions on Anthropic's side or in the network, and those can be retried. The table below lays out the codes the reference documents.
| Status | Error type | What it means | Recovery |
|---|---|---|---|
| 400 | invalid_request_error | Problem with the format or content of the request; also returned when you reach a spend limit you set | Change the request (or raise the spend limit) |
| 401 | authentication_error | API key malformed, revoked or expired | Fix the credentials |
| 402 | billing_error | Billing or payment information issue | Check payment details in the Claude Console |
| 403 | permission_error | Key lacks permission for the resource | Check organization access and workspace settings |
| 404 | not_found_error | Resource not found | Check the endpoint path and resource IDs |
| 409 | conflict_error | Conflicts with the current state of a resource, e.g. concurrent modification | Resolve the conflict, then retry |
| 413 | request_too_large | Request exceeds the maximum allowed bytes | Bring the request within the endpoint's size limit |
| 429 | rate_limit_error | Rate limit hit, tier spend cap reached, or Claude Code workspace spend limit | Depends on which of the three it is (see below) |
| 500 | api_error | Unexpected error internal to Anthropic's systems | Retry with exponential backoff; contact support with the request ID if it persists |
| 504 | timeout_error | Request timed out while processing | Consider the streaming Messages API for long requests |
| 529 | overloaded_error | API temporarily overloaded | Transient; retry |
Two rows need a closer look. A 413 never reaches the model. On the direct Claude API, Cloudflare rejects the request before it gets to the API servers. The limit depends on the endpoint: 32 MB for the Messages and Token Counting APIs, 256 MB for Batch and 500 MB for Files. Retrying the same payload cannot work. A 409 is a special case among the 4xx codes. The request itself was valid, but it collided with another write, such as two workers modifying the same resource. The documented fix is to resolve the conflict and then retry.
A production job calling the Messages API begins failing intermittently with HTTP 429 responses. Each failed response includes a `retry-after` header. The team's current retry code catches the exception and immediately resubmits the same request in a tight loop. What is the most effective change to make to the retry logic?
Correct answer: A — Parse the retry-after header and wait at least that many seconds before resubmitting, backing off further on repeated 429s
- A. Correct. rate_limit_error responses carry a retry-after header specifying the wait time; honoring it (and increasing backoff on repeated failures) is the documented handling for 429s and avoids hammering an already-throttled endpoint.
- B. Batch processing is a separate asynchronous workflow with its own rate limits and queue semantics; switching APIs does not fix a retry loop that ignores retry-after on the existing synchronous call.
- C. max_tokens affects output token consumption and OTPM accounting, not the request-per-minute or input-token limits that commonly trigger 429s, so this does not address the retry loop's core defect.
- D. Rate limits are enforced at the organization level, not per API key, so switching keys does not grant a separate quota and does not fix the missing backoff logic.
Sources1
3.What the SDK already retries for you
Usually it doesn't. The official SDKs retry transient failures on their own. That covers connection errors, rate limits and 5xx server errors. They use exponential backoff, retry twice by default, and honour the retry-after header when the response includes one. Your handler only sees the error once those retries have run out. You can raise, lower or turn off this behaviour with max_retries on the client. Turning it off makes sense when an outer job queue already handles retries and you don't want the two layers compounding.
Timeouts are the next thing to check. The SDKs refuse to send a non-streaming Messages request that is expected to run past a 10-minute timeout. The reference also warns that some networks drop idle connections, so a long non-streaming call with a large max_tokens can fail with no response at all. For a 504, the documented fix is to change how you call the API rather than to retry harder: use the streaming Messages API, or the Message Batches API, which lets you poll for results.
Sources1
4.Three things that look like a spend or rate problem
Rate limits deserve their own section, because 429 is the one status where "just retry" is sometimes right and sometimes wrong. Messages API rate limits are measured per model class in requests per minute (RPM), input tokens per minute (ITPM) and output tokens per minute (OTPM). Going over any of them returns a 429 that says which limit you hit and includes a retry-after header telling you how long to wait. That case is transient and the SDK's retries handle it. Two details affect how you tune for it. For most models, only uncached input tokens count toward ITPM, so prompt caching raises your effective throughput. And max_tokens has no effect on OTPM, which counts only the tokens actually generated.
Reaching your usage tier's monthly spend cap also returns a 429 with type rate_limit_error, but this one has no retry-after header, and every retry fails until access resumes. That includes the SDK's automatic ones. You tell it apart by the error details.
{
"type": "error",
"error": {
"type": "rate_limit_error",
"message": "You have reached your API usage limits: your organization has crossed its monthly API usage threshold, set based on your organization's API tier. You will regain access on 2026-09-01 at 00:00 UTC.",
"details": { "error_code": "enforced_spend_limit_reached" }
},
"request_id": "req_018EeWyXxfu5pfWkrYcMdjWG"
}A 400 invalid_request_error, not a 429. The message begins "You have reached your specified API usage limits" and says when access resumes. To restore access sooner, raise or remove the limit. The one exception is the Claude Code workspace, whose limits are checked separately and can return a 429 with a retry-after header.
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Every 429 is a temporary rate limit that the SDK's backoff will clear.Why is that wrong?
A 429 caused by the tier spend cap carries no retry-after header and keeps failing until access resumes. Check error.details.error_code for enforced_spend_limit_reached.
Covered in Three things that look like a spend or rate problem
2.A 413 request_too_large is a transient server problem worth retrying.Why is that wrong?
The request is larger than the endpoint allows, and on the direct API it is rejected before it reaches the API servers. Only a smaller request, or an endpoint with a larger limit, succeeds.
Covered in Sorting status codes: fix the request, or wait and retry
3.Parsing the error message string is a reliable way to classify API errors.Why is that wrong?
The SDKs raise typed exceptions and the documented practice is to catch those classes, most specific first. The set of type values can also grow over time.
Covered in What a Claude API error looks like
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/api/errorsOfficial docs
“with a top-level error object that always includes a type and message value”
↩︎ What a Claude API error looks like“Catch the SDK's typed classes rather than string-matching error messages, handling the most specific classes first.”
↩︎ What a Claude API error looks like“Retry the request with exponential backoff; if the error persists, contact support with the request ID.”
↩︎ Sorting status codes: fix the request, or wait and retry“Resolve the conflict, then retry the request.”
↩︎ Sorting status codes: fix the request, or wait and retry“On the direct Claude API, Cloudflare returns this error before the request reaches the API servers.”
↩︎ Sorting status codes: fix the request, or wait and retry“with exponential backoff, twice by default, honoring the retry-after header when present”
↩︎ What the SDK already retries for you“The SDK client accepts max_retries to configure or disable this behavior.”
↩︎ What the SDK already retries for you“Consider using the streaming Messages API for long-running requests.”
↩︎ What the SDK already retries for you“The official SDKs automatically retry transient failures (such as connection errors, rate limits, and 5xx server errors)”
↩︎ Key concept“A tier spend-cap 429 has no retry-after header and keeps failing until access resumes”
↩︎ Exam trap 1“On the direct Claude API, Cloudflare returns this error before the request reaches the API servers.”
↩︎ Exam trap 2“Catch the SDK's typed classes rather than string-matching error messages, handling the most specific classes first.”
↩︎ Exam trap 3 - 2.https://platform.claude.com/docs/en/api/rate-limitsOfficial docs
“you will get a 429 error describing which rate limit was exceeded, along with a retry-after header indicating how long to wait”
↩︎ Three things that look like a spend or rate problem“The error type is rate_limit_error, the same as for a rate limit, but the response has no retry-after header.”
↩︎ Three things that look like a spend or rate problem“When usage reaches a spend limit you set, requests return HTTP 400 with error type invalid_request_error.”
↩︎ Three things that look like a spend or rate problem“For most Claude models, only uncached input tokens count toward your ITPM rate limits.”
↩︎ Three things that look like a spend or rate problem