What you will be able to do
- Distinguish MCP protocol errors from tool execution errors and route each failure to the right channel
- Explain why tool failures must be returned in the result with isError: true so the model can self-correct
- Classify a tool failure as transient, validation, business, or permission and say what each implies for recovery
- Explain why a uniform 'Operation failed' message leaves the agent unable to choose a recovery path
- Tell a retryable error from a non-retryable one and return structured metadata (errorCategory, isRetryable, a readable description) so the agent does not waste retry attempts
- Return retriable: false with a customer-friendly explanation for a business rule violation so the agent can communicate it appropriately
- Decide which failures a subagent should recover from locally and what it should propagate to the coordinator, including partial results and what was attempted
- Separate an access failure, which needs a retry decision, from a valid empty result, which is a successful query with no matches
Key concept
Tool execution error (isError: true) — A failure that happens inside a tool is sent back as a normal tool result with isError set to true, not as a protocol-level error. Because it arrives as a result, the model can read what went wrong and decide what to do next.
1.Two channels for failure: protocol errors and tool execution errors
An MCP tool call can go wrong in two quite different places. In the first, the request never reaches a working tool. The tool name is unknown, the request is malformed, or the server itself fails. In the second, the tool runs but its work fails: an upstream API is down, an input value is out of range, or a business rule refuses the operation. The MCP specification reports these two cases through two different channels. Protocol errors are standard JSON-RPC errors, returned in the response's error member instead of a result. Tool execution errors come back as an ordinary result with its isError field set to true.
{
"jsonrpc": "2.0",
"id": 3,
"error": {
"code": -32602,
"message": "Unknown tool: invalid_tool_name"
}
}This response has no content array and nothing shaped like a tool result. The schema keeps this channel for a narrow set of conditions: errors in finding the tool, a server that does not support tool calls, and other exceptional conditions. Everything that goes wrong while the tool is doing its job belongs in the other channel. The current revision of the specification sorts failures like this:
| Failure | Channel | Shape of the response |
|---|---|---|
| Unknown tool | Protocol error | JSON-RPC error object |
| Malformed request (fails to satisfy the CallToolRequest schema) | Protocol error | JSON-RPC error object |
| Server error | Protocol error | JSON-RPC error object |
| API failure | Tool execution error | result with isError: true |
| Input validation error (date in wrong format, value out of range) | Tool execution error | result with isError: true |
| Business logic error | Tool execution error | result with isError: true |
Older revisions drew the line in a different place. The 2024-11-05 and 2025-06-18 specifications list "Invalid arguments" as a protocol error, next to unknown tools and server errors. The 2025-11-25 revision narrows the protocol side to "Malformed requests", meaning requests that fail the CallToolRequest schema. The schema defines that request as a tool name plus the arguments to use. Argument *values* that are wrong, such as a date in the wrong format or a number out of range, are now listed as input validation errors reported with isError: true.
An MCP server exposes `update_inventory`. When the request body is missing the required `sku` field, the server currently returns `isError: true` with `errorCategory: "transient"` and `isRetryable: true`. An agent retries the exact same malformed request repeatedly and never succeeds. What is wrong with this error classification?
Correct answer: A — A missing required field is a validation error, not a transient one, so marking it retryable causes the agent to resend an identically malformed request instead of correcting the input
- A. Correct. A missing required field is a validation problem the caller must fix by changing the request, not a condition that resolves with time; labeling it transient and retryable misleads the agent into resending the same broken payload.
- B. Incorrect. Whether an error is a missing-field validation problem doesn't dictate protocol-level vs. tool-result reporting; a well-formed request with invalid field values is exactly the kind of business-logic-adjacent case reported via isError in the tool result.
- C. Incorrect. The field being supplied in some unrelated future request doesn't make this particular failed call retryable; the current request itself needs correction before it can succeed.
- D. Incorrect. The scenario explicitly shows the agent wasting retries on a request that can never succeed unmodified, which demonstrates the classification is causing the exact problem structured metadata is meant to prevent.
2.Why tool failures belong in the result
The schema answers this directly. Tool-originated errors go in the result with isError: true because the model needs to see them. A protocol error is handled by the client. A tool result is what gets placed in front of the model. If the date problem is sent as a protocol error, the model may never learn that the date was the problem, so it cannot fix the date and try again. The specification's own example of a tool execution error shows what the model gets when it is done correctly:
{
"jsonrpc": "2.0",
"id": 4,
"result": {
"content": [
{
"type": "text",
"text": "Invalid departure date: must be in the future. Current date is 08/08/2025."
}
],
"isError": true
}
}Look at what the text includes. It names the field that failed, states the rule ("must be in the future"), and supplies the fact the model needs to fix it (the current date). With that, the model can build a corrected call on its next turn. The Claude API has the same pattern one level down: a tool_result block carries an optional is_error flag, set when tool execution resulted in an error. In the Agent SDK, a custom tool handler returns isError: true so that you write the message Claude reads, instead of letting a raw exception reach it.
3.Four kinds of tool failure
The exam guide sorts tool failures into four categories: transient, validation, business, and permission. The specification's list of tool execution errors covers three of them in its own words. API failures include the transient kind. Input validation errors are the validation kind. Business logic errors are the business kind. The categories matter because each one calls for a different next step from the agent:
| Category | Where it appears in the spec | Will the same call succeed later? |
|---|---|---|
| Transient (timeouts, service unavailability) | API failures | Possibly: the input was fine, the service was not |
| Validation (invalid input) | Input validation errors, e.g. date in wrong format | No, not until the input is corrected |
| Business (policy violations) | Business logic errors | No, the rule does not change on retry |
| Permission | Not listed separately in the tool error list | No, not with the same access |
Transient errors come from things like the rate-limit example in the specification ("API rate limit exceeded") and from timeouts, which the specification says clients should implement for tool calls. Permission errors appear because the specification requires servers to implement proper access controls. When an access check denies a call, that denial is a tool failure the agent has to understand. The specification does not name a permission category in its error list, however. The four-way split comes from the exam guide, not from the protocol.
One outcome that is not a failure at all deserves a place next to the four categories: the valid empty result. A search tool that runs a correct query and finds no matches has succeeded. It is a successful query with no matches, so its result carries the matches it found, which is none, and isError stays unset or false. The schema states that when isError is not set, the call is assumed to have been successful. An access failure is different in kind. The access control the specification requires servers to implement has refused the call, and the tool reports that refusal with isError: true. The agent then has a retry decision to make: try again with different access, ask the user, or give up. It should make no retry decision for an empty result, because nothing failed and there is nothing to retry. A tool that reports "no matches" as an error, or reports an access denial as an empty list, folds two different situations into one shape and steers the agent to the wrong next step. Distinguishing access failures from valid empty results is the tool's job, and the flag is how it does it. The specification's success example shows the shape an empty-but-successful result should share:
{
"jsonrpc": "2.0",
"id": 2,
"result": {
"content": [
{
"type": "text",
"text": "Current weather in New York:\nTemperature: 72°F\nConditions: Partly cloudy"
}
],
"isError": false
}
}4.What a generic "Operation failed" costs
Imagine a tool that sends every failure back as a uniform error response: isError: true with the generic text "Operation failed". It follows the letter of the protocol, since the error is in the result and the model can see it. But the model has nothing to reason with. Claude's documented behaviour depends on the error text. After a tool failure, Claude will incorporate the error into its response to the user. After an invalid request, it retries 2-3 times with corrections before apologizing to the user. A generic message undermines both. The model cannot tell whether to wait and retry, fix an argument, or stop and explain a policy to the user, so it cannot make an appropriate recovery decision. It may retry a policy violation that will never succeed, abandon a call that a short wait would have fixed, or tell the user something meaningless. Uniform error responses do not just lose detail. They prevent the agent from making appropriate recovery decisions at all, because every failure looks the same.
A raw exception is the opposite mistake. A stack trace is specific, but it is written for the tool's developer, not for the agent. This is why the Agent SDK documentation recommends returning isError: true with a message you write yourself. The goal is text that states the category, the cause, and what would change the outcome.
The cost is easiest to see in retries. The difference between a retryable and a non-retryable error is whether the same call can ever succeed later. Claude's corrections-and-retry loop is the right response to a retryable validation error and pure waste on a non-retryable business or permission error. With "Operation failed", the model cannot tell which it has, so the non-retryable case still gets retried, two or three times, each one a tool call and a model turn spent on an outcome that was settled before the first attempt. Returning structured metadata about the error, which the next section shows, is what prevents those wasted retry attempts.
5.Retryable or not: structured metadata that stops wasted retries
The four categories divide along one line that matters most to the agent: is the error retryable or non-retryable? A transient failure, such as the rate-limit example in the specification, is retryable. The same call may succeed after a wait, because nothing about the request was wrong. A validation error is not retryable as sent, though a corrected call may succeed. A business rule violation and a permission error are non-retryable: the policy and the access do not change between attempts. Claude's documented behaviour shows why this line has to be visible. After an invalid request, Claude will retry 2-3 times with corrections before apologizing to the user. That is the right move for a validation error and a wasted retry attempt for a policy violation, three times over, each costing a tool call and a turn. Only the tool knows which case it is in. Returning structured metadata about the error lets the agent skip the retries that cannot succeed.
The MCP tool result has a field for exactly this. Alongside the content array, a result may carry structuredContent, an optional JSON object that represents the structured result of the tool call. The Agent SDK exposes the same field on custom tool handlers: a JSON object holding the result as machine-readable data, returned alongside content. Neither the protocol nor the SDK prescribes the keys, so the shape is yours to design. The exam guide's pattern uses three parts. An errorCategory value of transient, validation, or permission (with business as a fourth for policy violations) tells the agent what kind of thing went wrong. An isRetryable boolean tells it whether trying again can help. A human-readable description in the text content gives the agent, and through it the user, the detail. The isError: true flag on the same result marks the whole thing as a failure, so a client that reads only the flag still handles it correctly.
Business rule violations need the most care, because the agent cannot fix them and neither can a retry. The specification lists business logic errors as tool execution errors, so they travel with isError: true like the rest. The metadata should carry a retriable: false flag (the guide writes isRetryable in the general case and retriable here; the name is yours, the meaning is the same) so the agent does not burn attempts on a call that policy will refuse every time. The description should be a customer-friendly explanation of the rule, not an internal code. Claude incorporates a tool error into its response to the user, as the documentation's own example shows: a connection error becomes "I'm sorry, I was unable to retrieve the current weather because the weather service API is not available." If the tool says "ERR_POLICY_4471", that is roughly what the user hears back. If it says "Refunds over 500 euros need a manager's approval; this order is 620 euros", the agent can communicate appropriately: tell the customer what the rule is and what would satisfy it.
6.Local error recovery in subagents: what to handle, what to propagate
In a multi-agent design, a coordinator delegates work to subagents, and each subagent calls tools. Where should a tool failure be handled? The retryability line answers it. Transient failures belong to local error recovery inside the subagent that made the call. It holds the context, it can wait and retry within a bounded budget, and the coordinator gains nothing from hearing about a timeout that a second attempt fixed. Non-retryable errors, and transient ones that outlasted the local retry budget, are the errors that cannot be resolved locally, and those are what get propagated to the coordinator. Anything else is noise in the coordinator's context.
The Claude platform already applies this layering to its own server tools, and that behaviour is the model for a subagent. When a server tool such as web search hits a network error, Claude transparently handles it and attempts an alternative response, and the caller does not need to handle is_error results for server tools at all. The failure was resolved at the layer closest to it. The same documentation bounds the retry: for an invalid request, Claude retries 2-3 times with corrections before apologizing to the user. A subagent's local recovery should be bounded in the same way, with timeouts on tool calls as the specification asks clients to implement, so that a transient failure cannot become an endless loop.
What to propagate when local recovery fails matters as much as when. A bare "subagent failed" reproduces the generic "Operation failed" problem one level up. The subagent should return three things: the partial results it did obtain, the errors that could not be resolved locally with their category and retryability, and what was attempted, meaning which calls it made and how many retries it spent. A tool result can carry all of this at once. The content array holds the partial results and the readable description, structuredContent holds the machine-readable metadata, and isError: true marks that the work is incomplete. With that, the coordinator can decide whether to reassign the remaining work, ask the user, or proceed with what it has, instead of re-running everything from the start.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A tool that fails while running should raise a JSON-RPC protocol error, so the client clearly knows the call failed.Why is that wrong?
Errors that originate in the tool belong in the result with isError: true. Sent as protocol errors, they are hidden from the model, which then cannot recover.
Covered in Why tool failures belong in the result
2.Any argument that is wrongly formatted or out of range is a JSON-RPC protocol error.Why is that wrong?
In the 2025-11-25 specification, wrong-format and out-of-range values are input validation errors reported with isError: true. Only requests that fail the CallToolRequest schema are protocol errors.
Covered in Two channels for failure: protocol errors and tool execution errors
3.A search tool that finds no matching records should return isError: true so the agent knows nothing came back.Why is that wrong?
A query that runs correctly and matches nothing is a success with empty content. isError stays unset or false; the schema treats an unset flag as a successful call. Marking it as an error makes the agent treat a valid empty result as an access failure that needs a retry decision.
Covered in Four kinds of tool failure
4.When a tool reports a business rule violation, the agent should retry the same call a few times in case it goes through.Why is that wrong?
Claude's documented retry-with-corrections loop is for invalid requests it can fix. A policy violation does not change on retry, so the tool should return retriable: false and a customer-friendly explanation, and the agent should explain the rule rather than retry.
Covered in Retryable or not: structured metadata that stops wasted retries
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Malformed requests (requests that fail to satisfy CallToolRequest schema)”
↩︎ Two channels for failure: protocol errors and tool execution errors“API failures Input validation errors (e.g., date in wrong format, value out of range) Business logic errors”
↩︎ Four kinds of tool failure“Business logic errors”
↩︎ Retryable or not: structured metadata that stops wasted retries“Input validation errors (e.g., date in wrong format, value out of range)”
↩︎ Exam trap 2 - 2.
“Protocol Errors: Standard JSON-RPC errors for issues like: Unknown tools Invalid arguments Server errors”
↩︎ Two channels for failure: protocol errors and tool execution errors“Implement proper access controls”
↩︎ Four kinds of tool failure“Implement timeouts for tool calls”
↩︎ Four kinds of tool failure“Failed to fetch weather data: API rate limit exceeded”
↩︎ Retryable or not: structured metadata that stops wasted retries“Implement timeouts for tool calls”
↩︎ Local error recovery in subagents: what to handle, what to propagate - 3.
“any errors in finding the tool, an error indicating that the server does not support tool calls”
↩︎ Two channels for failure: protocol errors and tool execution errors“Otherwise, the LLM would not be able to see that an error occurred and self-correct.”
↩︎ Why tool failures belong in the result“If not set, this is assumed to be false (the call was successful).”
↩︎ Four kinds of tool failure“Whether the tool call ended in an error.”
↩︎ Four kinds of tool failure“An optional JSON object that represents the structured result of the tool call.”
↩︎ Retryable or not: structured metadata that stops wasted retries“An optional JSON object that represents the structured result of the tool call.”
↩︎ Local error recovery in subagents: what to handle, what to propagate“Any errors that originate from the tool SHOULD be reported inside the result object, with isError set to true”
↩︎ Key concept“Otherwise, the LLM would not be able to see that an error occurred and self-correct.”
↩︎ Exam trap 1“If not set, this is assumed to be false (the call was successful).”
↩︎ Exam trap 3 - 4.
“is_error (optional): Set to true if the tool execution resulted in an error.”
↩︎ Why tool failures belong in the result“Claude will then incorporate this error into its response to the user.”
↩︎ What a generic "Operation failed" costs“If a tool request is invalid or missing parameters, Claude will retry 2-3 times with corrections before apologizing to the user.”
↩︎ What a generic "Operation failed" costs“If a tool request is invalid or missing parameters, Claude will retry 2-3 times with corrections before apologizing to the user.”
↩︎ Retryable or not: structured metadata that stops wasted retries“I'm sorry, I was unable to retrieve the current weather because the weather service API is not available.”
↩︎ Retryable or not: structured metadata that stops wasted retries“Claude will transparently handle these errors and attempt to provide an alternative response or explanation to the user.”
↩︎ Local error recovery in subagents: what to handle, what to propagate“Unlike client tools, you do not need to handle is_error results for server tools.”
↩︎ Local error recovery in subagents: what to handle, what to propagate“If a tool request is invalid or missing parameters, Claude will retry 2-3 times with corrections before apologizing to the user.”
↩︎ Local error recovery in subagents: what to handle, what to propagate“If a tool request is invalid or missing parameters, Claude will retry 2-3 times with corrections before apologizing to the user.”
↩︎ Exam trap 4 - 5.
“Return isError: true to compose the message instead of surfacing the raw exception”
↩︎ What a generic "Operation failed" costs“structuredContent (optional): a JSON object holding the result as machine-readable data, returned alongside content.”
↩︎ Retryable or not: structured metadata that stops wasted retries