What you will be able to do
- Tell an access failure (timeout, refused connection) apart from a valid empty result, and explain why the coordinator needs to know which one happened
- Design a subagent error return that includes the failure type, what was attempted, partial results and possible alternatives
- Decide which failures a subagent should recover from locally and which it should propagate to the coordinator
- Recognise silent suppression and whole-workflow termination as the two opposite anti-patterns
- Explain why silently suppressing errors (returning empty results as success) and terminating an entire workflow on a single failure both come from failing to report access failures separately from valid empty results
- Structure a synthesis with coverage annotations that separate well-supported findings from gaps caused by unavailable sources
Key concept
Structured error propagation — A coordinator knows only what a subagent's final report tells it. So when a subagent cannot finish its task, that report has to say what kind of failure happened, what was tried, what was recovered and what else could be tried. Then the coordinator can decide how to recover instead of guessing.
1.The coordinator only knows what the subagent tells it
In a coordinator–subagent design, each subagent works in its own context. It runs searches, reads pages, retries calls and hits errors, and the coordinator sees none of that. The coordinator gets back one thing: the subagent's final message. This isolation is why multi-agent systems scale, because the raw work never fills the coordinator's context. It also creates a reliability problem. If a failure isn't described in that final message, then as far as the coordinator is concerned, it never happened.
So error propagation is a design decision about what the final message contains. The coordinator is the only agent that sees the whole task. It knows which other subagents are running, which sources overlap and how much the missing piece matters. It is the right place to make recovery decisions, but only if the subagents give it enough to decide with. The rest of this lesson covers what 'enough' means.
Sources1
2.Access failures are not empty results
The two outcomes look alike on the surface but mean opposite things. A valid empty result is a successful query that matched nothing. That is real information: the discontinued SKU really has no rows, and no case law really matches the fact pattern. The coordinator can treat it as a finding and move on. An access failure (a timeout, an exhausted connection pool, a refused connection) means nothing is known about the answer. The data might exist. The right next step is a retry decision: try again, try another route, or record a gap.
If the subagent reports both as 'no results,' the coordinator will treat an unknown as a negative. It won't retry, and the final answer will state as fact something nobody checked. A success status does not prove that useful work happened, and a system that only watches for error statuses will miss this. Anthropic's own guidance on refusals makes the same point: a response can arrive with a success code and still carry no usable result, so it has to be flagged as a signal of its own.
This is why silently suppressing errors (returning empty results as success) is an anti-pattern. The coordinator receives a clean success, so it never makes the retry decision the access failure needed, and an unchecked gap becomes a false 'nothing found.' The opposite reaction is an anti-pattern too: terminating the entire workflow on a single failure because one query timed out. That throws away every query that did succeed. Both anti-patterns come from the same missing distinction. When the subagent reports access failures separately from valid empty results, the coordinator can keep what succeeded, retry or route around what failed, and mark the rest as a gap. Claude Code works this way: after a mid-stream server error it keeps what completed and carries on instead of failing the whole turn. The 'Two opposite anti-patterns' section below looks at each one in more depth.
A regulatory-filings subagent cannot reach its data provider because the provider's API key expired. Instead of surfacing the failure, the subagent returns an empty findings list with status 'success' so the pipeline continues cleanly. The coordinator later synthesizes a report claiming full regulatory coverage. What went wrong?
Correct answer: A — The subagent silently suppressed an access failure as success, so the coordinator reported false full coverage instead of a real gap
- A. Correct. Reporting a genuine access failure as an empty success is a classic anti-pattern: it hides the failure from the coordinator, which then has no way to know its synthesized output has an undisclosed coverage gap.
- B. An expired API key is not self-resolving and is exactly the kind of failure that must be surfaced, since the coordinator cannot distinguish it from a legitimately empty result once it is disguised as success.
- C. Requiring manual verification of every subagent result defeats the purpose of automated multi-agent coordination and does not address the actual defect, which is the mislabeled status.
- D. Terminating the whole workflow over one subagent's failure is the opposite anti-pattern; the fix is accurate error reporting, not an all-or-nothing shutdown.
In a research pipeline, a coordinator dispatches five subagents to gather sources on different aspects of a market analysis. One subagent's web-fetch tool raises an unrecoverable connection error. The coordinator's current implementation aborts the entire pipeline and discards the four completed subagent results. What is the better design?
Correct answer: A — Synthesize from the four completed results and use the failed subagent's error context to flag the coverage gap
- A. Correct. Terminating the whole workflow on one subagent failure discards useful completed work; the coordinator should use the partial results it has and rely on the failed subagent's error context to annotate what's missing.
- B. Aborting on a single subagent failure is the anti-pattern being illustrated, since it throws away four subagents' worth of valid, already-completed work.
- C. Re-running the four successful subagents wastes time and cost for no benefit, since their results are already valid and do not need to be regenerated.
- D. Fabricating content to hide a coverage gap is worse than aborting, because it presents unsupported information as if it were verified research.
3.What a useful error return contains
A generic status like 'search unavailable' is honest, but it gives the coordinator nothing to act on. It doesn't say whether the problem was a timeout or a permissions error, which query was attempted, whether any data came back before the failure, or whether another route exists. The coordinator's only options become 'give up' or 'blindly retry the same thing.' A structured error return fixes this with four pieces of context:
1. Failure type: access failure (timeout, connection refused) versus a failure the retry won't fix (bad query, permission denied). 2. What was attempted: the exact query, source and time window, so the coordinator doesn't repeat the same doomed call. 3. Partial results: whatever did come back, such as 3 of 4 pods' logs or the first pages of results, so the work isn't thrown away. 4. Potential alternatives: a read replica, a different index, a narrower query. The subagent is the one that knows these routes exist.
The idea of keeping partial work is not special to agents. Anthropic's streaming guidance starts recovery from an interrupted response by saving what already arrived. Anthropic's guidance on tool design likewise recommends error responses that tell the agent something specific it can act on, rather than an opaque code. A subagent's report to its coordinator is that same kind of error response, one level up.
if isinstance(message, ResultMessage):
if message.subtype == "success" and message.structured_output:
# Use the validated output
print(message.structured_output)
elif message.subtype == "error_max_structured_output_retries":
print("Could not produce valid output")
else:
print("Run ended without a structured output")A document-search subagent hits an authentication error against a knowledge base and simply returns the string 'search unavailable' to the coordinator. The coordinator has no other information to act on. What is the primary problem with this design, and what should replace it?
Correct answer: A — The generic status hides recovery detail; the subagent should return failure type, what was queried, and any partial results
- A. Correct. A bare status string like 'search unavailable' strips away exactly the information (why it failed, what was attempted, what alternatives exist) that lets a coordinator decide between retrying, rerouting, or proceeding with partial coverage.
- B. Hiding failure detail from the coordinator is the problem being described, not an acceptable design; the coordinator needs that context to make recovery decisions.
- C. A boolean flag carries even less information than the current string and would make the coordinator's recovery decision harder, not easier.
- D. Silently retrying without ever surfacing the failure risks masking a persistent auth problem that the subagent cannot resolve on its own, leaving the coordinator unaware anything is wrong.
4.Recover locally, propagate what you cannot resolve
Structured errors don't mean every hiccup goes to the coordinator. A brief timeout or a transient connection drop is usually best handled where it happened. The subagent has the context to retry, and bouncing each blip upward would clutter the coordinator's context and slow the whole team down. The rule: handle transient failures locally with a bounded number of retries. Propagate only what you can't resolve, and include what was attempted and any partial results.
Claude Code's own error handling follows this pattern, which makes it a useful reference. It retries some failures itself, reports others right away because retrying can't help, and when a failure arrives partway through a response, it keeps the work that already finished instead of throwing it away.
| Failure | What Claude Code does | Pattern it illustrates |
|---|---|---|
| Input plus max_tokens exceeds the context limit | Retries with a reduced max_tokens; stops retrying and compacts when no reduction can fit | Don't resend an unchanged request that will fail the same way |
| Expired or missing cloud credential | Retries up to two times, then reports the error | Bounded local recovery, then propagate |
| TLS certificate validation failure | Reports the error on the first attempt | Propagate at once when a retry cannot help |
| Server error after a completed block or tool call | Keeps what Claude completed and continues the turn from finished tool calls | Preserve partial results instead of discarding them |
It should include the failure type (access failure: timeout, still happening after N retries), the exact query and source it attempted, any partial results gathered before or between failures, and any alternative routes it knows of (a replica, a different index, a narrower query). That lets the coordinator choose between rerouting, re-delegating, or recording a gap, without repeating the retries the subagent already did.
5.Two opposite anti-patterns
Silent suppression returns an empty result marked as success when the subagent actually failed. It feels safe because the pipeline keeps running, but the failure has been turned into a false negative. The coordinator can't retry what it doesn't know failed, and the final output says 'nothing found' when nothing was actually checked.
Whole-workflow termination goes the other way: one subagent's failure aborts the entire run. Three sources that answered fine get discarded because a fourth timed out. That throws away good work and leaves the user with nothing, when a partial answer with a clearly marked gap would have been useful. Claude Code's documentation describes a fix for exactly this: earlier versions discarded the partial output and reported the whole turn as an error, and newer versions keep what completed and carry on.
The right behaviour sits between the two. The failure is made visible, so it isn't suppressed. The work continues with what succeeded, so it isn't terminated. The coordinator decides whether the gap matters enough to retry, reroute or flag.
Sources3
6.Coverage annotations in the final synthesis
Error propagation doesn't stop at the coordinator. When the coordinator writes the final synthesis, the gaps its subagents reported need to show up there too. Otherwise the user reads a confident answer and can't tell that one source was unreachable. The fix is a coverage annotation: the synthesis marks which findings are well-supported (several sources, or an authoritative one) and which topic areas have gaps because a source was unavailable, noting that the gap comes from a failure, not from evidence that nothing exists.
This is the same honesty rule Anthropic suggests for agents reporting on their own progress: claim only what the evidence supports, and say explicitly what isn't verified. For example, a research synthesis might say: 'Market-size figures are well-supported (three independent sources agree). Regulatory status in the EU is a gap: the legal database timed out after retries, so absence of findings here is not evidence of absence.'
Sources6
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A timeout and a query that matched zero rows can both be reported as 'no results found', because either way nothing came back.Why is that wrong?
They call for different decisions. A zero-match query is a real finding. A timeout means the answer is unknown and the coordinator must decide whether to retry or reroute. Retry logic only works if the failure type is known, because some requests fail the same way when resent unchanged and others are worth retrying.
Covered in Access failures are not empty results
2.When a source times out, returning an empty result with a success status is the safe choice, because the pipeline keeps running and nothing breaks.Why is that wrong?
That is silent suppression. A success status hides the failure, so the coordinator never retries and reports an unchecked gap as 'nothing found'. A result that arrives as a success but carries nothing usable has to be flagged as its own signal. Aborting the whole workflow instead is the opposite anti-pattern. The fix is to report the access failure explicitly and keep the results that succeeded.
Covered in Access failures are not empty results
3.A short, generic error such as 'search unavailable' is the cleanest thing to send the coordinator, because it keeps the coordinator's context small.Why is that wrong?
A generic status hides the failure type, what was attempted, the partial results and any alternatives, which is exactly what the coordinator needs to recover. Errors should communicate specific, actionable information, not opaque codes.
Covered in What a useful error return contains
4.For full transparency, a subagent should pass every transient error up to the coordinator as soon as it happens.Why is that wrong?
Transient failures should be retried locally, with a limit, and propagated only if they persist, together with what was attempted. Claude Code itself retries some failures a bounded number of times before reporting them.
Covered in Recover locally, propagate what you cannot resolve
5.If any subagent fails, the safest move is to fail the whole workflow so no incomplete answer reaches the user.Why is that wrong?
Aborting everything throws away the work that succeeded. The better pattern keeps completed work, continues, and marks the gap. Claude Code's documentation records moving away from discarding everything on a mid-stream error.
Covered in Two opposite anti-patterns
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://code.claude.com/docs/en/agent-sdk/subagentsOfficial docs
“intermediate tool calls and results stay inside the subagent; only its final message returns to the parent”
↩︎ The coordinator only knows what the subagent tells it“intermediate tool calls and results stay inside the subagent; only its final message returns to the parent”
↩︎ Key concept - 2.
“A refusal is an HTTP 200, so monitoring built on error rates or 5xx responses never sees it.”
↩︎ Access failures are not empty results“Budget retries per request, not per turn or per session.”
↩︎ Recover locally, propagate what you cannot resolve“A refusal is an HTTP 200, so monitoring built on error rates or 5xx responses never sees it.”
↩︎ Exam trap 2 - 3.https://code.claude.com/docs/en/errorsOfficial docs
“Re-sending it unchanged would fail the same way”
↩︎ Access failures are not empty results“It keeps what Claude completed, runs any tool calls Claude finished, and continues the turn from their results.”
↩︎ Access failures are not empty results“retries up to two times, then reports the error so you can re-authenticate right away”
↩︎ Recover locally, propagate what you cannot resolve“Before v2.1.199, Claude Code discarded the partial output and reported the whole turn as an error when a server error arrived mid-stream.”
↩︎ Two opposite anti-patterns“It keeps what Claude completed, runs any tool calls Claude finished, and continues the turn from their results.”
↩︎ Two opposite anti-patterns“Re-sending it unchanged would fail the same way”
↩︎ Exam trap 1“retries up to two times, then reports the error so you can re-authenticate right away”
↩︎ Exam trap 4“Before v2.1.199, Claude Code discarded the partial output and reported the whole turn as an error when a server error arrived mid-stream.”
↩︎ Exam trap 5 - 4.
“Capture the partial response: Save all content that was successfully received before the error occurred.”
↩︎ What a useful error return contains - 5.https://www.anthropic.com/engineering/writing-tools-for-agentsSecondary source
“prompt-engineer your error responses to clearly communicate specific and actionable improvements, rather than opaque error codes or tracebacks”
↩︎ What a useful error return contains“prompt-engineer your error responses to clearly communicate specific and actionable improvements, rather than opaque error codes or tracebacks”
↩︎ Exam trap 3 - 6.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5Official docs
“Only report work you can point to evidence for; if something is not yet verified, say so explicitly.”
↩︎ Coverage annotations in the final synthesis