CertSafari
    CLAUDE-CERTIFIED-ARCHITECT-FOUNDATIONS-CCAR-F · Lessons

    Domain 5 · Lesson 27/30

    Error Propagation in Multi-Agent Systems

    Implement error propagation strategies across multi-agent systems

    14 min read
    2.5% of exam
    6 sources
    Published 29 Sep 2026
    Docs as of 27 Sep 2026

    What you will be able to do

    • Tell an access failure (timeout, refused connection) apart from a valid empty result, and explain why the coordinator needs to know which one happened
    • Design a subagent error return that includes the failure type, what was attempted, partial results and possible alternatives
    • Decide which failures a subagent should recover from locally and which it should propagate to the coordinator
    • Recognise silent suppression and whole-workflow termination as the two opposite anti-patterns
    • Explain why silently suppressing errors (returning empty results as success) and terminating an entire workflow on a single failure both come from failing to report access failures separately from valid empty results
    • Structure a synthesis with coverage annotations that separate well-supported findings from gaps caused by unavailable sources

    Key concept

    Structured error propagation — A coordinator knows only what a subagent's final report tells it. So when a subagent cannot finish its task, that report has to say what kind of failure happened, what was tried, what was recovered and what else could be tried. Then the coordinator can decide how to recover instead of guessing.

    1.The coordinator only knows what the subagent tells it

    In a coordinator–subagent design, each subagent works in its own context. It runs searches, reads pages, retries calls and hits errors, and the coordinator sees none of that. The coordinator gets back one thing: the subagent's final message. This isolation is why multi-agent systems scale, because the raw work never fills the coordinator's context. It also creates a reliability problem. If a failure isn't described in that final message, then as far as the coordinator is concerned, it never happened.

    So error propagation is a design decision about what the final message contains. The coordinator is the only agent that sees the whole task. It knows which other subagents are running, which sources overlap and how much the missing piece matters. It is the right place to make recovery decisions, but only if the subagents give it enough to decide with. The rest of this lesson covers what 'enough' means.

    Sources1

    2.Access failures are not empty results

    The two outcomes look alike on the surface but mean opposite things. A valid empty result is a successful query that matched nothing. That is real information: the discontinued SKU really has no rows, and no case law really matches the fact pattern. The coordinator can treat it as a finding and move on. An access failure (a timeout, an exhausted connection pool, a refused connection) means nothing is known about the answer. The data might exist. The right next step is a retry decision: try again, try another route, or record a gap.

    If the subagent reports both as 'no results,' the coordinator will treat an unknown as a negative. It won't retry, and the final answer will state as fact something nobody checked. A success status does not prove that useful work happened, and a system that only watches for error statuses will miss this. Anthropic's own guidance on refusals makes the same point: a response can arrive with a success code and still carry no usable result, so it has to be flagged as a signal of its own.

    This is why silently suppressing errors (returning empty results as success) is an anti-pattern. The coordinator receives a clean success, so it never makes the retry decision the access failure needed, and an unchecked gap becomes a false 'nothing found.' The opposite reaction is an anti-pattern too: terminating the entire workflow on a single failure because one query timed out. That throws away every query that did succeed. Both anti-patterns come from the same missing distinction. When the subagent reports access failures separately from valid empty results, the coordinator can keep what succeeded, retry or route around what failed, and mark the rest as a gap. Claude Code works this way: after a mid-stream server error it keeps what completed and carries on instead of failing the whole turn. The 'Two opposite anti-patterns' section below looks at each one in more depth.

    A regulatory-filings subagent cannot reach its data provider because the provider's API key expired. Instead of surfacing the failure, the subagent returns an empty findings list with status 'success' so the pipeline continues cleanly. The coordinator later synthesizes a report claiming full regulatory coverage. What went wrong?

    In a research pipeline, a coordinator dispatches five subagents to gather sources on different aspects of a market analysis. One subagent's web-fetch tool raises an unrecoverable connection error. The coordinator's current implementation aborts the entire pipeline and discards the four completed subagent results. What is the better design?

    Sources23

    3.What a useful error return contains

    A generic status like 'search unavailable' is honest, but it gives the coordinator nothing to act on. It doesn't say whether the problem was a timeout or a permissions error, which query was attempted, whether any data came back before the failure, or whether another route exists. The coordinator's only options become 'give up' or 'blindly retry the same thing.' A structured error return fixes this with four pieces of context:

    1. Failure type: access failure (timeout, connection refused) versus a failure the retry won't fix (bad query, permission denied). 2. What was attempted: the exact query, source and time window, so the coordinator doesn't repeat the same doomed call. 3. Partial results: whatever did come back, such as 3 of 4 pods' logs or the first pages of results, so the work isn't thrown away. 4. Potential alternatives: a read replica, a different index, a narrower query. The subagent is the one that knows these routes exist.

    The idea of keeping partial work is not special to agents. Anthropic's streaming guidance starts recovery from an interrupted response by saving what already arrived. Anthropic's guidance on tool design likewise recommends error responses that tell the agent something specific it can act on, rather than an opaque code. A subagent's report to its coordinator is that same kind of error response, one level up.

    The Agent SDK's result handling already separates outcomes by subtype: a validated success, a specific failure, and a run that ended without output. A coordinator should get the same kind of distinction from a subagent.python
    if isinstance(message, ResultMessage):
                    if message.subtype == "success" and message.structured_output:
                        # Use the validated output
                        print(message.structured_output)
                    elif message.subtype == "error_max_structured_output_retries":
                        print("Could not produce valid output")
                    else:
                        print("Run ended without a structured output")

    A document-search subagent hits an authentication error against a knowledge base and simply returns the string 'search unavailable' to the coordinator. The coordinator has no other information to act on. What is the primary problem with this design, and what should replace it?

    Sources45

    4.Recover locally, propagate what you cannot resolve

    Structured errors don't mean every hiccup goes to the coordinator. A brief timeout or a transient connection drop is usually best handled where it happened. The subagent has the context to retry, and bouncing each blip upward would clutter the coordinator's context and slow the whole team down. The rule: handle transient failures locally with a bounded number of retries. Propagate only what you can't resolve, and include what was attempted and any partial results.

    Claude Code's own error handling follows this pattern, which makes it a useful reference. It retries some failures itself, reports others right away because retrying can't help, and when a failure arrives partway through a response, it keeps the work that already finished instead of throwing it away.

    How Claude Code handles different failures: local retry, immediate report, or keeping partial work
    FailureWhat Claude Code doesPattern it illustrates
    Input plus max_tokens exceeds the context limitRetries with a reduced max_tokens; stops retrying and compacts when no reduction can fitDon't resend an unchanged request that will fail the same way
    Expired or missing cloud credentialRetries up to two times, then reports the errorBounded local recovery, then propagate
    TLS certificate validation failureReports the error on the first attemptPropagate at once when a retry cannot help
    Server error after a completed block or tool callKeeps what Claude completed and continues the turn from finished tool callsPreserve partial results instead of discarding them

    Sources32

    5.Two opposite anti-patterns

    Silent suppression returns an empty result marked as success when the subagent actually failed. It feels safe because the pipeline keeps running, but the failure has been turned into a false negative. The coordinator can't retry what it doesn't know failed, and the final output says 'nothing found' when nothing was actually checked.

    Whole-workflow termination goes the other way: one subagent's failure aborts the entire run. Three sources that answered fine get discarded because a fourth timed out. That throws away good work and leaves the user with nothing, when a partial answer with a clearly marked gap would have been useful. Claude Code's documentation describes a fix for exactly this: earlier versions discarded the partial output and reported the whole turn as an error, and newer versions keep what completed and carry on.

    The right behaviour sits between the two. The failure is made visible, so it isn't suppressed. The work continues with what succeeded, so it isn't terminated. The coordinator decides whether the gap matters enough to retry, reroute or flag.

    Sources3

    6.Coverage annotations in the final synthesis

    Error propagation doesn't stop at the coordinator. When the coordinator writes the final synthesis, the gaps its subagents reported need to show up there too. Otherwise the user reads a confident answer and can't tell that one source was unreachable. The fix is a coverage annotation: the synthesis marks which findings are well-supported (several sources, or an authoritative one) and which topic areas have gaps because a source was unavailable, noting that the gap comes from a failure, not from evidence that nothing exists.

    This is the same honesty rule Anthropic suggests for agents reporting on their own progress: claim only what the evidence supports, and say explicitly what isn't verified. For example, a research synthesis might say: 'Market-size figures are well-supported (three independent sources agree). Regulatory status in the EU is a gap: the legal database timed out after retries, so absence of findings here is not evidence of absence.'

    Sources6

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A timeout and a query that matched zero rows can both be reported as 'no results found', because either way nothing came back.Why is that wrong?

      They call for different decisions. A zero-match query is a real finding. A timeout means the answer is unknown and the coordinator must decide whether to retry or reroute. Retry logic only works if the failure type is known, because some requests fail the same way when resent unchanged and others are worth retrying.

      Covered in Access failures are not empty results

    2. 2.When a source times out, returning an empty result with a success status is the safe choice, because the pipeline keeps running and nothing breaks.Why is that wrong?

      That is silent suppression. A success status hides the failure, so the coordinator never retries and reports an unchecked gap as 'nothing found'. A result that arrives as a success but carries nothing usable has to be flagged as its own signal. Aborting the whole workflow instead is the opposite anti-pattern. The fix is to report the access failure explicitly and keep the results that succeeded.

      Covered in Access failures are not empty results

    3. 3.A short, generic error such as 'search unavailable' is the cleanest thing to send the coordinator, because it keeps the coordinator's context small.Why is that wrong?

      A generic status hides the failure type, what was attempted, the partial results and any alternatives, which is exactly what the coordinator needs to recover. Errors should communicate specific, actionable information, not opaque codes.

      Covered in What a useful error return contains

    4. 4.For full transparency, a subagent should pass every transient error up to the coordinator as soon as it happens.Why is that wrong?

      Transient failures should be retried locally, with a limit, and propagated only if they persist, together with what was attempted. Claude Code itself retries some failures a bounded number of times before reporting them.

      Covered in Recover locally, propagate what you cannot resolve

    5. 5.If any subagent fails, the safest move is to fail the whole workflow so no incomplete answer reaches the user.Why is that wrong?

      Aborting everything throws away the work that succeeded. The better pattern keeps completed work, continues, and marks the gap. Claude Code's documentation records moving away from discarding everything on a mid-stream error.

      Covered in Two opposite anti-patterns

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “intermediate tool calls and results stay inside the subagent; only its final message returns to the parent”
      ↩︎ The coordinator only knows what the subagent tells it
      “intermediate tool calls and results stay inside the subagent; only its final message returns to the parent”
      ↩︎ Key concept
    2. 2.
      “A refusal is an HTTP 200, so monitoring built on error rates or 5xx responses never sees it.”
      ↩︎ Access failures are not empty results
      “Budget retries per request, not per turn or per session.”
      ↩︎ Recover locally, propagate what you cannot resolve
      “A refusal is an HTTP 200, so monitoring built on error rates or 5xx responses never sees it.”
      ↩︎ Exam trap 2
    3. 3.
      “Re-sending it unchanged would fail the same way”
      ↩︎ Access failures are not empty results
      “It keeps what Claude completed, runs any tool calls Claude finished, and continues the turn from their results.”
      ↩︎ Access failures are not empty results
      “retries up to two times, then reports the error so you can re-authenticate right away”
      ↩︎ Recover locally, propagate what you cannot resolve
      “Before v2.1.199, Claude Code discarded the partial output and reported the whole turn as an error when a server error arrived mid-stream.”
      ↩︎ Two opposite anti-patterns
      “It keeps what Claude completed, runs any tool calls Claude finished, and continues the turn from their results.”
      ↩︎ Two opposite anti-patterns
      “Re-sending it unchanged would fail the same way”
      ↩︎ Exam trap 1
      “retries up to two times, then reports the error so you can re-authenticate right away”
      ↩︎ Exam trap 4
      “Before v2.1.199, Claude Code discarded the partial output and reported the whole turn as an error when a server error arrived mid-stream.”
      ↩︎ Exam trap 5
    4. 4.
      “Capture the partial response: Save all content that was successfully received before the error occurred.”
      ↩︎ What a useful error return contains
    5. 5.
      “prompt-engineer your error responses to clearly communicate specific and actionable improvements, rather than opaque error codes or tracebacks”
      ↩︎ What a useful error return contains
      “prompt-engineer your error responses to clearly communicate specific and actionable improvements, rather than opaque error codes or tracebacks”
      ↩︎ Exam trap 3
    6. 6.
      “Only report work you can point to evidence for; if something is not yet verified, say so explicitly.”
      ↩︎ Coverage annotations in the final synthesis