What you will be able to do
- Describe how a coordinator agent sits at the hub of all subagent communication
- Explain what a subagent receives from its coordinator, and what it does not inherit
- Name the coordinator's jobs: decomposing a query, delegating subtasks, and aggregating results
- Justify routing every subagent exchange through the coordinator for observability, error handling and controlled information flow
- Design a coordinator that analyzes a query's complexity and selects which subagents to invoke instead of running the full pipeline every time
- Partition a research scope across subagents by subtopic or source type so that no two subagents duplicate work
- Explain why an overly narrow decomposition leaves a broad research topic incompletely covered
- Implement an iterative refinement loop in which the coordinator checks the synthesis for gaps, re-delegates targeted queries and re-invokes synthesis
Key concept
Hub-and-spoke with isolated context — A coordinator agent owns the plan and is the single point every subagent connects to. Each subagent works in its own separate context, sees only the task the coordinator hands it, and sends back only its result.
1.One coordinator at the hub
In the coordinator-subagent pattern, one agent sits at the centre and the others hang off it. Anthropic's Research system calls this an orchestrator-worker pattern: a lead agent coordinates the process and delegates to specialized subagents that run in parallel. Nothing in that description has workers talking to each other. The lead agent is the hub and each subagent is a spoke.
Claude Managed Agents makes the hub concrete. The coordinator runs in the session's primary thread, and additional threads are spawned at runtime whenever it delegates work. You don't build that plumbing yourself. Declaring a multiagent roster is what turns an agent into a coordinator:
---
name: Engineering Lead
model: claude-opus-5-5
tools:
- type: agent_toolset_20260401
multiagent:
type: coordinator
agents: # paths: ant apply substitutes {type: agent, id, version}
- ./reviewer.md
- ./test-writer.md
---
You coordinate engineering work. Delegate code review to the reviewer agent and test writing to the test agent.Once the roster exists, the platform gives each side a different set of tools, and that split shows where control lives.
| Role | Tools the server provides | What defines it |
|---|---|---|
| Coordinator | create_agent, send_to_agent, wait_for_agents, list_agents | A multiagent roster naming the worker(s) |
| Worker | submit_result, send_to_parent | An ordinary agent: a model, a scoped toolset and a system prompt |
Look at the worker's row. Its only outbound channels are submitting a result and sending to its parent. Starting agents, messaging them and waiting on them all belong to the coordinator, so every conversation between spokes has to pass through the hub.
That is the hub-and-spoke architecture in full. The coordinator manages all inter-subagent communication, because there is no other channel between spokes. It manages error handling, because every result and every failure arrives on its side. And it manages information routing, deciding which subagent's output becomes which other subagent's input. Anthropic's Research team found out what happens without that discipline: their early agents were distracting each other with excessive updates. A hub that owns the routing removes that failure mode, since a subagent can only report upward and the coordinator chooses what travels on.
A coordinator for a market-research assistant splits the broad query 'analyze the competitive landscape for electric vehicle charging networks' into narrow subtasks like 'find charger connector types' and 'list charging speeds,' each assigned to a separate subagent. The synthesized report ends up missing pricing models, regulatory incentives, and major competitors entirely. What went wrong?
Correct answer: D — The coordinator decomposed the query too narrowly, so the subtasks covered only a few facets and left broad areas uncovered
- A. The scenario describes subagents accurately answering their narrow assigned subtasks; the problem is which subtasks were assigned, not their ability to search for information.
- B. Nothing in the scenario indicates a timing or synchronization issue; the subagents' assigned subtasks themselves never covered pricing, incentives, or competitors.
- C. The two subtasks described are distinct, not duplicated, so overlapping effort isn't what caused the missing coverage.
- D. Correct. Overly narrow task decomposition by the coordinator leaves broad research topics incompletely covered, since the chosen subtasks addressed only a slice of the competitive landscape.
2.Subagents start with a clean context
The spokes share no memory with the hub. In the Agent SDK, each subagent runs in its own conversation, which starts fresh unless the subagent is a fork. Managed Agents says the same thing: each agent has its own model, system prompt and tools, and tools, MCP servers and context are not shared between agents.
The practical consequence is that the coordinator's delegation is the whole briefing. In Managed Agents, create_agent takes a bare agent name and a task string. The coordinator also can't see its roster agents' prompts, names or descriptions, so everything it believes about its workers comes from its own system prompt. Nothing on the server keeps that belief in step with the workers' real prompts.
The orchestrator-workers cookbook builds this into its prompts: each worker receives the original task *and* its own specific instructions, for better context. A worker handed only its slice of the problem has no other way to learn what the whole problem is.
Nothing, from the subagent's side. It never saw those twenty turns. It starts from a fresh context holding only the task string, so a constraint the coordinator leaves out of that string doesn't exist for the subagent. The fix is in the delegation: put the constraint in the task.
3.Decompose, delegate, aggregate
With isolation in place, the coordinator's work follows a clear cycle. When a query arrives, the lead agent analyzes it, develops a strategy and spawns subagents to explore different aspects at the same time. The orchestrator-workers cookbook splits this into two phases. In the analysis-and-planning phase, the orchestrator writes structured subtask descriptions. In the execution phase, each worker carries one out. The subtasks aren't fixed in advance: the orchestrator decides at runtime which ones to create, and that is what separates the pattern from plain, pre-defined parallelization.
Aggregation is where isolation pays off. A subagent's intermediate tool calls and results stay inside it, and only its final message returns to the parent. A research-assistant subagent can read dozens of files and hand back a short summary. In the Research system, subagents act as intelligent filters and return their findings for the lead agent to compile into a final answer. The coordinator ends up with compact results from several spokes, and turning them into one answer is its job, not any subagent's.
An architect delegates a task to a subagent using the Agent tool, expecting it to reference a decision the user made three turns earlier in the main conversation about which authentication provider to use. The subagent's response ignores that decision entirely and proposes a different provider. What is the most likely cause, and how should the architect fix it?
Correct answer: C — The subagent's context starts fresh each call, so the decision must be included directly in the Agent tool's prompt text
- A. Subagents don't cache or partially load prior turns; they simply never receive the parent conversation at all, so there is no stale copy to resume.
- B. Tool permissions govern which actions a subagent can take, not whether it receives conversation history, so widening tool access would not surface the missing decision.
- C. Correct. A subagent's context window starts fresh and does not automatically inherit the parent's conversation history, so any needed decisions must be passed explicitly in the Agent tool's prompt string.
- D. A custom system prompt defines the subagent's role and expertise; it doesn't override user decisions, and disabling it would remove needed specialization without fixing the missing context.
4.Route every exchange through the coordinator
Every spoke connects only to the hub, so the hub sees everything. That is the case for keeping the hub-and-spoke architecture intact: one coordinator agent that manages all inter-subagent communication, error handling and information routing, rather than a mesh in which subagents call each other.
Observability. In Managed Agents, the coordinator reports activity in the primary thread, which is the same as the session-level event stream. One stream shows the whole run. The plan-big-execute-small cookbook follows a delegation live through thread_created, thread_message_sent and thread_message_received events, and meters each thread with its own cumulative usage.
Consistent error handling. In the orchestrator-workers cookbook, error handling checks that workers return non-empty responses. That check lives in one place because every worker's result passes through the orchestrator. A subagent that fails, times out or returns nothing is handled by the same code path as every other subagent, and the coordinator decides whether to retry, re-delegate or proceed without it.
Controlled information flow. In the specialist-team cookbook, the coordinator decides the order and the hand-offs without doing any specialist work itself. It starts the researcher and the pricing modeler in parallel. It runs the case-study picker only after the researcher's findings come back, because the picker needs those priorities to judge relevance. The coordinator routes the researcher's output to the picker; nothing else reaches the picker, and the pricing modeler never sees either.
The coordinator would lose the one place where the sequence, the hand-off and any failure are visible. When the findings pass through the coordinator, it decides when the picker runs and what the picker receives, the exchange appears on the session's event stream, and a bad or empty result is caught by the same check as every other result.
5.Decide which subagents to invoke from the query's complexity
The orchestrator-workers cookbook frames the coordinator's first job as analysis. A central LLM analyzes each unique task and dynamically determines the best subtasks to delegate to specialized worker LLMs. The contrast it draws is with hardcoded parallelization that generates the same variations regardless of context. A coordinator that always routes every query through the full pipeline of search, analysis and synthesis subagents is that hardcoded pipeline with extra steps. Task decomposition, delegation and result aggregation should all scale with the query: a simple question gets answered directly or with a single subagent, a broad one gets a fan-out.
Managed Agents lists three shapes delegation takes, and each is a choice the coordinator makes per query. Parallelization fans out independent subtasks. Specialization routes to an agent with a domain-focused prompt and tools. Escalation consults a more capable agent or model for a subset of complex subtasks. Only a subset: the rest the coordinator handles itself or hands to a cheaper worker. In the Agent SDK the built-in general-purpose subagent can be invoked at any time via the Agent tool, so fanning out is a decision made in the moment, not a fixed stage in a pipeline.
Anthropic's Research team learned the cost of skipping the analysis step. Early agents made errors like spawning 50 subagents for simple queries. The multi-agent blog's support example shows the small-scale version of getting it right: the support agent calls the order-lookup agent only when the message needs order information, and otherwise answers in its own context.
The first is a breadth-first query: ten independent directions, each needing a search and a reason. It deserves a fan-out of search subagents partitioned by bank, an analysis pass and a synthesis. The second has one fact and no facets. Routing it through the full pipeline would be the 50-subagents-for-a-simple-query mistake. The coordinator should answer it directly, or at most invoke one search subagent.
6.Partition research scope, and don't cut it too narrow
Once a query is broad enough to fan out, how the coordinator cuts it up decides what the run covers and what it wastes. Managed Agents' parallelization pattern is to fan out independent subtasks (searching multiple sources, analyzing separate files) and have the coordinator synthesize the results. The Research system describes subagents exploring different aspects of the question simultaneously, each with distinct tools, prompts and exploration trajectories.
Two partitions recur. By subtopic: one facet of the question per subagent, as in the blog's decompose_query step below. By source type: the specialist-team cookbook gives the researcher web search, the librarian the case-study library, and the pricing modeler sees only the rules file and the seat count. Either way, each subagent owns a slice nobody else is also working. That is what minimizes duplication: no two subagents spend their budget finding the same thing, and the coordinator receives complementary results instead of overlapping ones. Duplication is not free. Multi-agent systems use about 15× more tokens than chats, so two subagents reading the same sources is the most expensive way to learn nothing new.
async def research_topic(query: str) -> dict:
# Lead agent breaks query into research facets
facets = await lead_agent.decompose_query(query)
# Spawn subagents to research each facet in parallel
tasks = [
research_subagent(facet)
for facet in facets
]
results = await asyncio.gather(*tasks)
# Lead agent synthesizes findings
return await lead_agent.synthesize(results)The risk runs the other way too. A subagent starts from only the task it is handed, so anything the coordinator leaves out of its decomposition is never searched. A coordinator that carves a broad research topic into a few narrow facets gets a thorough answer on those facets and silence on everything between them. The result reads as complete, because every subagent reported back, and is incomplete, because the facet list was. Anthropic's evaluations show what is at stake: multi-agent research excels especially for breadth-first queries that involve pursuing multiple independent directions simultaneously, and the S&P 500 board-members query was solved by decomposing it into tasks for subagents while a single agent failed. The primary benefit of parallelization is thoroughness, not speed, and thoroughness is only as wide as the decomposition that produced it.
7.Iterate: evaluate the synthesis, re-delegate, re-synthesize
Because a decomposition can be too narrow, the coordinator should not treat its first synthesis as final. The Research system's LeadResearcher runs an iterative research process. It creates subagents with specific research tasks, each returns its findings, and the LeadResearcher synthesizes these results and decides whether more research is needed. If so, it can create additional subagents or refine its strategy. Once sufficient information is gathered, the system exits the research loop and passes everything to a CitationAgent. The architecture is a multi-step search that adapts to new findings, not a single pass.
Written as a loop a coordinator prompt can follow:
1. Decompose and delegate the query to search and analysis subagents. 2. Invoke synthesis over their results. 3. Evaluate the synthesis output for gaps: facets with thin evidence, contradictions between subagents, and sub-questions the query implied that no subagent was assigned. 4. Re-delegate with targeted queries that name the gap, to the search and analysis subagents. 5. Re-invoke synthesis and return to step 3 until coverage is sufficient.
Managed Agents makes step 4 cheap. Threads are persistent, so the coordinator can send a follow-up to an agent it called earlier, and that agent retains everything from its previous turns. A targeted query to the search subagent that already read the relevant sources doesn't start from zero; it starts from the context that subagent built up, which the coordinator never had to hold.
The loop needs a defined exit. Anthropic's simulations caught agents continuing when they already had sufficient results, and the multi-agent blog describes teams that spent more tokens coordinating than executing. Put the definition of "sufficient" in the coordinator's prompt, as coverage criteria it checks in step 3, rather than leaving it to the coordinator's appetite for one more round.
In step 3 it compares the synthesis against the query and finds the gap: no coverage of the US, Japan, Korea, India or any other jurisdiction. In step 4 it re-delegates with targeted queries, one search subagent per missing region rather than a repeat of the original broad task, and asks the analysis subagent to compare the new findings with the EU and China material. In step 5 it re-invokes synthesis with the enlarged result set, then checks coverage again. It exits when the report answers "worldwide" to the standard the prompt set, not when the subagents stop finding things.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A subagent can see the coordinator's conversation so far, so a short delegation prompt is enough.Why is that wrong?
A subagent starts from a fresh context. It doesn't inherit the coordinator's history, so every constraint and piece of background it needs must be in the task the coordinator passes to it.
Covered in Subagents start with a clean context
2.Letting subagents message each other directly is a harmless shortcut that saves a hop.Why is that wrong?
In the coordinator pattern, workers report upward. Managed Agents gives workers only submit_result and send_to_parent, and keeps agent creation, messaging and waiting on the coordinator's side.
Covered in One coordinator at the hub
3.A careful coordinator routes every query through the full pipeline of subagents so nothing is missed.Why is that wrong?
The coordinator is supposed to analyze the query and decide at runtime which subagents, if any, to invoke. Running the whole roster on a simple question is the failure Anthropic's early research agents made, and multi-agent runs cost many times a single agent's tokens.
Covered in Decide which subagents to invoke from the query's complexity
4.Once the synthesis subagent has produced a report, the research is done and the coordinator returns it.Why is that wrong?
The lead agent's job includes judging the synthesis. It decides whether more research is needed and, if so, creates additional subagents or refines its strategy before synthesizing again. The loop exits on sufficient coverage, not on the first report.
Covered in Iterate: evaluate the synthesis, re-delegate, re-synthesize
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“an orchestrator-worker pattern, where a lead agent coordinates the process while delegating to specialized subagents that operate in parallel”
↩︎ One coordinator at the hub“distracting each other with excessive updates”
↩︎ One coordinator at the hub“the lead agent analyzes it, develops a strategy, and spawns subagents to explore different aspects simultaneously”
↩︎ Decompose, delegate, aggregate“the subagents act as intelligent filters”
↩︎ Decompose, delegate, aggregate“Early agents made errors like spawning 50 subagents for simple queries”
↩︎ Decide which subagents to invoke from the query's complexity“exploring different aspects of the question simultaneously before condensing the most important tokens for the lead research agent”
↩︎ Partition research scope, and don't cut it too narrow“distinct tools, prompts, and exploration trajectories—which reduces path dependency”
↩︎ Partition research scope, and don't cut it too narrow“multi-agent systems use about 15× more tokens than chats”
↩︎ Partition research scope, and don't cut it too narrow“multi-agent research systems excel especially for breadth-first queries that involve pursuing multiple independent directions simultaneously”
↩︎ Partition research scope, and don't cut it too narrow“the multi-agent system found the correct answers by decomposing this into tasks for subagents”
↩︎ Partition research scope, and don't cut it too narrow“The LeadResearcher synthesizes these results and decides whether more research is needed”
↩︎ Iterate: evaluate the synthesis, re-delegate, re-synthesize“if so, it can create additional subagents or refine its strategy.”
↩︎ Iterate: evaluate the synthesis, re-delegate, re-synthesize“Once sufficient information is gathered, the system exits the research loop”
↩︎ Iterate: evaluate the synthesis, re-delegate, re-synthesize“multi-step search that dynamically finds relevant information, adapts to new findings”
↩︎ Iterate: evaluate the synthesis, re-delegate, re-synthesize“agents continuing when they already had sufficient results”
↩︎ Iterate: evaluate the synthesis, re-delegate, re-synthesize“Early agents made errors like spawning 50 subagents for simple queries”
↩︎ Exam trap 3“The LeadResearcher synthesizes these results and decides whether more research is needed”
↩︎ Exam trap 4 - 2.
“additional threads are spawned at runtime when the coordinator delegates work.”
↩︎ One coordinator at the hub“The coordinator can only delegate to one level of agents”
↩︎ One coordinator at the hub“A maximum of 20 unique agents can be listed in multiagent.agents, but the coordinator can call multiple copies of each agent.”
↩︎ One coordinator at the hub“Tools, MCP servers, and context are not shared.”
↩︎ Subagents start with a clean context“have the coordinator synthesize the results.”
↩︎ Decompose, delegate, aggregate“The coordinator reports activity in the primary thread (which is the same as the session-level event stream)”
↩︎ Route every exchange through the coordinator“Escalation: Consult a more capable agent or model for a subset of complex subtasks.”
↩︎ Decide which subagents to invoke from the query's complexity“Fan out independent subtasks simultaneously (searching multiple sources, analyzing separate files) and have the coordinator synthesize the results.”
↩︎ Partition research scope, and don't cut it too narrow“Threads are persistent: the coordinator can send a follow-up to an agent it called earlier, and that agent retains everything from its previous turns.”
↩︎ Iterate: evaluate the synthesis, re-delegate, re-synthesize“each agent runs in its own session thread, a context-isolated event stream with its own conversation history.”
↩︎ Key concept - 3.
“the server automatically gives it create_agent, send_to_agent, wait_for_agents, and list_agents”
↩︎ One coordinator at the hub“Its create_agent tool takes a bare agent name and task string.”
↩︎ Subagents start with a clean context“Everything the coordinator believes about its workers comes from its own system prompt”
↩︎ Subagents start with a clean context“thread_created, thread_message_sent, thread_message_received”
↩︎ Route every exchange through the coordinator“workers get submit_result and send_to_parent the same way.”
↩︎ Exam trap 2 - 4.https://code.claude.com/docs/en/agent-sdk/subagentsOfficial docs
“each subagent runs in its own conversation, which starts fresh unless the subagent is a fork.”
↩︎ Subagents start with a clean context“intermediate tool calls and results stay inside the subagent; only its final message returns to the parent.”
↩︎ Decompose, delegate, aggregate“Claude can invoke the built-in general-purpose subagent at any time via the Agent tool without you defining anything”
↩︎ Decide which subagents to invoke from the query's complexity“each subagent runs in its own conversation, which starts fresh unless the subagent is a fork.”
↩︎ Exam trap 1 - 5.
“Workers receive both the original task AND their specific instructions for better context”
↩︎ Subagents start with a clean context“The orchestrator decides at runtime what subtasks to create”
↩︎ Decompose, delegate, aggregate“Error handling validates that workers return non-empty responses”
↩︎ Route every exchange through the coordinator“dynamically determine the best subtasks to delegate to specialized worker LLMs”
↩︎ Decide which subagents to invoke from the query's complexity“hardcoded parallelization that generates the same variations regardless of context”
↩︎ Decide which subagents to invoke from the query's complexity - 6.
“the coordinator gets to decide the order and the hand-offs without doing any of the specialist work itself”
↩︎ Route every exchange through the coordinator“A pricing modeler sees only the rules file and the seat count.”
↩︎ Partition research scope, and don't cut it too narrow - 7.
“if needs_order_info(user_message):”
↩︎ Decide which subagents to invoke from the query's complexity“decomposes a question into independent facets, runs subagents concurrently, then synthesizes the results.”
↩︎ Partition research scope, and don't cut it too narrow“The primary benefit of parallelization is thoroughness, not speed.”
↩︎ Partition research scope, and don't cut it too narrow“spent more tokens coordinating than executing”
↩︎ Iterate: evaluate the synthesis, re-delegate, re-synthesize