What you will be able to do
- Describe how an orchestrator-subagent (manager/supervisor) hierarchy is structured, and the limits Claude Managed Agents places on it
- Explain the ways subagents improve task execution: context isolation, parallelism, specialization, tool restriction and model choice
- Weigh the token and coordination costs of multiple agents against a well-equipped single agent
1.Supervisor hierarchies: a lead agent and its subagents
A multi-agent system runs several LLM instances, each with its own conversation context, coordinated through code. The most common topology, and the one Anthropic recommends as a starting point for teams new to multi-agent work, is the orchestrator-subagent pattern. It is a hierarchy: a lead agent creates specialized subagents for specific subtasks and manages them. Other coordination patterns exist, such as agent swarms and message-bus architectures, but the hierarchy has the simplest coordination model.
Anthropic's Research feature shows what this looks like in practice. The lead agent analyzes the query, works out a strategy and creates subagents to explore different aspects in parallel. Each subagent acts as an intelligent filter: it searches repeatedly and returns a condensed result, and the lead agent compiles the final answer. This works because search is a form of compression. Each subagent uses its own context window to explore, then passes back only the most important tokens.
Claude Managed Agents expresses the hierarchy as configuration. A coordinator agent lists, in a multiagent roster, the agents it is allowed to delegate to:
---
name: Engineering Lead
model: claude-opus-5-5
tools:
- type: agent_toolset_20260401
multiagent:
type: coordinator
agents: # paths: ant apply substitutes {type: agent, id, version}
- ./reviewer.md
- ./test-writer.md
---
You coordinate engineering work. Delegate code review to the reviewer agent and test writing to the test agent.| Rule | What the docs say |
|---|---|
| Depth | The coordinator delegates to one level only. Listing an agent that has its own multiagent.agents roster fails validation. |
| Breadth | At most 20 unique agents in multiagent.agents, but the coordinator can call multiple copies of each. |
| Isolation | Each agent runs in its own session thread with its own conversation history. Tools, MCP servers and context are not shared. |
| Shared environment | All agents share the same sandbox, filesystem and vault credentials. |
| Persistence | The coordinator can send a follow-up to an agent it called earlier, and that agent keeps its previous turns. |
| Versioning | The roster is snapshotted when the coordinator is created or updated, so referenced agents stay pinned to those versions. |
The supervisor does the same three things in every version: it breaks the task down, delegates the pieces and combines the results. Workers never see the whole picture, and the coordinator never sees the workers' raw exploration. The docs also list escalation as a pattern that works well: consult a more capable agent or model for a subset of complex subtasks.
2.What subagents add to task execution
The Claude Agent SDK gives you subagents in three ways. You can pass them programmatically in the agents parameter of your query() options, write them as markdown files in .claude/agents/ directories, or rely on the built-in general-purpose subagent, which Claude can call through the Agent tool without you defining anything. Here is the programmatic form:
"code-reviewer": AgentDefinition(
# description tells Claude when to use this subagent
description="Expert code review specialist. Use for quality, security, and maintainability reviews.",
# prompt defines the subagent's behavior and expertise
prompt="""You are a code review specialist with expertise in security, performance, and best practices.
When reviewing code:
- Identify security vulnerabilities
- Check for performance issues
- Verify adherence to coding standards
- Suggest specific improvements
Be thorough but concise in your feedback.""",
# tools restricts what the subagent can do (read-only here)
tools=["Read", "Grep", "Glob"],
# model overrides the default model for this subagent
model="sonnet",
),The most important benefit is context isolation. Each subagent runs in its own conversation. Its intermediate tool calls and results stay inside it, and only its final message goes back to the parent. A research subagent can explore dozens of files, and the parent receives a short summary instead of every file the subagent read. The other benefits come from fields in the definition above.
| Benefit | What it buys you | Field |
|---|---|---|
| Specialized instructions | A tailored system prompt with domain expertise that would be noise in the main agent's instructions | prompt |
| Tool restrictions | Less risk of unintended actions, e.g. a reviewer that can analyze but never modify | tools / disallowedTools |
| Model choice | Route a subtask to a different model, such as 'haiku' | model |
| Bounded runs | Stop after a set number of agentic turns and return output marked as partial | maxTurns |
| Delegation choice | Tells Claude when to use this subagent | description |
Subagents also bring parallelism. Several can run at the same time, so independent subtasks finish in the time of the slowest one rather than the sum of all of them. A code review can run a style checker, a security scanner and a test-coverage checker simultaneously. The Claude Code docs add cost control to the list of benefits: you can send tasks to faster, cheaper models such as Haiku.
Only the subagent's final message. Every file read and tool result stays in the subagent's own context, so the main conversation gets the summary without the thousands of tokens that produced it. You can also restrict the subagent to read-only tools, so the audit cannot change anything.
A content team needs to produce localized marketing copy. Every request goes through the same three steps in the same order: draft the copy, run a fixed compliance check on the draft, then translate the approved draft into the target language. Each step consumes the previous step's output, and the sequence never changes. Which workflow pattern fits this process?
Correct answer: D — Prompt chaining, where each LLM call processes the fixed output of the step directly before it in an unchanging sequence.
- A. Incorrect -- the three steps always run in the same fixed order for every request, so there's no runtime decision for an orchestrator to make about which steps apply.
- B. Incorrect -- evaluator-optimizer implies an iterative critique-and-revise loop until criteria are satisfied, but this process is a single fixed pass through three steps, not a repeated refinement loop.
- C. Incorrect -- there's only one pipeline used for every request; routing would require multiple distinct pipelines chosen by classification, which isn't described here.
- D. Correct -- prompt chaining describes a fixed, unchanging sequence where each call processes the prior step's output, exactly matching draft, then check, then translate.
3.The cost side: when more agents make things worse
Every benefit above has a price. Each extra agent is another possible point of failure, another set of prompts to maintain and another source of unexpected behaviour. Anthropic has seen teams build separate agents for planning, execution, review and iteration, only to lose context at every handoff and spend more tokens coordinating than doing the work. In Anthropic's testing, multi-agent implementations typically used 3 to 10 times more tokens than a single agent for the same task. The overhead comes from duplicated context, coordination messages and summaries at each handoff. Measured against plain chat, agents used about 4 times the tokens and multi-agent systems about 15 times. Some teams spent months building elaborate architectures, then found that better prompting of a single agent got the same results.
Anthropic names three situations where multiple agents consistently beat one. The first is context pollution: one subtask fills the context with material the next subtask does not need. The second is parallelizable work: a larger search space can be covered at once. The third is specialization: an agent with too many tools, often 20 or more, or tools from unrelated domains, starts choosing the wrong one, and adding tools makes existing tasks worse. Outside those three situations, coordination costs usually outweigh the benefits.
Be precise about what parallelism buys. Concurrent subagents finish faster than the same subagents run one after another. Compared with a single agent, though, a multi-agent system often takes longer overall because it does far more total computation. Its main benefit is thoroughness, not speed. The fit also depends on the domain: tasks where all agents need the same context, or with many dependencies between agents, suit multi-agent systems poorly today. Most coding tasks have fewer truly parallel pieces than research does.
A security team reviews every pull request by running several independent prompts against the same diff at the same time -- one focused on injection vulnerabilities, one on authentication issues, one on dependency risks -- and then combines the separate flagged findings into one report. Which pattern is this?
Correct answer: A — Parallelization by sectioning, where independent subtasks examine the same input simultaneously and their separate results are aggregated.
- A. Correct -- sectioning splits the work into independent subtasks that run in parallel over the same input, with results aggregated afterward, matching the three simultaneous, differently-focused reviews.
- B. Incorrect -- voting runs the identical task multiple times for diverse outputs on the same question; here the three prompts each examine a different aspect, not the same question repeated.
- C. Incorrect -- evaluator-optimizer requires an iterative critique loop between two LLMs, but these three reviews run independently and simultaneously with no feedback loop between them.
- D. Incorrect -- the three review categories are fixed and predetermined, not dynamically chosen at runtime by a central orchestrating LLM.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A subagent's tool results flow back into the parent conversation, so delegating a file-heavy task still fills the main context.Why is that wrong?
Everything the subagent does in between stays in its own conversation. Only its final message returns, so the parent gets a concise summary.
Covered in What subagents add to task execution
2.Splitting work across parallel agents makes a job both cheaper and faster than a single agent.Why is that wrong?
Multi-agent systems typically use 3 to 10 times more tokens and often take longer overall. What they buy is broader coverage.
Covered in The cost side: when more agents make things worse
3.A Managed Agents coordinator can delegate to agents that are coordinators themselves, building a hierarchy as deep as you like.Why is that wrong?
Managed Agents allows only one level. Referencing an agent that has its own multiagent.agents roster fails the create or update request with a validation error.
Covered in Supervisor hierarchies: a lead agent and its subagents
Practise it for real
Run an Agent SDK query in which Claude delegates to subagents you defined, then change what one subagent is allowed to do.
1.Save the documented Python example, which calls query() with a code-reviewer and a test-runner in agents and allowed_tools set to Read, Grep, Glob and Agent, and run it in a repository that has an authentication module.
Why: No code in the example picks a subagent. Claude chooses from each subagent's description, which is agent behaviour rather than workflow routing.
You should see: The script prints the final result message. The review work happens inside the subagent's own conversation, and only its final message returns to the main agent.
2.Delete the tools line from the code-reviewer definition and run the script again.
Why: If tools is omitted, the subagent inherits every tool available to subagents, so the read-only restriction is gone.
You should see: code-reviewer is no longer limited to Read, Grep and Glob. Put the line back to restore the guarantee that it can analyze but not modify.
3.Add model="haiku" to the test-runner definition.
Why: Sending a narrow subtask to a cheaper, faster model is one of the documented cost levers for subagents.
You should see: test-runner runs on Haiku, while the main agent and code-reviewer keep their own models.
Stuck? Get a nudge
If Claude never delegates, check that "Agent" is still in allowed_tools. The example auto-approves it together with the read-only tools.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“a hierarchical model where a lead agent spawns and manages specialized subagents for specific subtasks”
↩︎ Supervisor hierarchies: a lead agent and its subagents“multi-agent implementations typically use 3-10x more tokens than single-agent approaches for equivalent tasks”
↩︎ The cost side: when more agents make things worse“The primary benefit of parallelization is thoroughness, not speed.”
↩︎ Exam trap 2 - 2.
“Threads are persistent: the coordinator can send a follow-up to an agent it called earlier, and that agent retains everything from its previous turns.”
↩︎ Supervisor hierarchies: a lead agent and its subagents“Tools, MCP servers, and context are not shared.”
↩︎ Supervisor hierarchies: a lead agent and its subagents“Agents can act in parallel with their own isolated context, which helps improve output quality and can also improve time to completion.”
↩︎ The cost side: when more agents make things worse“The coordinator can only delegate to one level of agents”
↩︎ Exam trap 3 - 3.https://code.claude.com/docs/en/agent-sdk/subagentsOfficial docs
“multiple subagents can run concurrently, so independent subtasks finish in the time of the slowest one rather than the sum of all of them”
↩︎ What subagents add to task execution“intermediate tool calls and results stay inside the subagent; only its final message returns to the parent”
↩︎ Exam trap 1 - 4.https://code.claude.com/docs/en/sub-agentsOfficial docs
“Control costs by routing tasks to faster, cheaper models like Haiku”
↩︎ What subagents add to task execution - 5.
“LLM agents are not yet great at coordinating and delegating to other agents in real time”
↩︎ The cost side: when more agents make things worse