What you will be able to do
- Recognise when subtasks cannot be predicted and the split must be decided at runtime
- Use context isolation to decompose tasks that would otherwise pollute one agent's context
- Weigh the token and latency cost of multi-agent decomposition against the value of the task
1.When the split has to be decided at runtime
Some problems cannot be cut in advance. In a coding change, for example, how many files need editing and what each edit looks like depend on the request itself. For cases like this, the orchestrator-workers pattern hands the decomposition to a model. A central LLM breaks the task down on the fly, delegates the pieces to worker LLMs, and synthesises what they return. It resembles parallelization in shape, but the orchestrator creates the subtasks at runtime to fit each input.
Anthropic's cookbook runs this in two phases. In the analysis and planning phase, the orchestrator reads the task and writes structured subtask descriptions in XML. In the execution phase, each worker receives the original task for context along with its own subtask type and description. Giving workers both is a deliberate design decision. Anthropic's Research feature applies the same approach at scale. Research is open-ended and path-dependent, so a fixed pipeline cannot handle it. A lead agent works out a strategy and spawns subagents to cover different aspects in parallel.
async def research_topic(query: str) -> dict:
# Lead agent breaks query into research facets
facets = await lead_agent.decompose_query(query)
# Spawn subagents to research each facet in parallel
tasks = [
research_subagent(facet)
for facet in facets
]
results = await asyncio.gather(*tasks)
# Lead agent synthesizes findings
return await lead_agent.synthesize(results)A model-driven split can go wrong in ways a hard-coded one cannot. Anthropic's early research agents spawned 50 subagents for simple queries and distracted each other with excessive updates, and they fixed this mainly through prompt engineering. In Claude Managed Agents, the coordinator cannot inspect the workers on its roster. What it knows about them comes only from its own system prompt, so that description has to stay in agreement with the workers' real prompts.
A support-automation pipeline is decomposed into two subtasks: classifying a high volume of incoming tickets, and drafting detailed technical remediation plans for the subset that need engineering follow-up. How should model selection be applied across these decomposed subtasks?
Correct answer: A — Assign Claude Haiku 4.5 to the high-volume classification subtask and Claude Opus 4.8 to the remediation subtask, matching model to complexity
- A. Correct. Decomposition allows each subtask to use the model best suited to it: a fast, economical model for high-volume classification, and a more capable model for complex remediation reasoning, per Anthropic's model selection guidance.
- B. Incorrect. Lowering effort on Opus reduces reasoning depth uniformly but still runs the larger, costlier model for simple classification, missing the cost and latency benefit of matching model to subtask.
- C. Incorrect. A smaller model cannot be reliably compensated for on complex reasoning subtasks purely through longer prompts; capability gaps for nuanced remediation work require a more capable model.
- D. Incorrect. Random model assignment ignores the actual complexity differences between subtasks and produces inconsistent quality and unpredictable cost.
2.Decomposing to keep context clean
Decomposition also manages what each model sees. In Managed Agents, all agents share the sandbox, filesystem and vault credentials, but each one runs in its own session thread. Threads persist, so the coordinator can send a follow-up to an earlier agent and that agent still has its previous turns. Tools, MCP servers and context are not shared between agents.
That isolation fixes context pollution. Pollution happens when material from one subtask sits in the context while the agent works on a different subtask. A support agent that pulls thousands of tokens of order history reasons worse about the technical fault it is supposed to diagnose. The fix is to hand the lookup to a separate agent that reads the full history and returns only what the main agent needs.
class SupportAgent:
def handle_issue(self, user_message: str):
if needs_order_info(user_message):
order_id = extract_order_id(user_message)
# Get only what's needed, not full history
order_summary = OrderLookupAgent().lookup_order(order_id)
# Inject compact summary, not full context
context = f"Order {order_id}: {order_summary['status']}, purchased {order_summary['date']}"This cut pays off when a subtask produces a lot of context, more than about 1,000 tokens, most of which the main task does not need. It also needs a clear rule for what to extract. Anthropic's research post frames search the same way, as compression: each subagent explores in its own context window and passes back only the most important tokens.
3.Decomposing by difficulty and cost
A task can also be split by how much intelligence each part needs. Managed Agents lists escalation as a pattern: a subset of hard subtasks goes to a more capable agent or model. The reverse split is the cookbook's plan-big, execute-small coordinator. A frontier model plans and synthesises, and cheap workers do the token-heavy reading in their own parallel context windows. In the authors' runs, 84-98% of the team's input tokens were billed at the worker rate. The same economics apply to document review, log analysis and codebase sweeps.
Mixing models can also raise quality. On Anthropic's internal research eval, a system with a Claude Opus 4 lead agent and Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2%. It found every board member of the S&P 500 IT companies by giving each piece to a subagent, where the single agent failed with slow, sequential searches.
4.When not to decompose further
Every extra agent is another point of failure, another set of prompts to maintain, and another source of surprises. Anthropic's blog reports that multi-agent implementations typically use 3-10x more tokens than single-agent approaches for the same task. The research post measures from a different baseline: agents use about 4 times the tokens of a chat interaction, and multi-agent systems about 15 times. Anthropic has seen teams spend months on elaborate multi-agent designs, only to get equivalent results from better prompting on a single agent.
Speed is a subtle point. Managed Agents says isolated parallel agents can improve time to completion. The blog adds that multi-agent systems often take longer overall than a single agent because they do so much more total work, and that the main benefit is thoroughness. Tasks where every agent needs the same context, or where subtasks depend heavily on each other, are poor fits. Most coding tasks offer fewer truly parallel pieces than research does.
| Situation | Signal that decomposition helps |
|---|---|
| Context pollution | A subtask generates high context volume, mostly irrelevant to the main task |
| Parallelizable work | A large search space or many independent angles to investigate |
| Specialization | Too many tools (often 20+), domain confusion, or new tools degrading existing tasks |
A decomposed data-processing workflow requires Claude to invoke dozens of small tools in a tight loop, and issuing each call through the standard conversational tool-use turn is adding significant latency and token overhead. Which technique addresses this?
Correct answer: A — Use programmatic tool calling so Claude invokes tools directly from a code execution container, cutting round trips and token overhead per call
- A. Correct. Programmatic tool calling lets Claude call tools directly from code running in a code execution container, avoiding a full conversational round trip per call and reducing latency and token consumption for tight, multi-tool loops.
- B. Incorrect. Wrapping every small tool call in its own subagent adds orchestration and reporting overhead rather than removing the per-call round-trip cost within a tight loop.
- C. Incorrect. Message Batches is for asynchronous, large-volume request processing, not for reducing round trips within a single live, tightly looped tool-calling workflow.
- D. Incorrect. Raising effort increases reasoning depth per turn but does not change how many separate conversational turns are needed to issue a large number of small tool calls.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Orchestrator-workers is just parallelization with a model in charge, so either works when the subtasks are known in advance.Why is that wrong?
What sets orchestrator-workers apart is that it creates the subtasks at runtime. If the subtasks are predictable, the cookbook recommends simpler parallelization.
Covered in When the split has to be decided at runtime
2.Fanning a task out to parallel subagents makes the whole system finish faster than a single agent.Why is that wrong?
Parallelism shortens the time compared with doing the same work sequentially, but multi-agent systems often take longer overall because they do much more total computation. The main gain is thoroughness.
Covered in When not to decompose further
3.Subagents in Managed Agents share the coordinator's conversation history, so they see everything it has seen.Why is that wrong?
Each agent runs in its own context-isolated session thread. Only the sandbox, filesystem and vault credentials are shared.
Covered in Decomposing to keep context clean
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Multiagent orchestration lets one agent coordinate with others to complete complex work.”
↩︎ When the split has to be decided at runtime“each agent runs in its own session thread, a context-isolated event stream with its own conversation history.”
↩︎ Decomposing to keep context clean“Escalation: Consult a more capable agent or model for a subset of complex subtasks.”
↩︎ Decomposing by difficulty and cost“best suited for complex tasks that either require work across a variety of surfaces, or where multiple well-scoped tasks contribute”
↩︎ When not to decompose further“Tools, MCP servers, and context are not shared.”
↩︎ Exam trap 3 - 2.https://www.anthropic.com/engineering/building-effective-agentsSecondary source
“a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results.”
↩︎ When the split has to be decided at runtime - 3.
“Workers receive both the original task AND their specific instructions for better context”
↩︎ When the split has to be decided at runtime“Subtasks are predictable and can be pre-defined (use simpler parallelization)”
↩︎ When not to decompose further“The orchestrator decides at runtime what subtasks to create, making this more adaptive than pre-defined parallel workflows.”
↩︎ Exam trap 1 - 4.
“Early agents made errors like spawning 50 subagents for simple queries”
↩︎ When the split has to be decided at runtime“The essence of search is compression: distilling insights from a vast corpus.”
↩︎ Decomposing to keep context clean“a multi-agent system with Claude Opus 4 as the lead agent and Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2%”
↩︎ Decomposing by difficulty and cost“most coding tasks involve fewer truly parallelizable tasks than research”
↩︎ When not to decompose further - 5.
“Everything the coordinator believes about its workers comes from its own system prompt”
↩︎ When the split has to be decided at runtime“A frontier model plans the research and synthesizes the answer, but it never touches a raw web page”
↩︎ Decomposing by difficulty and cost - 6.
“Subagents provide isolation, with each operating in its own clean context focused on its specific task.”
↩︎ Decomposing to keep context clean“In our testing, multi-agent implementations typically use 3-10x more tokens than single-agent approaches for equivalent tasks.”
↩︎ When not to decompose further“The primary benefit of parallelization is thoroughness, not speed.”
↩︎ Exam trap 2