CertSafari
    CCAR-P · Lessons

    Domain 1 · Lesson 4/38

    Multi-Agent Systems: When Orchestrators and Subagents Pay Off

    Design multi-agent systems and orchestration strategies

    9 min read
    2.83% of exam
    3 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Explain what separates a multi-agent system from a single agent, including what Claude Managed Agents shares and isolates between agents
    • Name the three situations where multiple agents reliably beat one agent, and recognise when coordination costs outweigh the gain
    • Weigh a multi-agent design against cost, latency and quality before recommending it
    • Use context isolation and parallel subagents for the problems each one actually solves

    Key concept

    Context isolation — Each agent in a multi-agent system works in its own conversation history. A subagent can take in a large volume of material and hand back only a distilled result, so the coordinating agent's context stays focused on its own task.

    1.What makes a system multi-agent

    A multi-agent system runs several LLM instances, each with its own conversation context, and coordinates them through code. Each agent handles a distinct slice of the task. Several coordination patterns exist, including agent swarms, capability-based systems and message bus architectures. The one Anthropic recommends starting with is orchestrator-subagent: a lead agent spawns specialised subagents for specific subtasks and manages them. It is hierarchical and has a simple coordination model.

    Claude Managed Agents builds this pattern into the platform. One agent coordinates the others, and the others can work in parallel, each in its own isolated context. The docs say this can improve output quality and can also shorten time to completion. The coordinator reports its activity in the primary thread, which is the same as the session-level event stream. New threads are spawned at runtime whenever the coordinator delegates work.

    Sources12

    2.Deciding whether multiple agents are worth it

    The starting point is one agent. Anthropic reports that some teams spent months building elaborate multi-agent architectures, only to find that better prompting of a single agent gave the same results. One design split planning, execution, review and iteration into separate agents. It lost context at every handoff and spent more tokens coordinating than executing. Every added agent is another possible point of failure and another set of prompts to maintain.

    Anthropic names three situations where multiple agents consistently beat one: context pollution is degrading performance, the tasks can run in parallel, or specialisation improves tool selection or task focus. Outside those situations, coordination usually costs more than it returns. The Managed Agents docs describe the same fit: complex tasks that need work across many surfaces, or several well-scoped tasks that each contribute to an overall goal. The research-system post adds where it fits poorly. Domains where every agent needs the same context, or where agents depend heavily on each other, are poor candidates today. Most coding tasks offer fewer truly parallelisable pieces than research does.

    How a multi-agent design changes the business trade-offs, compared with a single agent
    DimensionEffect of going multi-agent
    CostTypically 3–10x more tokens for an equivalent task, from duplicated context, coordination messages and handoff summaries
    LatencyParallelism beats running the same work one step after another, but total time is often longer than a single agent's
    QualityBetter results when context pollution, parallel search or tool overload was the real constraint; otherwise better prompting of one agent can match it
    Reliability and upkeepEvery agent adds a possible point of failure and another set of prompts to maintain

    A code-review agent needs to investigate dozens of files to find the root cause of a bug, but you want to keep that investigation's file contents and intermediate reasoning out of the main conversation's context window, receiving only a concise final finding. Which capability best fits this need?

    Sources123

    3.Context isolation: keeping the main agent's reasoning clean

    Context windows are finite, and response quality can drop as context grows. Context pollution happens when material gathered for one subtask stays in the context and is irrelevant to the subtasks that follow. Anthropic's example is a customer support agent that has to look up order history while diagnosing a technical fault.

    A separate lookup agent reads the full order history in its own context and returns only a summarypython
    class OrderLookupAgent:
        def lookup_order(self, order_id: str) -> dict:
            # Separate agent with its own context
            messages = [
                {"role": "user", "content": f"Get essential details for order {order_id}"}
            ]
            response = client.messages.create(
                model="claude-sonnet-4-5",
                max_tokens=1024,
                messages=messages,
                tools=[get_order_details_tool]
            )
            # Returns only essential information
            return extract_summary(response)

    Isolation works best under three conditions. The subtask generates a lot of context (over 1,000 tokens) and most of it is irrelevant to the main task. The subtask is well defined, with clear criteria for what to extract. Or it is a lookup or retrieval that needs filtering before its result is used. Managed Agents enforces this separation: each agent runs in its own thread, so a subagent's reading never lands in the coordinator's history.

    Sources12

    4.Parallel subagents: buying thoroughness with tokens

    The second reason to split is to cover a larger search space than one agent can. In Anthropic's Research feature, a lead agent analyses the query, plans a strategy and spawns subagents to investigate different facets at the same time. The post calls search a problem of compression. Each subagent works in its own context window and condenses the most important tokens for the lead agent. Each also has its own tools, prompt and exploration path, which reduces path dependency.

    The core of parallel research: decompose the query, run subagents concurrently, then synthesizepython
    async def research_topic(query: str) -> dict:
        # Lead agent breaks query into research facets
        facets = await lead_agent.decompose_query(query)
    
        # Spawn subagents to research each facet in parallel
        tasks = [
            research_subagent(facet)
            for facet in facets
        ]
        results = await asyncio.gather(*tasks)
    
        # Lead agent synthesizes findings
        return await lead_agent.synthesize(results)

    On Anthropic's internal research eval, a system with Claude Opus 4 as lead and Claude Sonnet 4 subagents beat single-agent Claude Opus 4 by 90.2%. The gains were largest on breadth-first queries. One example was listing the board members of every IT company in the S&P 500: the single agent failed with slow, sequential searches. The post's explanation is budget. In its BrowseComp analysis, token usage alone explained 80% of the performance variance, and spreading the work across separate context windows adds capacity. The cost is steep: multi-agent systems used about 15× more tokens than chat interactions. They only make economic sense when the task is valuable enough to pay for the gain.

    During a PR review you want a style-checker, a security-scanner, and a test-coverage subagent to each analyze the same pull request, with all three finishing in roughly the time of the slowest one rather than the sum of all three. What should you do?

    Sources132

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Splitting a task into separate planning, execution, review and iteration agents will improve quality.Why is that wrong?

      Teams that built exactly this found it lost context at each handoff and spent more tokens coordinating than doing the work. Split only for context pollution, parallel work or specialisation.

      Covered in Deciding whether multiple agents are worth it

    2. 2.Running subagents in parallel makes a multi-agent system faster end to end than a single agent.Why is that wrong?

      Parallelism only beats running the same work in sequence. Because total computation goes up so much, multi-agent systems often take longer overall. What you gain is thoroughness.

      Covered in Parallel subagents: buying thoroughness with tokens

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Agents can act in parallel with their own isolated context, which helps improve output quality and can also improve time to completion.”
      ↩︎ What makes a system multi-agent
      “Tools, MCP servers, and context are not shared.”
      ↩︎ What makes a system multi-agent
      “All agents share the same sandbox, filesystem, and vault credentials”
      ↩︎ What makes a system multi-agent
      “best suited for complex tasks that either require work across a variety of surfaces, or where multiple well-scoped tasks contribute to an overall goal”
      ↩︎ Deciding whether multiple agents are worth it
      “each agent runs in its own session thread, a context-isolated event stream with its own conversation history”
      ↩︎ Context isolation: keeping the main agent's reasoning clean
      “Fan out independent subtasks simultaneously (searching multiple sources, analyzing separate files) and have the coordinator synthesize the results.”
      ↩︎ Parallel subagents: buying thoroughness with tokens
      “each agent runs in its own session thread, a context-isolated event stream with its own conversation history”
      ↩︎ Key concept
    2. 2.
      “A multi-agent system is an architecture where multiple LLM instances run with separate conversation contexts, coordinated through code.”
      ↩︎ What makes a system multi-agent
      “when context pollution degrades performance, when tasks can run in parallel, and when specialization improves tool selection or task focus”
      ↩︎ Deciding whether multiple agents are worth it
      “Outside these situations, the coordination costs typically exceed the benefits.”
      ↩︎ Deciding whether multiple agents are worth it
      “multi-agent implementations typically use 3-10x more tokens than single-agent approaches for equivalent tasks”
      ↩︎ Deciding whether multiple agents are worth it
      “The main agent receives only the 50-100 tokens it actually needs, keeping context focused.”
      ↩︎ Context isolation: keeping the main agent's reasoning clean
      “when subtasks generate high context volume (more than 1000 tokens) but most of that information is irrelevant to the main task”
      ↩︎ Context isolation: keeping the main agent's reasoning clean
      “The primary benefit of parallelization is thoroughness, not speed.”
      ↩︎ Parallel subagents: buying thoroughness with tokens
      “lost context at each handoff and spent more tokens coordinating than executing”
      ↩︎ Exam trap 1
      “multi-agent systems often take longer overall than single-agent systems because of the sheer increase in total computation”
      ↩︎ Exam trap 2
    3. 3.
      “most coding tasks involve fewer truly parallelizable tasks than research”
      ↩︎ Deciding whether multiple agents are worth it
      “outperformed single-agent Claude Opus 4 by 90.2% on our internal research eval”
      ↩︎ Parallel subagents: buying thoroughness with tokens
      “token usage by itself explains 80% of the variance”
      ↩︎ Parallel subagents: buying thoroughness with tokens
      “multi-agent systems use about 15× more tokens than chats”
      ↩︎ Parallel subagents: buying thoroughness with tokens

    Continue to page 2 of 2

    Orchestrator-Worker Design: Task Decomposition, Specialist Rosters and Synthesis