CertSafari
    CLAUDE-CERTIFIED-ARCHITECT-FOUNDATIONS-CCAR-F · Lessons

    Domain 1 · Lesson 4/30

    Multi-Concern Requests: Parallel Investigation, Synthesis and Human Handoff

    Implement multi-step workflows with enforcement and handoff patterns

    8 min read
    3.86% of exam
    4 sources
    Published 28 Sep 2026
    Docs as of 27 Sep 2026

    What you will be able to do

    • Split a multi-concern customer request into independent items and decide which investigations may run in parallel
    • Use SubagentStart and SubagentStop to track parallel investigations and collect their results before synthesizing
    • Compile a self-contained handoff summary for a human agent who cannot see the transcript

    1.Decompose the request before you investigate it

    A customer writes: "I was charged twice last month, my shipping address is wrong, and I want a refund for the duplicate." That is one message but three concerns: a billing dispute, an account-data fix, and a refund that depends on the outcome of the first. An agent that handles it as one blob tends to answer the loudest concern and drop the rest.

    Anthropic's agent-patterns guidance gives the reason to split: "For complex tasks with multiple considerations, LLMs generally perform better when each consideration is handled by a separate LLM call". The matching pattern is sectioning: "Breaking a task into independent subtasks run in parallel." When the concerns can't be known in advance and a model has to identify them from the input, the guidance describes orchestrator-workers: "a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results". Customer requests fit that description, because the coordinator only learns how many concerns there are by reading the message.

    The same post lists routing customer queries as a separate idea: "Directing different types of customer service queries (general questions, refund requests, technical support) into different downstream processes, prompts, and tools." Routing picks one path for a whole request. Decomposition gives each concern in a mixed request its own path, then brings the paths back together.

    The shared context is what makes the separate investigations add up. Each item is investigated against the same customer record, the one the verified get_customer call established. The billing investigation and the refund decision therefore describe the same account, and the final synthesis can resolve dependencies, such as the refund depending on whether the duplicate charge is confirmed.

    Sources12

    2.Which investigations may run in parallel

    At the tool level, parallelism is already built in: "By default, Claude may call multiple tools in a single response." The API leaves execution order to you, and the parallel-tool-use page gives the rule for choosing: "Independent, read-only operations are usually safe to run in parallel for lower latency." By contrast: "Tools with side effects, shared state, or ordering requirements might be better run sequentially."

    That rule splits the example ticket cleanly. Looking up the charge history and reading the current address are independent reads, so they can fan out. Issuing the refund has a side effect and a dependency on the billing finding, so it waits until the investigations have finished. A prerequisite gate on the refund tool still applies there.

    Classifying the items from one multi-concern ticket
    ItemOperation typeRun
    Investigate duplicate chargeIndependent, read-onlyIn parallel
    Check shipping address on fileIndependent, read-onlyIn parallel
    Issue refund for duplicateSide effects, ordering requirementSequentially, after the billing finding

    One formatting rule keeps this working: "return one tool_result for each tool_use block, all together in the next user message". If you decide not to run one call, for example because an earlier call in a sequential batch failed, you still return a result for it with is_error: true and a short explanation. The model then sees every branch accounted for.

    Sources2

    3.Knowing the investigations are done: SubagentStart and SubagentStop

    When the coordinator sends each concern to its own subagent, it needs a reliable signal that every one has finished before it writes the unified answer. Guessing from elapsed time or from the model's own narration is the prompt-based approach again. Hooks give you lifecycle events for this instead.

    Lifecycle hooks relevant to fan-out and fan-in
    HookWhat triggers itExample use case
    SubagentStartSubagent initializationTrack parallel task spawning
    SubagentStopSubagent completionAggregate results from parallel tasks
    PostToolBatchA full batch of tool calls resolves, once per batch before the next model callInject conventions once for the whole batch
    StopAgent execution stopSave session state before exit

    Record each spawn in SubagentStart and each completion in SubagentStop, and you have a ledger. Synthesis waits until every started investigation has a matching stop. Stop is a different event: in Claude Code it fires "When Claude finishes responding", which is the end of the agent's turn, not the end of one worker. PostToolBatch is about batches of parallel tool calls, not subagents. Note that the hooks table marks PostToolBatch as TypeScript-only in the SDK, while SubagentStart and SubagentStop are available in both Python and TypeScript.

    A team registers three independent PreToolUse hooks for process_refund: one checks identity verification, one checks fraud score, and one logs the attempt for audit. On one ticket, the verification hook returns "deny" while the fraud-score hook returns "allow" and the logging hook returns an empty object. What happens to the tool call?

    Sources34

    4.Escalating mid-process: the structured handoff

    Sometimes an investigation ends with a decision the agent isn't allowed to make, such as a policy exception only a supervisor can approve. The Agent SDK lists this as a use of hooks: "Require human approval for sensitive actions like database writes or API calls". A gate on the exception-granting action is how the escalation gets triggered reliably.

    What the sources do not cover: none of the documentation in this lesson defines a handoff-summary format for human agents. The content below comes from the exam guide's own statement of the skill, not from vendor docs. The guide names the fields: customer ID, root cause, refund amount, and recommended action, together with the customer details and the root-cause analysis gathered so far.

    The design reason is the constraint the guide states: the human agent cannot see the conversation transcript. The summary is therefore the only thing they get, the same way a subagent receives only what is passed to it. It has to stand alone. "Customer disputes charge, see above" is useless. A usable summary gives the verified customer ID, what the investigation found and why (the root cause), the exact amount at stake, and the specific action the agent recommends and needs approved. The human should be able to act without re-running the investigation.

    A support agent has two tools, get_customer and process_refund. The team's system prompt says "Always call get_customer and confirm the identity before calling process_refund." During testing, the agent occasionally calls process_refund first when a ticket is phrased as an urgent complaint. The architect wants this ordering to hold every time, not just most of the time. What should they implement?

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Once a request is decomposed, every resulting tool call should run concurrently for speed, including the refund.Why is that wrong?

      Only independent, read-only work is the safe parallel case. Calls with side effects or ordering requirements, like a refund that depends on the billing finding, should run sequentially.

      Covered in Which investigations may run in parallel

    2. 2.The Stop hook is the right place to learn that each parallel subagent has finished and to collect its findings.Why is that wrong?

      Stop fires when the agent's execution stops. SubagentStop fires on each subagent's completion and is the documented place to aggregate results from parallel tasks.

      Covered in Knowing the investigations are done: SubagentStart and SubagentStop

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “For complex tasks with multiple considerations, LLMs generally perform better when each consideration is handled by a separate LLM call”
      ↩︎ Decompose the request before you investigate it
      “a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results”
      ↩︎ Decompose the request before you investigate it
      “Directing different types of customer service queries (general questions, refund requests, technical support) into different downstream processes, prompts, and tools.”
      ↩︎ Decompose the request before you investigate it
    2. 2.
      “By default, Claude may call multiple tools in a single response.”
      ↩︎ Decompose the request before you investigate it
      “Independent, read-only operations are usually safe to run in parallel for lower latency.”
      ↩︎ Which investigations may run in parallel
      “return one tool_result for each tool_use block, all together in the next user message”
      ↩︎ Which investigations may run in parallel
      “Tools with side effects, shared state, or ordering requirements might be better run sequentially.”
      ↩︎ Exam trap 1
    3. 3.
      “Require human approval for sensitive actions like database writes or API calls”
      ↩︎ Escalating mid-process: the structured handoff
      “Aggregate results from parallel tasks”
      ↩︎ Exam trap 2