What you will be able to do
- Split a multi-concern customer request into independent items and decide which investigations may run in parallel
- Use SubagentStart and SubagentStop to track parallel investigations and collect their results before synthesizing
- Compile a self-contained handoff summary for a human agent who cannot see the transcript
1.Decompose the request before you investigate it
A customer writes: "I was charged twice last month, my shipping address is wrong, and I want a refund for the duplicate." That is one message but three concerns: a billing dispute, an account-data fix, and a refund that depends on the outcome of the first. An agent that handles it as one blob tends to answer the loudest concern and drop the rest.
Anthropic's agent-patterns guidance gives the reason to split: "For complex tasks with multiple considerations, LLMs generally perform better when each consideration is handled by a separate LLM call". The matching pattern is sectioning: "Breaking a task into independent subtasks run in parallel." When the concerns can't be known in advance and a model has to identify them from the input, the guidance describes orchestrator-workers: "a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results". Customer requests fit that description, because the coordinator only learns how many concerns there are by reading the message.
The same post lists routing customer queries as a separate idea: "Directing different types of customer service queries (general questions, refund requests, technical support) into different downstream processes, prompts, and tools." Routing picks one path for a whole request. Decomposition gives each concern in a mixed request its own path, then brings the paths back together.
The shared context is what makes the separate investigations add up. Each item is investigated against the same customer record, the one the verified get_customer call established. The billing investigation and the refund decision therefore describe the same account, and the final synthesis can resolve dependencies, such as the refund depending on whether the duplicate charge is confirmed.
2.Which investigations may run in parallel
At the tool level, parallelism is already built in: "By default, Claude may call multiple tools in a single response." The API leaves execution order to you, and the parallel-tool-use page gives the rule for choosing: "Independent, read-only operations are usually safe to run in parallel for lower latency." By contrast: "Tools with side effects, shared state, or ordering requirements might be better run sequentially."
That rule splits the example ticket cleanly. Looking up the charge history and reading the current address are independent reads, so they can fan out. Issuing the refund has a side effect and a dependency on the billing finding, so it waits until the investigations have finished. A prerequisite gate on the refund tool still applies there.
| Item | Operation type | Run |
|---|---|---|
| Investigate duplicate charge | Independent, read-only | In parallel |
| Check shipping address on file | Independent, read-only | In parallel |
| Issue refund for duplicate | Side effects, ordering requirement | Sequentially, after the billing finding |
One formatting rule keeps this working: "return one tool_result for each tool_use block, all together in the next user message". If you decide not to run one call, for example because an earlier call in a sequential batch failed, you still return a result for it with is_error: true and a short explanation. The model then sees every branch accounted for.
A tool_result with the refund call's tool_use_id, is_error: true, and a brief explanation such as "Not executed: the preceding lookup failed." It goes in the same user message as the other results. Leaving it out breaks the one-result-per-call rule.
Sources2
3.Knowing the investigations are done: SubagentStart and SubagentStop
When the coordinator sends each concern to its own subagent, it needs a reliable signal that every one has finished before it writes the unified answer. Guessing from elapsed time or from the model's own narration is the prompt-based approach again. Hooks give you lifecycle events for this instead.
| Hook | What triggers it | Example use case |
|---|---|---|
| SubagentStart | Subagent initialization | Track parallel task spawning |
| SubagentStop | Subagent completion | Aggregate results from parallel tasks |
| PostToolBatch | A full batch of tool calls resolves, once per batch before the next model call | Inject conventions once for the whole batch |
| Stop | Agent execution stop | Save session state before exit |
Record each spawn in SubagentStart and each completion in SubagentStop, and you have a ledger. Synthesis waits until every started investigation has a matching stop. Stop is a different event: in Claude Code it fires "When Claude finishes responding", which is the end of the agent's turn, not the end of one worker. PostToolBatch is about batches of parallel tool calls, not subagents. Note that the hooks table marks PostToolBatch as TypeScript-only in the SDK, while SubagentStart and SubagentStop are available in both Python and TypeScript.
A team registers three independent PreToolUse hooks for process_refund: one checks identity verification, one checks fraud score, and one logs the attempt for audit. On one ticket, the verification hook returns "deny" while the fraud-score hook returns "allow" and the logging hook returns an empty object. What happens to the tool call?
Correct answer: D — The call is blocked, because when multiple hooks disagree, a "deny" from any hook overrides "allow" results from the others
- A. Incorrect. Deny outranks ask in the priority order, so a mixed result does not fall back to a manual-approval prompt when one hook already denied it.
- B. Incorrect. All matching hooks' outputs are considered together; an empty object from one hook does not remove it from the resolution, and deny still overrides allow.
- C. Incorrect. Decisions are not resolved by majority vote; a single deny is sufficient to block the call regardless of how many other hooks allowed it.
- D. Correct. When several hooks fire for the same event, the most restrictive decision wins: deny takes priority over defer, which takes priority over ask, which takes priority over allow.
4.Escalating mid-process: the structured handoff
Sometimes an investigation ends with a decision the agent isn't allowed to make, such as a policy exception only a supervisor can approve. The Agent SDK lists this as a use of hooks: "Require human approval for sensitive actions like database writes or API calls". A gate on the exception-granting action is how the escalation gets triggered reliably.
What the sources do not cover: none of the documentation in this lesson defines a handoff-summary format for human agents. The content below comes from the exam guide's own statement of the skill, not from vendor docs. The guide names the fields: customer ID, root cause, refund amount, and recommended action, together with the customer details and the root-cause analysis gathered so far.
The design reason is the constraint the guide states: the human agent cannot see the conversation transcript. The summary is therefore the only thing they get, the same way a subagent receives only what is passed to it. It has to stand alone. "Customer disputes charge, see above" is useless. A usable summary gives the verified customer ID, what the investigation found and why (the root cause), the exact amount at stake, and the specific action the agent recommends and needs approved. The human should be able to act without re-running the investigation.
A support agent has two tools, get_customer and process_refund. The team's system prompt says "Always call get_customer and confirm the identity before calling process_refund." During testing, the agent occasionally calls process_refund first when a ticket is phrased as an urgent complaint. The architect wants this ordering to hold every time, not just most of the time. What should they implement?
Correct answer: B — A PreToolUse hook matched to process_refund that checks for a verified customer ID in session state and returns permissionDecision "deny" when none exists
- A. Incorrect. A larger context window helps retention of information but does not create a hard block on calling process_refund out of order.
- B. Correct. A PreToolUse hook is programmatic enforcement: it inspects state before the tool executes and can deny the call outright, so the ordering holds regardless of how the prompt is phrased.
- C. Incorrect. Rewording the same prompt-based instruction still leaves compliance dependent on the model following it; it does not close the non-zero failure rate the architect is trying to eliminate.
- D. Incorrect. An injected reminder is still a prompt-based nudge the model can override under unusual phrasing; it is not a gate on the tool call itself.
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Once a request is decomposed, every resulting tool call should run concurrently for speed, including the refund.Why is that wrong?
Only independent, read-only work is the safe parallel case. Calls with side effects or ordering requirements, like a refund that depends on the billing finding, should run sequentially.
Covered in Which investigations may run in parallel
2.The Stop hook is the right place to learn that each parallel subagent has finished and to collect its findings.Why is that wrong?
Stop fires when the agent's execution stops. SubagentStop fires on each subagent's completion and is the documented place to aggregate results from parallel tasks.
Covered in Knowing the investigations are done: SubagentStart and SubagentStop
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://www.anthropic.com/engineering/building-effective-agentsSecondary source
“For complex tasks with multiple considerations, LLMs generally perform better when each consideration is handled by a separate LLM call”
↩︎ Decompose the request before you investigate it“a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results”
↩︎ Decompose the request before you investigate it“Directing different types of customer service queries (general questions, refund requests, technical support) into different downstream processes, prompts, and tools.”
↩︎ Decompose the request before you investigate it - 2.
“By default, Claude may call multiple tools in a single response.”
↩︎ Decompose the request before you investigate it“Independent, read-only operations are usually safe to run in parallel for lower latency.”
↩︎ Which investigations may run in parallel“return one tool_result for each tool_use block, all together in the next user message”
↩︎ Which investigations may run in parallel“Tools with side effects, shared state, or ordering requirements might be better run sequentially.”
↩︎ Exam trap 1 - 3.https://code.claude.com/docs/en/agent-sdk/hooksOfficial docs
“Aggregate results from parallel tasks”
↩︎ Knowing the investigations are done: SubagentStart and SubagentStop“Require human approval for sensitive actions like database writes or API calls”
↩︎ Escalating mid-process: the structured handoff“Aggregate results from parallel tasks”
↩︎ Exam trap 2 - 4.https://code.claude.com/docs/en/hooksOfficial docs
“When Claude finishes responding”
↩︎ Knowing the investigations are done: SubagentStart and SubagentStop