What you will be able to do
- Match a business requirement to prompt chaining, routing, parallelization, orchestrator-workers or evaluator-optimizer
- Tell sectioning from voting, and orchestrator-workers from plain parallelization
- Name the conditions under which evaluator-optimizer is a good fit
- Account for the token and latency cost of multi-call and multi-agent designs when picking a pattern
1.Fixed paths: prompt chaining and routing
Once one augmented LLM call is not enough, the simplest step up is a workflow whose path you write yourself. Two patterns fix the path in advance. They differ in whether the path is the same for every input.
Prompt chaining splits a task into a sequence of steps, and each LLM call processes the output of the previous one. You can put a programmatic "gate" on any intermediate step to check that the process is still on track. The main goal is to trade latency for higher accuracy by making each call an easier task. It fits when the task splits cleanly into fixed subtasks. Examples: writing marketing copy and then translating it, or writing an outline, checking it against criteria, and then writing the document. Because each step is a separate API call, you can log, evaluate, or branch at any point.
Routing classifies an input and sends it to a specialized follow-up task. It separates concerns, which matters because optimizing one prompt for one kind of input can hurt performance on other inputs. Routing works when there are distinct categories that are better handled separately, and when classification is accurate, whether done by an LLM or a traditional classifier. Sending customer-service queries (general questions, refunds, technical support) to different prompts and tools is one example. Sending easy questions to Claude Haiku 4.5 and hard ones to Claude Sonnet 4.5, to balance cost and performance, is another.
Routing. The inputs fall into distinct categories that are handled differently, so a classifier sends each one to the right prompt or model. Chaining would run every input through the same fixed sequence.
A customer support system receives billing questions, technical bugs, and account cancellations. Each category needs a differently tuned prompt, and an initial step must first determine which category a given ticket belongs to before further processing. Which pattern should the team implement?
Correct answer: B — Routing, where an initial LLM call classifies the ticket and directs it to the specialized prompt for that category
- A. Incorrect. Parallelization runs independent subtasks or repeated attempts concurrently; it does not classify an input into one of several distinct categories.
- B. Correct. An initial classification step that directs the ticket to a category-specific specialized prompt is the defining shape of routing.
- C. Incorrect. Prompt chaining passes every ticket through every step in sequence; here only one category-specific prompt should run per ticket, not all three.
- D. Incorrect. Orchestrator-workers is for tasks whose subtasks are unpredictable and must be dynamically decomposed, not for a single classification-to-specialized-prompt decision.
2.Parallelization: sectioning and voting
In parallelization, several LLM calls work on the task at the same time and code aggregates their outputs. There are two variations. Sectioning breaks a task into independent subtasks that run in parallel. Voting runs the same task several times to get diverse outputs.
Choose sectioning for speed, or when a task has several separate considerations. LLMs generally do better when each consideration gets its own call and its own focused attention. One example is guardrails: one instance answers the user while another screens the query for inappropriate content. This tends to work better than one call doing both. Choose voting when you need several attempts or perspectives for higher confidence. Examples are several prompts reviewing code for vulnerabilities, or several evaluators judging content with different vote thresholds to balance false positives against false negatives.
Claude's own guidance on delegation uses the same conditions: use subagents when tasks can run in parallel, need isolated context, or are independent workstreams that don't need to share state. When the subtasks are not independent, parallelization is the wrong choice.
3.Orchestrator-workers: when you can't list the subtasks in advance
In orchestrator-workers, a central LLM breaks the task down at runtime, delegates the pieces to worker LLMs, and combines their results. Its topology looks like parallelization. The difference is flexibility: the subtasks are not predefined, and the orchestrator decides them from the specific input. That is why it suits coding changes, where the number of files to change and the nature of each change depend on the task. It also suits search tasks that gather and analyze information from many sources.
async def research_topic(query: str) -> dict:
# Lead agent breaks query into research facets
facets = await lead_agent.decompose_query(query)
# Spawn subagents to research each facet in parallel
tasks = [
research_subagent(facet)
for facet in facets
]
results = await asyncio.gather(*tasks)
# Lead agent synthesizes findings
return await lead_agent.synthesize(results)The cookbook also lists when not to use it. Skip it for simple tasks with a single output, when latency is critical (each extra LLM call adds overhead), or when the subtasks are predictable and can be defined in advance. In that last case, plain parallelization is simpler. Current Claude models orchestrate subagents natively and will delegate to well-described subagent tools without explicit instruction, so the pattern is cheap to reach for. That makes it more important to check that it is actually needed.
A security team wants one LLM call to check generated code for vulnerabilities while a second, independent LLM call simultaneously checks the same code for licensing issues, with both outputs combined into a single report. Which pattern does this describe?
Correct answer: A — Parallelization by sectioning, running the vulnerability and licensing checks as independent, simultaneous LLM calls
- A. Correct. Two distinct subtasks running independently and simultaneously, with their outputs combined afterward, is parallelization by sectioning.
- B. Incorrect. Voting repeats the same prompt multiple times to build confidence in one answer; here the two checks are different tasks, not repeated attempts at the same task.
- C. Incorrect. Evaluator-optimizer involves one LLM iteratively critiquing and refining another's output toward a quality target, not two independent checks combined into a report.
- D. Incorrect. Prompt chaining feeds one step's output into the next sequential step; the scenario describes two checks running independently and simultaneously, not sequentially.
4.Evaluator-optimizer: a generate–critique loop
In evaluator-optimizer, one LLM call generates a response and another evaluates it and gives feedback, in a loop. It is especially effective when there are clear evaluation criteria and iterative refinement adds measurable value. There are two signs of a good fit. First, responses clearly improve when a human states their feedback. Second, the LLM can give that kind of feedback itself. Literary translation is Anthropic's example: an evaluator can catch nuances the translator missed on its first attempt.
Claude's prompting guidance names the most common chaining pattern: self-correction, where Claude drafts, reviews the draft against criteria, and then refines it based on the review. That is an evaluator-optimizer loop built from chained calls. Without clear criteria to judge against, the loop has nothing to push the output toward.
5.Picking the pattern, and what more agents cost
| Pattern | Requirement signal | Trade-off accepted |
|---|---|---|
| Prompt chaining | Task decomposes cleanly into fixed subtasks | Latency for higher accuracy |
| Routing | Distinct input categories; classification can be accurate | A classification step before the specialized path |
| Parallelization (sectioning) | Independent subtasks, or separate considerations each needing focus | More calls, aggregated programmatically |
| Parallelization (voting) | Multiple attempts needed for higher confidence | The same task run several times |
| Orchestrator-workers | Subtasks can't be predicted; they depend on the input | A central LLM plans and synthesizes |
| Evaluator-optimizer | Clear evaluation criteria; iteration gives measurable value | Repeated generate–evaluate loops |
Designs that spread work across several agents have a price. In Anthropic's testing, multi-agent implementations typically used 3–10x more tokens than single-agent approaches on equivalent tasks. The extra comes from duplicated context, coordination messages, and summaries passed at each handoff. They are often slower overall too. Parallel runs cut time compared with doing the same work in sequence, but total computation goes up. The main benefit of parallel agents is thoroughness, not speed. Anthropic also reports teams that spent months on elaborate multi-agent architectures, only to find that better prompting of a single agent gave the same results.
Models can over-delegate too. Claude Opus 4.6 has a strong predilection for subagents and may spawn them where a simpler, direct approach would do. For example, it may explore code with subagents when a direct grep would be faster. Whichever pattern you choose, keep it only as long as it earns its cost.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Orchestrator-workers is just parallelization with a nicer name, so the two can be swapped freely.Why is that wrong?
They look the same, but in parallelization you define the subtasks in advance. In orchestrator-workers, a central LLM decides them for each input. When the subtasks are predictable, use the simpler parallelization.
Covered in Orchestrator-workers: when you can't list the subtasks in advance
2.Splitting work across parallel agents is the way to cut latency and cost.Why is that wrong?
Multi-agent systems typically use 3–10x more tokens and often take longer overall. The main benefit of parallel agents is covering more ground.
3.Adding an evaluator loop improves any generation task.Why is that wrong?
Evaluator-optimizer fits only when there are clear evaluation criteria and iterative refinement adds measurable value. Without those, the loop only adds calls.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://www.anthropic.com/engineering/building-effective-agentsSecondary source
“Prompt chaining decomposes a task into a sequence of steps, where each LLM call processes the output of the previous one.”
↩︎ Fixed paths: prompt chaining and routing“The main goal is to trade off latency for higher accuracy, by making each LLM call an easier task.”
↩︎ Fixed paths: prompt chaining and routing“Without this workflow, optimizing for one kind of input can hurt performance on other inputs.”
↩︎ Fixed paths: prompt chaining and routing“LLMs generally perform better when each consideration is handled by a separate LLM call, allowing focused attention on each specific aspect.”
↩︎ Parallelization: sectioning and voting“a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results.”
↩︎ Orchestrator-workers: when you can't list the subtasks in advance“one LLM call generates a response while another provides evaluation and feedback in a loop.”
↩︎ Evaluator-optimizer: a generate–critique loop“subtasks aren't pre-defined, but determined by the orchestrator based on the specific input.”
↩︎ Exam trap 1“when we have clear evaluation criteria, and when iterative refinement provides measurable value”
↩︎ Exam trap 3 - 2.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practicesOfficial docs
“Each step is a separate API call so you can log, evaluate, or branch at any point.”
↩︎ Fixed paths: prompt chaining and routing“Use subagents when tasks can run in parallel, require isolated context, or involve”
↩︎ Parallelization: sectioning and voting“Claude's latest models orchestrate subagents natively.”
↩︎ Orchestrator-workers: when you can't list the subtasks in advance“generate a draft → have Claude review it against criteria → have Claude refine based on the review”
↩︎ Evaluator-optimizer: a generate–critique loop“may spawn them in situations where a simpler, direct approach would suffice”
↩︎ Picking the pattern, and what more agents cost - 3.
“Subtasks are predictable and can be pre-defined (use simpler parallelization)”
↩︎ Orchestrator-workers: when you can't list the subtasks in advance - 4.
“In our testing, multi-agent implementations typically use 3-10x more tokens than single-agent approaches for equivalent tasks.”
↩︎ Picking the pattern, and what more agents cost“multi-agent systems are often applied in situations where a single agent would perform better”
↩︎ Picking the pattern, and what more agents cost“The primary benefit of parallelization is thoroughness, not speed.”
↩︎ Exam trap 2