CertSafari
    CCAR-P · Lessons

    Domain 1 · Lesson 3/38

    Workflow Patterns: Chaining, Routing, Parallelization, Orchestrator-Workers, Evaluator-Optimizer

    Select appropriate architectural patterns

    9 min read
    2.83% of exam
    4 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Match a business requirement to prompt chaining, routing, parallelization, orchestrator-workers or evaluator-optimizer
    • Tell sectioning from voting, and orchestrator-workers from plain parallelization
    • Name the conditions under which evaluator-optimizer is a good fit
    • Account for the token and latency cost of multi-call and multi-agent designs when picking a pattern

    1.Fixed paths: prompt chaining and routing

    Once one augmented LLM call is not enough, the simplest step up is a workflow whose path you write yourself. Two patterns fix the path in advance. They differ in whether the path is the same for every input.

    Prompt chaining splits a task into a sequence of steps, and each LLM call processes the output of the previous one. You can put a programmatic "gate" on any intermediate step to check that the process is still on track. The main goal is to trade latency for higher accuracy by making each call an easier task. It fits when the task splits cleanly into fixed subtasks. Examples: writing marketing copy and then translating it, or writing an outline, checking it against criteria, and then writing the document. Because each step is a separate API call, you can log, evaluate, or branch at any point.

    Routing classifies an input and sends it to a specialized follow-up task. It separates concerns, which matters because optimizing one prompt for one kind of input can hurt performance on other inputs. Routing works when there are distinct categories that are better handled separately, and when classification is accurate, whether done by an LLM or a traditional classifier. Sending customer-service queries (general questions, refunds, technical support) to different prompts and tools is one example. Sending easy questions to Claude Haiku 4.5 and hard ones to Claude Sonnet 4.5, to balance cost and performance, is another.

    A customer support system receives billing questions, technical bugs, and account cancellations. Each category needs a differently tuned prompt, and an initial step must first determine which category a given ticket belongs to before further processing. Which pattern should the team implement?

    Sources12

    2.Parallelization: sectioning and voting

    In parallelization, several LLM calls work on the task at the same time and code aggregates their outputs. There are two variations. Sectioning breaks a task into independent subtasks that run in parallel. Voting runs the same task several times to get diverse outputs.

    Choose sectioning for speed, or when a task has several separate considerations. LLMs generally do better when each consideration gets its own call and its own focused attention. One example is guardrails: one instance answers the user while another screens the query for inappropriate content. This tends to work better than one call doing both. Choose voting when you need several attempts or perspectives for higher confidence. Examples are several prompts reviewing code for vulnerabilities, or several evaluators judging content with different vote thresholds to balance false positives against false negatives.

    Claude's own guidance on delegation uses the same conditions: use subagents when tasks can run in parallel, need isolated context, or are independent workstreams that don't need to share state. When the subtasks are not independent, parallelization is the wrong choice.

    Sources12

    3.Orchestrator-workers: when you can't list the subtasks in advance

    In orchestrator-workers, a central LLM breaks the task down at runtime, delegates the pieces to worker LLMs, and combines their results. Its topology looks like parallelization. The difference is flexibility: the subtasks are not predefined, and the orchestrator decides them from the specific input. That is why it suits coding changes, where the number of files to change and the nature of each change depend on the task. It also suits search tasks that gather and analyze information from many sources.

    Lead agent decomposes the query, runs subagents concurrently, then synthesizespython
    async def research_topic(query: str) -> dict:
        # Lead agent breaks query into research facets
        facets = await lead_agent.decompose_query(query)
    
        # Spawn subagents to research each facet in parallel
        tasks = [
            research_subagent(facet)
            for facet in facets
        ]
        results = await asyncio.gather(*tasks)
    
        # Lead agent synthesizes findings
        return await lead_agent.synthesize(results)

    The cookbook also lists when not to use it. Skip it for simple tasks with a single output, when latency is critical (each extra LLM call adds overhead), or when the subtasks are predictable and can be defined in advance. In that last case, plain parallelization is simpler. Current Claude models orchestrate subagents natively and will delegate to well-described subagent tools without explicit instruction, so the pattern is cheap to reach for. That makes it more important to check that it is actually needed.

    A security team wants one LLM call to check generated code for vulnerabilities while a second, independent LLM call simultaneously checks the same code for licensing issues, with both outputs combined into a single report. Which pattern does this describe?

    Sources132

    4.Evaluator-optimizer: a generate–critique loop

    In evaluator-optimizer, one LLM call generates a response and another evaluates it and gives feedback, in a loop. It is especially effective when there are clear evaluation criteria and iterative refinement adds measurable value. There are two signs of a good fit. First, responses clearly improve when a human states their feedback. Second, the LLM can give that kind of feedback itself. Literary translation is Anthropic's example: an evaluator can catch nuances the translator missed on its first attempt.

    Claude's prompting guidance names the most common chaining pattern: self-correction, where Claude drafts, reviews the draft against criteria, and then refines it based on the review. That is an evaluator-optimizer loop built from chained calls. Without clear criteria to judge against, the loop has nothing to push the output toward.

    Sources12

    5.Picking the pattern, and what more agents cost

    Workflow patterns against the signal in the requirement that points to each
    PatternRequirement signalTrade-off accepted
    Prompt chainingTask decomposes cleanly into fixed subtasksLatency for higher accuracy
    RoutingDistinct input categories; classification can be accurateA classification step before the specialized path
    Parallelization (sectioning)Independent subtasks, or separate considerations each needing focusMore calls, aggregated programmatically
    Parallelization (voting)Multiple attempts needed for higher confidenceThe same task run several times
    Orchestrator-workersSubtasks can't be predicted; they depend on the inputA central LLM plans and synthesizes
    Evaluator-optimizerClear evaluation criteria; iteration gives measurable valueRepeated generate–evaluate loops

    Designs that spread work across several agents have a price. In Anthropic's testing, multi-agent implementations typically used 3–10x more tokens than single-agent approaches on equivalent tasks. The extra comes from duplicated context, coordination messages, and summaries passed at each handoff. They are often slower overall too. Parallel runs cut time compared with doing the same work in sequence, but total computation goes up. The main benefit of parallel agents is thoroughness, not speed. Anthropic also reports teams that spent months on elaborate multi-agent architectures, only to find that better prompting of a single agent gave the same results.

    Models can over-delegate too. Claude Opus 4.6 has a strong predilection for subagents and may spawn them where a simpler, direct approach would do. For example, it may explore code with subagents when a direct grep would be faster. Whichever pattern you choose, keep it only as long as it earns its cost.

    Sources42

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Orchestrator-workers is just parallelization with a nicer name, so the two can be swapped freely.Why is that wrong?

      They look the same, but in parallelization you define the subtasks in advance. In orchestrator-workers, a central LLM decides them for each input. When the subtasks are predictable, use the simpler parallelization.

      Covered in Orchestrator-workers: when you can't list the subtasks in advance

    2. 2.Splitting work across parallel agents is the way to cut latency and cost.Why is that wrong?

      Multi-agent systems typically use 3–10x more tokens and often take longer overall. The main benefit of parallel agents is covering more ground.

      Covered in Picking the pattern, and what more agents cost

    3. 3.Adding an evaluator loop improves any generation task.Why is that wrong?

      Evaluator-optimizer fits only when there are clear evaluation criteria and iterative refinement adds measurable value. Without those, the loop only adds calls.

      Covered in Evaluator-optimizer: a generate–critique loop

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Prompt chaining decomposes a task into a sequence of steps, where each LLM call processes the output of the previous one.”
      ↩︎ Fixed paths: prompt chaining and routing
      “The main goal is to trade off latency for higher accuracy, by making each LLM call an easier task.”
      ↩︎ Fixed paths: prompt chaining and routing
      “Without this workflow, optimizing for one kind of input can hurt performance on other inputs.”
      ↩︎ Fixed paths: prompt chaining and routing
      “LLMs generally perform better when each consideration is handled by a separate LLM call, allowing focused attention on each specific aspect.”
      ↩︎ Parallelization: sectioning and voting
      “a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results.”
      ↩︎ Orchestrator-workers: when you can't list the subtasks in advance
      “one LLM call generates a response while another provides evaluation and feedback in a loop.”
      ↩︎ Evaluator-optimizer: a generate–critique loop
      “subtasks aren't pre-defined, but determined by the orchestrator based on the specific input.”
      ↩︎ Exam trap 1
      “when we have clear evaluation criteria, and when iterative refinement provides measurable value”
      ↩︎ Exam trap 3
    2. 2.
      “Each step is a separate API call so you can log, evaluate, or branch at any point.”
      ↩︎ Fixed paths: prompt chaining and routing
      “Use subagents when tasks can run in parallel, require isolated context, or involve”
      ↩︎ Parallelization: sectioning and voting
      “Claude's latest models orchestrate subagents natively.”
      ↩︎ Orchestrator-workers: when you can't list the subtasks in advance
      “generate a draft → have Claude review it against criteria → have Claude refine based on the review”
      ↩︎ Evaluator-optimizer: a generate–critique loop
      “may spawn them in situations where a simpler, direct approach would suffice”
      ↩︎ Picking the pattern, and what more agents cost
    3. 3.
      “Subtasks are predictable and can be pre-defined (use simpler parallelization)”
      ↩︎ Orchestrator-workers: when you can't list the subtasks in advance
    4. 4.
      “In our testing, multi-agent implementations typically use 3-10x more tokens than single-agent approaches for equivalent tasks.”
      ↩︎ Picking the pattern, and what more agents cost
      “multi-agent systems are often applied in situations where a single agent would perform better”
      ↩︎ Picking the pattern, and what more agents cost
      “The primary benefit of parallelization is thoroughness, not speed.”
      ↩︎ Exam trap 2

    Ready to test yourself?

    Practise the 12 questions on this subdomain.