CertSafari
    CCAR-P · Lessons

    Domain 1 · Lesson 2/38

    Claude Workflow and Agent Patterns: Choosing the Architecture

    Design end-to-end architectures

    7 min read
    2.83% of exam
    7 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Tell workflows apart from agents and justify when either one is worth its cost and latency
    • Match prompt chaining, routing and parallelization to the task shapes they fit
    • Choose orchestrator-workers when subtasks cannot be predicted, and evaluator-optimizer when iterative feedback measurably helps
    • Link each pattern choice to its effect on cost, latency and quality

    1.Workflows, agents, and the case for neither

    Every Claude architecture sits somewhere on a scale of how much control the model has. Anthropic calls everything on that scale an agentic system, but separates two kinds. In workflows, your code fixes the path the LLM calls and tools follow. In agents, the model decides its own next step and which tools to use.

    The first design rule is to use the simplest option that works. Agentic systems usually spend more latency and cost to get better task performance, and that trade is not always worth it. For many applications, a single well-tuned LLM call with retrieval and in-context examples is enough. When more complexity is justified, workflows bring predictability and consistency to well-defined tasks. Agents fit better when you need flexibility and model-driven decisions at scale.

    The same reasoning applies to tools. A tool round trip earns its place when the task needs side effects, fresh or external data, a guaranteed output shape, or access to existing systems. For summarising, translating or answering from general knowledge, the extra round trip adds cost without adding value.

    An architect is building an agent that must read internal wiki pages, query a ticketing database, and post replies to a chat tool. The team wants to avoid writing and maintaining custom tool-execution code for each integration and wants the agent to autonomously decide when to call each system. Which architecture choice best satisfies this?

    Sources12

    2.Predefined paths: chaining, routing, parallelization

    All the patterns are built from the same block: an LLM with retrieval, tools and memory added. Three workflow patterns arrange that block in fixed ways. Anthropic's cookbook shows each one on a business task: a chain for structured data extraction and formatting, parallelization for stakeholder impact analysis, and routing for customer support tickets.

    The three fixed-path workflow patterns and what each one trades
    PatternStructureUse whenWhat it trades
    Prompt chainingEach LLM call processes the previous call's output; programmatic gates can check intermediate stepsThe task splits cleanly into fixed subtasksLatency, in return for higher accuracy from easier individual calls
    RoutingClassify the input, then send it to a specialised follow-up prompt, tools or modelDistinct categories are better handled separately and can be classified accuratelyA classification step, in return for prompts that don't hurt each other
    Parallelization: sectioningIndependent subtasks run at the same time and are combined in codeSubtasks are independent, or each consideration deserves its own focused callMore calls, in return for speed or focus
    Parallelization: votingThe same task runs several times to get varied outputsYou need more confidence, e.g. several reviewers flagging vulnerable codeMore calls, in return for confidence

    Routing is where architecture ties most directly to business value. A customer support conversation is itself a mix of different tasks, so sending refund requests, technical support and general questions to separate processes lets each prompt be tuned without degrading the others. The same mechanism can cut cost: send easy or common questions to a smaller, cheaper model and hard or unusual ones to a more capable model.

    Sectioning also improves guardrails. Having one model instance answer the user while another screens the query for inappropriate content tends to work better than asking a single call to do both.

    Sources134

    3.Dynamic delegation and built-in feedback loops

    Some tasks cannot be split up ahead of time. In a coding change, for example, which files need editing, and how, depends on the request itself. The orchestrator-workers pattern handles this. A central LLM breaks the task down at runtime, hands the pieces to worker LLMs, and combines what they return. It looks like parallelization, but the subtasks are chosen by the orchestrator for each specific input rather than fixed in advance.

    The cookbook implementation has two phases. In the first, the orchestrator analyses the task and writes structured subtask descriptions in XML. In the second, each worker gets the original task, its own subtask, and any extra context. Giving workers the full task, not just their piece, is a deliberate decision, and the implementation also rejects empty worker outputs. The cookbook is equally clear about when not to use the pattern: for simple single-output tasks, when latency is critical, or when the subtasks are predictable enough for plain parallelization.

    Production systems scope delegates tightly. In Anthropic's commerce reference design, the merchant agent can hand work to an optional analysis delegate that only runs read-only queries, within a time and size budget. Delegation happens inside a limit the architecture sets.

    The evaluator-optimizer pattern builds the feedback loop into the processing stage itself. One LLM call generates a response, another evaluates it and gives feedback, and the cycle repeats. It fits when there are clear evaluation criteria and when repeated refinement adds measurable value. Two signs point to it: responses clearly get better with feedback, and an LLM can give that feedback. Literary translation is the standard example. This is where the success criteria and benchmarks defined at the requirements stage are put to work.

    A production agent built with the Claude Agent SDK runs multi-hour autonomous coding sessions. The architecture needs a feedback loop that automatically writes an audit-log entry to disk every time the agent edits or writes a file, without the model having to be asked to do this in every prompt. Which mechanism should the architect use?

    Sources1567

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A production Claude solution should start as an autonomous agent, because agents perform best.Why is that wrong?

      The guidance is to find the simplest solution first. Agentic systems spend latency and cost to gain performance, and sometimes the right answer is no agentic system at all.

      Covered in Workflows, agents, and the case for neither

    2. 2.Orchestrator-workers is just parallelization with a coordinator; both use the same fixed set of subtasks.Why is that wrong?

      They look alike, but in orchestrator-workers the subtasks are decided at runtime for each input. If the subtasks can be defined in advance, the simpler parallelization pattern is the right choice.

      Covered in Dynamic delegation and built-in feedback loops

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Workflows are systems where LLMs and tools are orchestrated through predefined code paths.”
      ↩︎ Workflows, agents, and the case for neither
      “Agents, on the other hand, are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks.”
      ↩︎ Workflows, agents, and the case for neither
      “Agentic systems often trade latency and cost for better task performance, and you should consider when this tradeoff makes sense.”
      ↩︎ Workflows, agents, and the case for neither
      “For many applications, however, optimizing single LLM calls with retrieval and in-context examples is usually enough.”
      ↩︎ Workflows, agents, and the case for neither
      “The main goal is to trade off latency for higher accuracy, by making each LLM call an easier task.”
      ↩︎ Predefined paths: chaining, routing, parallelization
      “Without this workflow, optimizing for one kind of input can hurt performance on other inputs.”
      ↩︎ Predefined paths: chaining, routing, parallelization
      “Routing easy/common questions to smaller, cost-efficient models like Claude Haiku 4.5 and hard/unusual questions to more capable models like Claude Sonnet 4.5”
      ↩︎ Predefined paths: chaining, routing, parallelization
      “a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results.”
      ↩︎ Dynamic delegation and built-in feedback loops
      “one LLM call generates a response while another provides evaluation and feedback in a loop.”
      ↩︎ Dynamic delegation and built-in feedback loops
      “This might mean not building agentic systems at all.”
      ↩︎ Exam trap 1
      “subtasks aren't pre-defined, but determined by the orchestrator based on the specific input.”
      ↩︎ Exam trap 2
    2. 2.
      “Every tool call is at least one extra round trip; for lightweight tasks the overhead can exceed the work.”
      ↩︎ Workflows, agents, and the case for neither
    3. 5.
      “Workers receive both the original task AND their specific instructions for better context”
      ↩︎ Dynamic delegation and built-in feedback loops
      “Latency is critical (multiple LLM calls add overhead)”
      ↩︎ Dynamic delegation and built-in feedback loops
    4. 6.
      “An optional analysis delegate runs read-only queries under a time and size budget.”
      ↩︎ Dynamic delegation and built-in feedback loops
    5. 7.
      “LLM responses can be demonstrably improved when feedback is provided”
      ↩︎ Dynamic delegation and built-in feedback loops

    Ready to test yourself?

    Practise the 12 questions on this subdomain.