What you will be able to do
- Decompose a business problem into steps before deciding what Claude handles at each one
- Choose between sequential, category-based and parallel cuts for a task whose subtasks are known in advance
- Explain the latency, cost and accuracy trade-off each fixed decomposition makes
Key concept
Fixed vs. dynamic decomposition — The first decomposition question is whether you can name the subtasks before the input arrives. If you can, the split is written into code (a chain, a router, a fan-out); if you cannot, a model has to decide the split at runtime.
1.Break the problem down before you pick a pattern
Decomposition starts with the problem, not with the model. Anthropic's builder course reframes delegation: the question is not whether to use AI somewhere, but how a customer problem breaks into steps and what role AI plays in each one. In its worked scenario, a community health clinic wants patients to check wait times, and the learner writes no code in the first session. They pin down who the users are, what outcome those users need, and what the real constraints are: budget, the clinic's technical skill, patients' access to devices, and privacy requirements.
Only after that do they map each piece of the build to a collaboration mode. Automation means AI does it and you check it. Augmentation means you and AI work on it together. Agency means AI operates with latitude inside boundaries you set. This is decomposition done at the level of the business problem. Each step then gets its own decision about how much the model is trusted to do, and those decisions tell you which technical pattern fits.
Sources1
2.Cutting by sequence and by input category
Once the steps are known, the simplest cut is sequential. Prompt chaining splits a task into steps, and each LLM call processes the output of the one before it. You can put programmatic checks, which Anthropic calls gates, on any intermediate step to confirm the process is still on track. Examples are writing marketing copy and then translating it, or drafting an outline, checking it against criteria, and then writing the document. Chaining deliberately gives up speed: every call becomes an easier task, and the price is extra round-trips.
The second cut is by input category. Routing classifies an input and sends it to a specialised follow-up task. You split along the categories because a single prompt tuned for one kind of input can hurt performance on the others. Customer-service traffic splits naturally into general questions, refund requests and technical support, each with its own prompts and tools. The same cut can be made by difficulty: send easy or common questions to Claude Haiku 4.5 and hard or unusual ones to Claude Sonnet 4.5. Routing only works if the classification itself is accurate, whether an LLM or a traditional classifier does it.
Claude Managed Agents documents the same idea at the multi-agent level as specialisation. Instead of one agent that carries every capability, you route work to agents built around a single domain.
Latency goes up, because there are four sequential round-trips. Accuracy per step should go up, because each call has one narrow job. Anthropic describes chaining as trading latency for accuracy. The Anthropic cookbook's chain example runs these same four steps on a quarterly metrics summary to produce a sorted markdown table.
3.Cutting into parallel pieces: sectioning and voting
When subtasks do not depend on each other, you can run them at the same time and combine the outputs in code. Parallelization comes in two forms. Sectioning splits the task into different independent pieces. Voting runs the same task several times to get diverse outputs you can compare. The two cuts answer different needs. Sectioning is for speed, or for giving each consideration its own focused call. Voting is for confidence.
Sectioning has a quality argument as well as a speed one. When a task has several considerations, LLMs generally do better if each consideration gets its own call. Anthropic's guardrail example has one model instance answer the user while another screens the query for inappropriate content, which tends to beat one call doing both. Automated evals work the same way, with one call per aspect being graded. Voting fits review tasks such as several prompts checking code for vulnerabilities, where you can tune the vote threshold to balance false positives against false negatives.
Managed Agents expresses the same cut as fan-out. The coordinator sends independent subtasks, such as searching several sources or analysing separate files, to other agents at once and then synthesises what comes back. In every version of the pattern, the recombination step is part of the design and not an afterthought.
A team is building an agent that must audit a large codebase across three distinct dimensions: security vulnerabilities, code quality, and test coverage. Each dimension requires a different set of tools and a different area of focus. Which approach best decomposes this problem using the Agent SDK?
Correct answer: A — Define an orchestrator agent that invokes specialized subagents through the Agent tool, each configured with a scoped tool set and description matching its subtask domain
- A. Correct. The Agent SDK's subagent capability lets an orchestrator delegate focused subtasks to specialized agents, each with its own scoped tools and description, then aggregate their results — the intended decomposition pattern for multi-dimensional review work.
- B. Incorrect. A single undifferentiated prompt forces one agent to hold all three review contexts and tool needs at once, losing the isolation and specialization benefits of decomposition.
- C. Incorrect. Exposing checks as tools on one MCP server still leaves a single agent responsible for orchestrating and interpreting all three domains itself, rather than delegating focused reasoning to specialized subagents.
- D. Incorrect. A larger context window and extended thinking add reasoning depth but do not decompose the task into independently scoped, delegatable units of work.
4.Separating generation from judgement
One more fixed cut splits making something from judging it. In the evaluator-optimizer workflow, one call produces a response and another evaluates it and sends feedback, round after round. It pays off when the evaluation criteria are clear and each refinement pass adds measurable value. Two signs of a good fit: a human's written feedback demonstrably improves the output, and an LLM can give that same kind of feedback. Literary translation is Anthropic's example, because an evaluator can catch nuances the translator missed.
The builder course makes the same move before any code exists: write the acceptance tests first. Those tests are your evaluation criteria. If you cannot write them, an evaluator loop has nothing to measure against. The table below gathers the fixed cuts in one place.
| Pattern | How the task is cut | Use when | Trade-off |
|---|---|---|---|
| Prompt chaining | Into a sequence; each call consumes the previous output | Task cleanly decomposes into fixed subtasks | Higher latency for higher accuracy |
| Routing | By input category or difficulty | Distinct categories exist and classification is accurate | Misclassification sends input down the wrong path |
| Parallelization: sectioning | Into independent subtasks run concurrently | Subtasks are independent, or each consideration needs focused attention | More calls; outputs must be aggregated |
| Parallelization: voting | Same task run multiple times | Higher confidence needed from multiple attempts | Cost multiplies with each vote |
| Evaluator-optimizer | Generator and evaluator in a feedback loop | Clear evaluation criteria; refinement gives measurable value | Extra iterations per output |
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Splitting one prompt into a prompt chain is a way to make the pipeline faster.Why is that wrong?
Chaining adds sequential calls. Its purpose is accuracy, because each call gets an easier task, and it pays for that in latency.
Covered in Cutting by sequence and by input category
2.Voting and sectioning are both ways to split a task into different pieces.Why is that wrong?
Only sectioning splits the task into different independent subtasks. Voting runs the same task several times and compares the outputs to gain confidence.
Covered in Cutting into parallel pieces: sectioning and voting
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://academy.claude.com/courses/ai-fluency-for-builders/delegation-the-builder-s-toolkitOfficial docs
“Delegation means decomposing the problem first, then deciding what AI handles at each step.”
↩︎ Break the problem down before you pick a pattern“Think budget, technical skill at the clinic, patient access to devices, privacy requirements.”
↩︎ Break the problem down before you pick a pattern“Write acceptance tests before code. They give you and AI a shared definition of done.”
↩︎ Separating generation from judgement - 2.
“Route to agents with domain-focused system prompts and tools, such as a security agent or a documentation agent”
↩︎ Cutting by sequence and by input category“Fan out independent subtasks simultaneously (searching multiple sources, analyzing separate files) and have the coordinator synthesize the results.”
↩︎ Cutting into parallel pieces: sectioning and voting - 3.https://www.anthropic.com/engineering/building-effective-agentsSecondary source
“Prompt chaining decomposes a task into a sequence of steps, where each LLM call processes the output of the previous one.”
↩︎ Cutting by sequence and by input category“Without this workflow, optimizing for one kind of input can hurt performance on other inputs.”
↩︎ Cutting by sequence and by input category“For complex tasks with multiple considerations, LLMs generally perform better when each consideration is handled by a separate LLM call”
↩︎ Cutting into parallel pieces: sectioning and voting“This workflow is particularly effective when we have clear evaluation criteria, and when iterative refinement provides measurable value.”
↩︎ Separating generation from judgement“This workflow is ideal for situations where the task can be easily and cleanly decomposed into fixed subtasks.”
↩︎ Key concept“The main goal is to trade off latency for higher accuracy, by making each LLM call an easier task.”
↩︎ Exam trap 1“Voting: Running the same task multiple times to get diverse outputs.”
↩︎ Exam trap 2