What you will be able to do
- Recognise the signs that a single request is carrying several distinct jobs
- Restructure a complex request as ordered, numbered steps inside one prompt
- Judge when splitting a task adds cost without adding quality
Key concept
Task decomposition — Task decomposition means splitting a complex request into smaller subtasks. These can be ordered steps inside one prompt or separate stages. Each piece gets focused attention and produces an output you can check. It helps when a request bundles distinct jobs. It is not something to do by default.
1.Why a single prompt can carry too much
Take a request like this one: "Read these supplier contracts, pull out the payment terms, compare them against our procurement policy, flag the risky ones, and draft an email to the finance director recommending which to renegotiate." It reads as one ask, but it holds at least four different jobs. There is extraction, then comparison against a standard, then judgement, then drafting for a particular reader. Each job has its own definition of done.
Anthropic's guidance on consistency gives the fix directly. Break a complex task into smaller subtasks, because each one then gets Claude's full attention and the results become more consistent. Anthropic's guidance on building agents makes the same point from the other side. When a task involves several considerations, models generally do better when each one is handled separately rather than all at once.
So length is not the sign that a prompt is doing too much. A long prompt that gives rich context for one job is fine. The sign is that the request bundles jobs with different outputs or different success criteria, or that one job needs another's finished result before it can start. The contract example has all three. You cannot judge which terms are risky until they have been extracted, and the email cannot recommend anything until that judgement is made.
Assigning owners needs the finished list of decisions. The executive update needs both the decisions and the owners. That makes three kinds of output: an extracted list, a judgement about owners, and a draft for a specific audience. They are linked by dependencies, so a prompt that asks for all of them at once is doing too much.
2.The lightest form: numbered steps in one prompt
Decomposing a task doesn't always mean splitting it across several conversations. The lightest form keeps one prompt but writes the work out as an explicit sequence. Claude's prompting guidance recommends exactly this whenever order or completeness matters: give the instructions as numbered steps or bullet points.
Rewritten this way, the contract request becomes: 1. From each contract below, list the payment terms (due days, penalties, early-payment discounts). 2. Compare each set of terms with the procurement policy below and mark any that breach it. 3. For each breach, say how serious it is and why. 4. Using only the breaches you marked, draft a short email to the finance director recommending which contracts to renegotiate first.
Two things changed. The order is now explicit, so drafting cannot start before the comparison exists. And completeness is now checkable: you can confirm that step 1 covered every contract before you trust step 4. Numbered steps also make follow-ups precise, because you can point at a single step instead of the whole answer.
Each step should say what it produces. The prompting guidance asks you to be specific about output format and constraints, and that applies to every step. Step 1 produces a list with named fields, and step 4 produces an email for a named reader. A step with no defined output is where a decomposed prompt quietly slides back into one vague request.
Sources3
3.When decomposition costs more than it gives
Splitting has a price. Every extra stage is another round trip, which means more time, more cost and more output for you to review. Anthropic's advice for systems built on Claude is to find the simplest solution that works and add complexity only when it's needed. It also notes that, for many applications, a single well-built call with good context and examples is enough.
The guidance for current models gives the same warning. Handing work to separate agents pays off for large, genuinely independent pieces of work, but on small tasks it multiplies cost and time. Anthropic's cookbook lists simple, single-output tasks as a case where a multi-call pattern only adds unnecessary complexity.
| Question to ask | If yes | If no |
|---|---|---|
| Does the request bundle jobs with different outputs or success criteria? | Separate them, at least as numbered steps | One well-specified prompt is likely enough |
| Does a later job need the finished, checked result of an earlier one? | Run them in order and look at the earlier result before continuing | The jobs may be independent and can be handled side by side |
| Is the task small, with a single output? | Don't split it; extra stages add overhead without improving quality | Consider how far to decompose |
A legal team needs to compare two lengthy contracts to highlight differences. What is the safest decomposition approach with Claude?
Correct answer: B — Extract the key clauses from each contract first, then compare them point by point.
- A. Incorrect. Asking Claude to output a comparison table with both contracts in a single request risks exceeding the context window limit, leading to truncated input or incomplete comparisons. This approach also makes it harder to manage nuanced differences safely.
- B. Correct. Extracting key clauses first reduces complexity, focuses on relevant parts, and minimizes context overflow. This step-by-step decomposition improves accuracy by breaking the task into manageable, point-by-point comparisons.
- C. Incorrect. Summarizing each contract separately may omit important legal details and precise wording differences necessary for accurate comparison. Comparing only summaries is less reliable than directly comparing extracted clauses.
- D. Incorrect. Loading both lengthy contracts into a single context window to generate a detailed line-by-line diff risks exceeding the token limit, resulting in truncated input or an incomplete diff. This approach is less safe than decomposing the task for thorough comparison.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Breaking a request into more stages always improves the result.Why is that wrong?
Decomposition trades time and cost for quality. It only pays off when the work is large or genuinely separable. On small tasks it just multiplies the overhead.
Covered in When decomposition costs more than it gives
2.A prompt is doing too much simply because it is long.Why is that wrong?
Length is not the signal. The signal is a request that bundles several considerations with different outputs. Models do better when each consideration is handled separately.
Covered in Why a single prompt can carry too much
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/increase-consistencyOfficial docs
“Break down complex tasks into smaller, consistent subtasks. Each subtask gets Claude's full attention, reducing inconsistency errors across scaled workflows.”
↩︎ Why a single prompt can carry too much“Break down complex tasks into smaller, consistent subtasks. Each subtask gets Claude's full attention, reducing inconsistency errors across scaled workflows.”
↩︎ Key concept - 2.https://www.anthropic.com/engineering/building-effective-agentsSecondary source
“For complex tasks with multiple considerations, LLMs generally perform better when each consideration is handled by a separate LLM call”
↩︎ Why a single prompt can carry too much“we recommend finding the simplest solution possible, and only increasing complexity when needed.”
↩︎ When decomposition costs more than it gives“For many applications, however, optimizing single LLM calls with retrieval and in-context examples is usually enough.”
↩︎ When decomposition costs more than it gives“For complex tasks with multiple considerations, LLMs generally perform better when each consideration is handled by a separate LLM call”
↩︎ Exam trap 2 - 3.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practicesOfficial docs
“Provide instructions as sequential steps using numbered lists or bullet points when the order or completeness of steps matters.”
↩︎ The lightest form: numbered steps in one prompt“Be specific about the desired output format and constraints.”
↩︎ The lightest form: numbered steps in one prompt - 4.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5Official docs
“Delegation pays off on genuinely independent, sizeable tracks of work, but it multiplies cost and time when applied to small tasks.”
↩︎ When decomposition costs more than it gives“Delegation pays off on genuinely independent, sizeable tracks of work, but it multiplies cost and time when applied to small tasks.”
↩︎ Exam trap 1 - 5.
“You have simple, single-output tasks (unnecessary complexity)”
↩︎ When decomposition costs more than it gives