What you will be able to do
- Explain why the simplest workable setup is the default and when extra complexity is justified
- Rewrite a vague request into a scoped one that avoids rounds of correction
- Break a large task into small, checkable steps instead of one oversized request
- Match a task to the lightest pattern that fits, including routing easy work to smaller models
Key concept
Simplest-first — Start with the least elaborate setup that could do the job, such as one well-specified request, and add steps, agents or bigger models only when that simpler version demonstrably falls short. Every added layer costs time and money, so it has to earn its place.
1.Efficiency starts with doing less
Most people think optimizing a workflow means adding things: more steps, more agents, a bigger model. Anthropic's guidance for building with language models points the other way. Find the simplest approach that works, and add complexity only when the result actually needs it. Sometimes that means building no multi-step system at all.
The reason is a trade-off you always pay. Multi-step and agentic setups often buy better task performance with extra latency and cost, so the question is never "would more machinery help?" but "is the gain worth what it costs here?". For many applications, the guidance adds, improving a single call with the right retrieved material and a few in-context examples is usually enough.
In practice, efficiency and effectiveness are the same discipline. An effective workflow gets the right result. An efficient one gets it without wasted round trips. A common source of waste is not the model or the tooling at all. It is a request that was under-specified, came back wrong, and had to be regenerated. So optimization starts with the smallest unit, the individual request, before anyone reaches for a pipeline.
2.Scope the request so it lands the first time
Claude Code's best-practices guide makes the case with before-and-after pairs. The "before" prompts are short and feel efficient. Each one leaves Claude guessing at the file, the scenario or the success condition, so the first answer is likely to miss and the real cost shows up later as rework. The "after" prompts are longer, but they do their work in one pass.
| Strategy | Before | After |
|---|---|---|
| Scope the task | add tests for foo.py | write a test for foo.py covering the edge case where the user is logged out. avoid mocks. |
| Point to sources | why does ExecutionFactory have such a weird api? | look through ExecutionFactory’s git history and summarize how its api came to be |
| Describe the symptom | fix the login bug | users report that login fails after session timeout. check the auth flow in src/auth/, especially token refresh. write a failing test that reproduces the issue, then fix it |
The same guide also covers how to supply context cheaply. Rather than describing where material lives, reference the file directly so Claude reads it before it responds. You can also paste images, give documentation URLs, pipe data straight in, or tell Claude to fetch what it needs itself. Each of these swaps your paraphrase, which can drift from the real thing, for context the model reads at the source.
Sources2
3.Break big jobs into small, checkable steps
Once a single request is well scoped, the next lever is shape. A large job crammed into one request asks the model to do everything at once, and any error spreads through the whole output. The simplest multi-step pattern, prompt chaining, splits the task into a sequence in which each call works on the previous call's output. You can place a check, or "gate", between steps to confirm things are still on track before continuing. The stated trade is deliberate: accept a little more latency in exchange for higher accuracy, because each call now has an easier job. It fits best when the task divides cleanly into fixed subtasks, such as outlining a document, checking the outline against criteria, and then writing from it.
Claude Code's own recipes follow the same rhythm. They recommend doing refactoring in small, testable increments, and planning before editing so you can review changes before they touch disk. Its best-practices flow of explore, then plan, then implement, then commit is the same idea applied to everyday work. A cheap planning step catches a wrong direction before the expensive implementation step commits to it.
An analyst begins every chat in her weekly reporting project by typing the same six lines about house style, the sections required and how figures should be rounded. What is the most effective first change?
Correct answer: A — Move those standing lines into the project's instructions so every chat there starts with them.
- A. Correct. A requirement repeated in every chat of the same project belongs in the project instructions, where it applies automatically and stops depending on her remembering it.
- B. A note reduces typing but still relies on her pasting it correctly every single time.
- C. Asking Claude to recall earlier style is unreliable and does not make the requirement durable.
- D. Uploaded documents are background knowledge; a standing rule about how to work belongs in the instructions.
4.Right-size the machinery, including the model
When a single call really isn't enough, the patterns in Anthropic's agent-building guidance form a ladder of increasing complexity. Each rung has a condition that justifies it. Climb only as far as the task requires.
| Pattern | Use it when | Example from the guidance |
|---|---|---|
| Single optimized call | Retrieval and in-context examples already get the result | Usually enough for many applications |
| Prompt chaining | The task decomposes cleanly into fixed subtasks | Write marketing copy, then translate it |
| Routing | Inputs fall into distinct categories that can be classified accurately | Easy questions to Claude Haiku 4.5, hard ones to Claude Sonnet 4.5 |
| Parallelization | Subtasks are independent, or several attempts raise confidence | One call screens content while another answers |
| Orchestrator-workers | You cannot predict the subtasks in advance | Coding changes across an unknown number of files |
Routing carries the model-choice lesson. Easy, common questions go to smaller, cost-efficient models, and only hard or unusual ones go to more capable models. The capable model is kept for the work that needs it rather than used everywhere by default. Routing also stops one prompt from being tuned for every kind of input at once, which the guidance warns can hurt performance on the others.
Scale matters too. Claude Code's comparison of delegation tools runs from subagents handling a few delegated tasks per turn to workflows that run dozens to hundreds of agents per run. That range is a menu, not a target. Anthropic's orchestrator-workers cookbook explicitly advises against the pattern for simple, single-output tasks. The same thinking applies to small habits: the Claude Code guide's sample project instructions prefer running single tests rather than the whole suite, for performance.
A four-person team runs out of its weekly Claude allowance every Thursday and work stalls until the limit resets. What should they do first?
Correct answer: B — Look at which tasks are actually consuming the allowance before changing plans or habits.
- A. Forcing the fastest model on everything trades quality away without knowing what actually consumed the allowance.
- B. Correct. Until the team knows where the allowance is actually going, every remedy is a guess; checking which work is consuming it points at the habit or task to change.
- C. Short messages can cost more overall if they cause long back-and-forth exchanges to reach the same result.
- D. Upgrading may be the right answer eventually, but paying to avoid a diagnosis usually hides the real cause.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Adding agents or parallel runs always makes a workflow more effective, so the most elaborate setup is the safest choice.Why is that wrong?
Extra machinery buys performance with latency and cost. It is justified only when the task needs it, and many tasks are served by one well-built call.
2.Sending every request to the most capable model is the efficient default.Why is that wrong?
The guidance routes easy, common questions to smaller, cost-efficient models and keeps more capable models for hard or unusual ones.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://www.anthropic.com/engineering/building-effective-agentsSecondary source
“For many applications, however, optimizing single LLM calls with retrieval and in-context examples is usually enough.”
↩︎ Efficiency starts with doing less“The main goal is to trade off latency for higher accuracy, by making each LLM call an easier task.”
↩︎ Break big jobs into small, checkable steps“Without this workflow, optimizing for one kind of input can hurt performance on other inputs.”
↩︎ Right-size the machinery, including the model“When building applications with LLMs, we recommend finding the simplest solution possible, and only increasing complexity when needed.”
↩︎ Key concept“Agentic systems often trade latency and cost for better task performance, and you should consider when this tradeoff makes sense.”
↩︎ Exam trap 1“Routing easy/common questions to smaller, cost-efficient models like Claude Haiku 4.5”
↩︎ Exam trap 2 - 2.https://code.claude.com/docs/en/best-practicesOfficial docs
“Scope the task. Specify which file, what scenario, and testing preferences.”
↩︎ Efficiency starts with doing less“Reference files with @ instead of describing where code lives. Claude reads the file before responding.”
↩︎ Scope the request so it lands the first time“Let Claude fetch what it needs. Tell Claude to pull context itself using Bash commands, MCP tools, or by reading files.”
↩︎ Scope the request so it lands the first time“Prefer running single tests, and not the whole test suite, for performance”
↩︎ Right-size the machinery, including the model - 3.https://code.claude.com/docs/en/common-workflowsOfficial docs
“Do refactoring in small, testable increments”
↩︎ Break big jobs into small, checkable steps“Plan before editing to review changes before they touch disk”
↩︎ Break big jobs into small, checkable steps - 4.https://code.claude.com/docs/en/workflowsOfficial docs
“Dozens to hundreds of agents per run”
↩︎ Right-size the machinery, including the model - 5.
“You have simple, single-output tasks (unnecessary complexity)”
↩︎ Right-size the machinery, including the model