What you will be able to do
- Turn a business requirement into a functional requirement that sets how much LLM complexity the solution needs
- Tell a workflow from an agent and pick the one a requirement calls for
- Turn capability, speed and cost requirements into a starting model and an effort setting
- Explain why an evaluation set is what decides whether a requirement has been met
Key concept
Simplest sufficient solution — Choose the least complex architecture that meets the business requirement. Add model calls, workflows or autonomy only when a measured gap shows you need them, because each step up costs latency and money.
1.From a business need to a functional requirement
A business requirement says what the organisation wants, such as "answer customer questions faster" or "review every pull request". A functional requirement says what the system must do to deliver that. With Claude, the gap between the two is mostly an architecture decision: one model call, a fixed chain of calls, or an agent that decides its own steps.
Anthropic's guidance starts from restraint. The recommended default is the simplest solution that works, and sometimes that means no agentic system at all. The reason is a trade-off every requirement has to justify: agentic systems often trade latency and cost for better task performance. When a requirement names a response-time target or a budget, that trade-off is part of the requirement, not something to settle later.
Anthropic also notes that for many applications, improving a single LLM call with retrieval and in-context examples is usually enough. So the first functional question is always whether a well-built single call already meets the need. Then turn the business need into three concrete questions from the model-selection guidance. What capabilities must the model have? How fast must it respond? What can you spend on development and on production use?
2.Workflow or agent: matching architecture to the requirement
When a single call is not enough, the next choice is structural. Anthropic groups both options as agentic systems but separates them clearly. In a workflow, your code fixes the path through the LLM calls. In an agent, the model directs its own process and tool use. The requirement tells you which one to use. If the task is well defined and the business needs the same behaviour every time, choose a workflow. If the steps can't be known in advance and the model has to decide what to do, choose an agent.
| Pattern | Requirement signal | Example from the guidance |
|---|---|---|
| Prompt chaining | Task splits cleanly into fixed subtasks; accuracy matters more than latency | Write an outline, check it meets criteria, then write the document |
| Routing | Distinct input categories that are better handled separately | Send refund requests, general questions and technical support to different processes |
| Parallelization | Independent subtasks for speed, or several attempts for higher confidence | One instance answers queries while another screens them |
| Orchestrator-workers | Subtasks can't be predicted in advance | Coding changes whose file count depends on the task |
| Evaluator-optimizer | Clear evaluation criteria, and iterative refinement adds measurable value | Literary translation reviewed by an evaluator LLM |
Routing shows how architecture and cost requirements meet. The guidance gives the example of sending easy, common questions to a smaller, cheaper model like Claude Haiku 4.5 and hard or unusual ones to a more capable model. The model-selection docs describe this as a multi-model strategy: pair a lower-cost model with a frontier model so that most tokens are billed at the lower rate.
A workflow using routing. The categories are distinct and known in advance, and the business wants the refund path to behave the same way every time. That calls for a predefined code path, not model-driven decisions.
3.Turning speed, cost and capability into a model choice
Once the architecture is set, the capability, speed and cost questions become a model decision. The docs describe two starting strategies, and the requirement decides which one fits.
| Strategy | Start with | Best when the requirement is |
|---|---|---|
| Cost-first, upgrade if needed | Claude Haiku 4.5 | Prototyping, tight latency, cost-sensitive, high-volume straightforward tasks |
| Capability-first, optimize down | Claude Opus 5.5 | Complex reasoning, nuanced understanding, accuracy over cost, high-autonomy agentic work |
| Escalate further | Claude Fable 5.1 | Evals at xhigh or max effort still fall short on demanding reasoning or long-horizon agentic work |
There is a lever inside a single model too. Several Claude models support an effort parameter that trades intelligence for latency and cost, and the docs say tuning effort is often a better lever than switching models. Defaults vary by model. On Claude Opus 5.5 the default is medium, and on Claude Fable 5.1 and Claude Opus 5 it is high. In both cases you start at the default and adjust based on your evals. For a pure speed requirement, Claude Opus 5.5, Claude Opus 5 and Claude Opus 4.8 support fast mode (research preview), which delivers up to 2.5x higher output speed at premium pricing.
Sources2
4.Evals are how a requirement gets verified
None of these choices is final until it is measured. The docs make the evaluation set the centre of the process. You build benchmark tests specific to your use case and run them with your real prompts and data. Then you compare models on accuracy, response quality and edge-case handling, and weigh performance against cost. This is what makes development iterative. The cost-first path begins with Haiku 4.5, tests the use case, and upgrades only to close a specific capability gap. The capability-first path begins high and lowers effort or moves to a smaller model as the workflow matures.
In practice, a requirement you can't test with an eval is not finished. "Answers must be accurate" turns into a benchmark set with a pass threshold. That threshold, not intuition, decides whether a simpler model, lower effort or a simpler architecture is enough.
A data team runs nightly analysis over millions of records. Results are only needed by the next business morning, and minimizing per-token cost is the top infrastructure priority. Which processing approach should the architecture use?
Correct answer: B — The Message Batches API
- A. Synchronous Messages API calls are priced at standard rates and are built for immediate, per-request responses, not for minimizing cost on a large overnight-tolerant workload.
- B. The Message Batches API processes large volumes of requests asynchronously at 50% lower cost than standard API calls, matching an overnight, cost-first workload with no immediate latency need.
- C. Streaming responses reduce perceived latency for interactive use, which is irrelevant here since results are only needed by the next morning, and it does not reduce per-token cost.
- D. Re-uploading files repeatedly adds unnecessary overhead and is meant to avoid redundant uploads across separate requests, not to reduce per-token processing cost for a batch analysis job.
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A more agentic, multi-step design is always the stronger answer to a business requirement.Why is that wrong?
Agentic systems cost latency and money. The guidance recommends the simplest solution that works, and sometimes that means no agentic system at all.
2.To make a solution cheaper or faster, you must switch to a smaller model.Why is that wrong?
Several Claude models expose an effort parameter that trades intelligence for latency and cost within the same model. The docs call it often the better lever.
Covered in Turning speed, cost and capability into a model choice
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://www.anthropic.com/engineering/building-effective-agentsSecondary source
“Agentic systems often trade latency and cost for better task performance, and you should consider when this tradeoff makes sense.”
↩︎ From a business need to a functional requirement“workflows offer predictability and consistency for well-defined tasks, whereas agents are the better option when flexibility and model-driven decision-making are needed at scale.”
↩︎ Workflow or agent: matching architecture to the requirement“Workflows are systems where LLMs and tools are orchestrated through predefined code paths.”
↩︎ Workflow or agent: matching architecture to the requirement“we recommend finding the simplest solution possible, and only increasing complexity when needed.”
↩︎ Key concept“Agentic systems often trade latency and cost for better task performance, and you should consider when this tradeoff makes sense.”
↩︎ Exam trap 1 - 2.
“Capabilities: What specific features or capabilities will you need the model to have to meet your needs?”
↩︎ From a business need to a functional requirement“Cost: What's your budget for both development and production usage?”
↩︎ From a business need to a functional requirement“Multi-model strategies pair a lower-cost model with a frontier model so that most tokens are billed at the lower rate.”
↩︎ Workflow or agent: matching architecture to the requirement“For complex tasks where intelligence and advanced capabilities are paramount, you may want to start capability-first”
↩︎ Turning speed, cost and capability into a model choice“starting with a faster, more cost-effective model like Claude Haiku 4.5 can be the optimal approach”
↩︎ Turning speed, cost and capability into a model choice“Tuning effort is often a better lever than switching models.”
↩︎ Turning speed, cost and capability into a model choice“Create benchmark tests specific to your use case - having a good evaluation set is the most important step in the process.”
↩︎ Evals are how a requirement gets verified“Weigh performance and cost tradeoffs.”
↩︎ Evals are how a requirement gets verified“Tuning effort is often a better lever than switching models.”
↩︎ Exam trap 2