What you will be able to do
- Choose between the Messages API, the Claude Agent SDK and Claude Managed Agents based on how much control the business needs to keep
- Pick a starting model and justify it with cost figures
- Use multi-model strategies and parallel calls to meet cost and time targets
- Design guardrails so an agent's writes need human approval outside the conversation
1.Choose how much of the stack to own
Once the tasks and targets are defined, the first architecture decision is who runs the agent loop and the tools. Anthropic describes three paths, which differ "in how much control you keep and how much of the implementation you offload to Anthropic."
| Path | Who runs the agent loop | Who runs tools and infrastructure |
|---|---|---|
| Messages API | You write it | You run your own tools and infrastructure |
| Claude Agent SDK | Provided by the SDK | Tool execution provided, in a process you operate |
| Claude Managed Agents | Anthropic hosts it | Anthropic hosts tool execution and the runtime |
The choice also limits where the solution can run. In Anthropic's open-source commerce blueprint, the Messages API and Agent SDK runtimes can run against the Claude API, Amazon Bedrock, Google Cloud's Agent Platform, Microsoft Foundry, or your own gateway. The Managed Agents path runs on the Claude API. A business that is committed to one cloud should weigh that before picking the most managed option.
2.Pick a model and tie it to cost
Model choice is where architecture meets the budget. The guidance weighs capabilities, speed and cost, and it describes two ways to start. The cost-first route is to "Begin implementation with Claude Haiku 4.5," test thoroughly, and "Upgrade only if necessary for specific capability gaps." It suits prototyping, tight latency budgets, cost-sensitive deployments and high-volume, straightforward tasks. The capability-first route starts with Claude Opus 5.5 and optimises downward later. It suits complex reasoning, nuanced understanding, and cases where accuracy matters more than cost. The model-comparison guidance notes that most workloads start with Claude Opus 5.5.
Before you switch models, check whether effort can be tuned. Several models support an effort parameter that trades intelligence for latency and cost within one model, and the guidance says "Tuning effort is often a better lever than switching models."
Cost follows volume, so estimate it from the business's real numbers. The moderation guide sizes a platform that receives one billion 100-character posts a month. That comes to about 28.6B input tokens, with 50 output tokens for each of the 3% of posts that get flagged:
| Model | Input cost | Output cost | Monthly total |
|---|---|---|---|
| Claude Haiku 4.5 | $28,600 | $7,500 | $36,100 |
| Claude Opus 5 | $143,000 | $37,500 | $180,500 |
Legal summarisation lands the other way. Accuracy has a legal cost, and the whole job is 1,000 agreements, not billions. The guide calls Claude Opus 5 "an excellent choice" there, at an estimated $438.75 compared with $87.75 on Haiku 4.5. What decides the model is the business problem, not a default.
Either way, the decision should be checked against evidence. When you consider upgrading, the guidance says "having a good evaluation set is the most important step in the process." That is the per-task evaluation you defined while scoping.
An enterprise engineering organization needs an agent that can autonomously plan and execute a multi-hour refactor across a large, complex legacy codebase with minimal human checkpoints, prioritizing correctness over cost. Which model is the best fit for this workload?
Correct answer: C — Claude Opus 4.8, because it is built for complex agentic coding and large-scale refactoring with minimal supervision at enterprise scale.
- A. Incorrect. Running many cheap parallel attempts does not substitute for the sustained reasoning depth an autonomous multi-hour refactor with minimal checkpoints requires.
- B. Incorrect. Sonnet 5 handles many agentic coding tasks well, but it is not the model Anthropic positions for the most complex, high-autonomy, enterprise-scale refactoring work described here.
- C. Correct. Opus 4.8 is explicitly positioned for complex agentic coding and enterprise work, including multi-hour autonomous coding agents and large-scale refactoring.
- D. Incorrect. Context window size is independent of reasoning depth; extending Haiku 4.5's context does not give it Opus-level reasoning for a correctness-prioritized, low-supervision refactor.
3.Mix models and parallelise to meet cost and time targets
You don't have to run the whole solution on one model. "Multi-model strategies pair a lower-cost model with a frontier model so that most tokens are billed at the lower rate." The guidance names two patterns. In the first, an executor escalates hard decisions to an advisor. In the second, an orchestrator hands bulk work to lower-cost workers. The model table also lists sub-agent tasks among Claude Haiku 4.5's uses. In both patterns, the frontier model's price is paid only where its intelligence is actually needed.
Throughput is part of the design as well. Long documents can take up to a minute each to summarise, so for large collections the legal guide suggests you "send API calls to Claude in parallel so that the summaries can be completed in a reasonable timeframe," within Anthropic's rate limits. The same guide places an ingestion step in front of the model: "Before you begin summarizing documents, you need to prepare your data." That means extracting text from PDFs and cleaning it, and making sure the pipeline can handle every file format the business will actually send.
4.Build guardrails into the design
Once an agent can take actions, the design has to decide what it may do on its own. Anthropic's commerce blueprint answers this in three ways. First, facts come from tools. The shopping agent "states products, prices, availability, and store terms only from tool results in the conversation," and cart writes only accept product IDs that a tool returned in that session.
Second, writes are staged. Every change the merchant agent proposes, whether a price move, a restock or a campaign, becomes a staged change with a server-generated ID and a preview card. Guardrails such as maximum price move and campaign budget are checked twice: when the change is staged and again when it is applied. Third, a person approves outside the chat. The approval mechanism differs by build path: a portal button on the Messages API path, a console confirmation prompt in the Agent SDK, or an always-ask permission policy on the apply tool in Managed Agents. The rule is blunt: "An approval typed in chat approves nothing."
The other guides adapt the same idea to their own domain. Moderation can use risk levels instead of a yes/no verdict, because "Creating multiple risk levels allows you to adjust the aggressiveness of your moderation." High-risk posts can be blocked automatically while medium-risk patterns go to human review. Legal summaries should come with disclaimers saying the output is AI-generated and should be reviewed by legal professionals.
A company wants a single default model for most of its internal tooling: code generation, data analysis, content drafting, and agentic tool use across many teams, seeking the best combination of frontier intelligence and everyday speed without paying premium enterprise pricing. Which model should they standardize on?
Correct answer: B — Claude Sonnet 5, because it offers frontier intelligence at scale for coding, agents, and everyday workflows at moderate cost.
- A. Incorrect. Opus 4.8 targets complex agentic coding and enterprise work at premium pricing, which is more than most everyday drafting and analysis tasks need.
- B. Correct. Sonnet 5 is described as frontier intelligence at scale built for code generation, data analysis, content creation, and agentic tool use at moderate, introductory pricing.
- C. Incorrect. Haiku 4.5 is a strong economical choice for high-volume or latency-sensitive tasks, but it is not positioned as the best all-around fit for mixed coding, analysis, and drafting work across many teams.
- D. Incorrect. Fable 5 is priced and designed for next-generation, long-running agent intelligence, which is a costlier and narrower fit than the broad, moderate-cost standardization the company wants.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If a Claude solution is too slow or too expensive, the first fix is to switch to a smaller model.Why is that wrong?
Several models support an effort parameter that trades intelligence for latency and cost within the same model, and Anthropic says tuning it is often the better lever.
Covered in Pick a model and tie it to cost
2.If the operator confirms a proposed price change in the chat, the agent can apply it.Why is that wrong?
Staged writes apply only after a person approves them through a channel outside the conversation, such as a portal button, a console prompt or an always-ask permission policy.
Covered in Build guardrails into the design
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The Messages API gives you the most control: you write the agent loop and run your own tools and infrastructure.”
↩︎ Choose how much of the stack to own“With Claude Managed Agents, you offload the most: Anthropic hosts the agent loop, tool execution, and runtime for you.”
↩︎ Choose how much of the stack to own - 2.
“The Claude Managed Agents path runs on the Claude API.”
↩︎ Choose how much of the stack to own“The agent states products, prices, availability, and store terms only from tool results in the conversation”
↩︎ Build guardrails into the design“An approval typed in chat approves nothing.”
↩︎ Build guardrails into the design“The change applies only after a person approves it outside the conversation”
↩︎ Exam trap 2 - 3.
“Upgrade only if necessary for specific capability gaps.”
↩︎ Pick a model and tie it to cost“having a good evaluation set is the most important step in the process.”
↩︎ Pick a model and tie it to cost“The two common patterns are an executor that escalates hard decisions to an advisor, and an orchestrator that delegates bulk work to lower-cost workers.”
↩︎ Mix models and parallelise to meet cost and time targets“Tuning effort is often a better lever than switching models.”
↩︎ Exam trap 1 - 4.
“If costs are a concern, a smaller model such as Claude Haiku 4.5 is an excellent choice because of its cost-effectiveness.”
↩︎ Pick a model and tie it to cost“Creating multiple risk levels allows you to adjust the aggressiveness of your moderation.”
↩︎ Build guardrails into the design - 5.
“Claude Opus 5 is an excellent choice for use cases such as this where high accuracy is required.”
↩︎ Pick a model and tie it to cost“you may want to send API calls to Claude in parallel so that the summaries can be completed in a reasonable timeframe.”
↩︎ Mix models and parallelise to meet cost and time targets