What you will be able to do
- Name the three requirements (capability, speed, cost) that drive a model choice and say what each one asks of the task
- Place Haiku, Sonnet and Opus on the speed, price and reasoning-depth trade-off
- Use task difficulty, latency needs and volume to pick a tier
- Explain why the most capable tier is the wrong answer for some workloads
Key concept
The cost–speed–quality trade-off — Claude's tiers differ mainly in how hard a problem each can reliably handle and what that costs in money and time. A good choice is the least expensive, fastest tier that still meets the quality the task needs, not the strongest tier available.
1.Three questions to ask before you pick a model
Anthropic's model-selection guidance begins with a short list of factors to weigh, not with a ranking of models. Three of them match this objective: capabilities (what the model has to be able to do for this task), speed (how quickly it has to respond) and cost (what you can spend, both while you are setting the work up and once it runs routinely).
They pull against each other. Higher capability usually means a slower, pricier model, and the cheapest, fastest option may not be good enough. The guidance on cost says both halves directly. The most capable model can cost too much once work runs at scale, and the least expensive one can miss the quality bar. So 'which model is best?' has no single answer. The useful question is which model is good enough for *this* task at a speed and price the task can bear.
A fourth lever sits beside them: many models let you trade some intelligence for lower latency and cost *within the same model*, through an effort setting. You will meet it again when it comes to tuning a choice. For now, the point is that picking a model is an alignment exercise. You start from what the task demands and choose the tier that meets it.
2.How the tiers trade speed and price against capability
Every Claude model can do coding, agentic work and knowledge work, and Anthropic does not recommend one model class for finance and another for science. What separates the tiers is how hard a problem each can reliably handle and what that costs in price and speed. The comparison table in the models overview shows the ladder.
| Model | How Anthropic describes it | Comparative latency | Input / output price per million tokens |
|---|---|---|---|
| Claude Haiku 4.5 | The fastest model with near-frontier intelligence | Fastest | $1 / $5 |
| Claude Sonnet 5 | The best combination of speed and intelligence | Fast | $2 / $10 |
| Claude Opus 5.5 | For long-running agentic coding and knowledge work | Moderate | $4 / $20 |
| Claude Fable 5.1 | For demanding reasoning and long-horizon agentic work | Slower | $10 / $50 |
Read the rows from top to bottom and two columns move together. Each step up adds reasoning depth, and each step up is slower and more expensive than the one below. Haiku is the lowest-cost, fastest class, built for high-frequency work where latency and cost matter. Sonnet is the versatile everyday class: it balances performance, cost and speed across the widest range of general-purpose work. Opus is for reasoning-heavy work. Fable sits above it, for cases where Opus is shown to struggle.
3.Reading the task: difficulty, latency and volume
With the ladder in mind, Anthropic suggests three questions about the task itself.
How hard is it? Work that usually takes a long time, has many steps or hasn't been solved before calls for a more capable class. What latency does it need? Anthropic's latency guidance names choosing the right model as one of the most direct ways to cut response time, and it points speed-critical applications to Haiku. What are the unit economics? Higher production volumes can suit lower model classes, especially when evaluations show those tasks are done well enough.
The cost-first starting approach lists the situations where a fast, low-cost model is often enough: early prototyping, applications with tight latency requirements, cost-sensitive implementations and high-volume, straightforward tasks. The capability-first approach lists the opposite profile: complex reasoning, scientific or mathematical work, tasks that need nuanced understanding, and cases where accuracy matters more than cost.
A fast, low-cost tier such as Haiku. The deciding requirements are volume (tens of thousands of short, straightforward items) and latency (near real time). The task isn't hard, so extra reasoning depth would add cost and delay without a matching gain in quality.
4.When the strongest tier is the wrong answer
It's tempting to pick the deepest-reasoning tier 'to be safe'. The trade-off table shows the cost of that habit. That tier is the slowest and the most expensive per token, and on a large batch both penalties grow with every item. If the work has a deadline or a budget, the extra reasoning depth can put both at risk while adding little to quality, since the task never needed it.
The reverse mistake is to write off the fast tier as fit only for trivial work. Anthropic's own summary lists Haiku for real-time applications, high-volume intelligent processing and cost-sensitive deployments needing strong reasoning. Speed and low cost don't mean the model can't reason.
The strongest tier earns its place when the task really is hard, when accuracy matters more than cost, or when a cheaper tier has been tried and shown to fall short.
An operations lead has two hundred short, standard customer replies to prepare today and is worried about her weekly allowance. What should she do first?
Correct answer: B — Separate the replies that genuinely need judgment from the routine bulk of the batch.
- A. Incorrect. Applying the most expensive tier to two hundred routine items is the fastest way to exhaust the allowance she is worried about.
- B. Correct. Once the small number of awkward cases is identified, the routine bulk can go to the fastest and cheapest tier and the allowance is reserved for the cases that need it.
- C. Incorrect. Spreading the work delays the replies without reducing the total allowance the batch consumes.
- D. Incorrect. An estimate of consumption does not change what the batch costs or tell her which parts need care.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Choosing the deepest-reasoning tier is always the safe choice, because it gives the best quality.Why is that wrong?
The most capable tier is also the slowest and the most expensive. On high-volume or deadline-bound work that can blow the budget or the time window without a quality gain the task needed.
Covered in When the strongest tier is the wrong answer
2.The fastest, cheapest tier is only suitable for trivial tasks that need no reasoning.Why is that wrong?
Anthropic lists Haiku for high-volume intelligent processing and for cost-sensitive deployments that still need strong reasoning.
Covered in When the strongest tier is the wrong answer
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Speed: How quickly does the model need to respond in your application?”
↩︎ Three questions to ask before you pick a model“Cost: What's your budget for both development and production usage?”
↩︎ Three questions to ask before you pick a model“an effort parameter that trades intelligence for latency and cost within a single model”
↩︎ Three questions to ask before you pick a model“High-volume, straightforward tasks”
↩︎ Reading the task: difficulty, latency and volume“Applications where accuracy outweighs cost considerations”
↩︎ When the strongest tier is the wrong answer“cost-sensitive deployments needing strong reasoning”
↩︎ When the strongest tier is the wrong answer“Real-time applications, high-volume intelligent processing, cost-sensitive deployments needing strong reasoning, sub-agent tasks”
↩︎ Exam trap 2 - 2.https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligenceOfficial docs
“The most capable model can be too expensive at scale, and the least expensive model can fall short on quality.”
↩︎ Three questions to ask before you pick a model“The most capable model can be too expensive at scale, and the least expensive model can fall short on quality.”
↩︎ Key concept“The most capable model can be too expensive at scale, and the least expensive model can fall short on quality.”
↩︎ Exam trap 1 - 3.https://platform.claude.com/docs/en/models/overviewOfficial docs
“Comparative latency | Slower | Moderate | Fast | Fastest”
↩︎ How the tiers trade speed and price against capability“$10 / input MTok$50 / output MTok”
↩︎ When the strongest tier is the wrong answer - 4.https://claude.com/blog/claude-models-explained-choosing-the-best-model-for-your-use-caseSecondary source
“how hard a problem they can reliably carry, and what that capability costs in price and speed”
↩︎ How the tiers trade speed and price against capability“Haiku models are designed for high-frequency workloads where latency and cost matter.”
↩︎ How the tiers trade speed and price against capability“Sonnet provides a balance of performance, cost, and speed for the widest set of general purpose use cases”
↩︎ How the tiers trade speed and price against capability“If it typically takes a lot of time, involves multiple steps, or is previously unsolved then a more capable model class is appropriate.”
↩︎ Reading the task: difficulty, latency and volume“Higher volumes of production may be more appropriate for lower classes of models”
↩︎ Reading the task: difficulty, latency and volume - 5.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-latencyOfficial docs
“One of the most direct ways to reduce latency is to select the appropriate model for your use case.”
↩︎ Reading the task: difficulty, latency and volume“For speed-critical applications, Claude Haiku 4.5 offers the fastest response times while maintaining high intelligence”
↩︎ Reading the task: difficulty, latency and volume