What you will be able to do
- Compare Claude Fable 5.1, Opus 5.5, Sonnet 5 and Haiku 4.5 on capability, latency, price, context window and thinking mode
- Apply the four selection questions (task difficulty, latency, access, unit economics) to a workload
- Judge model cost per completed task instead of per token
- Choose between a Haiku-first and a capability-first starting strategy, and know when to move up to Fable 5.1
Key concept
Cost per completed task — Judge a model by what it costs to finish a task you care about, not by its price per token. A more capable model can cost more per token and still cost less overall, because it needs fewer turns and less thinking to get the task right.
1.The four current models and where they sit
Four current Claude models are open to all customers. They differ on three things: capability, latency and price. The models overview gives a plain default: if you are unsure, start with Claude Opus 5.5 for most workloads. The other three sit around it. Claude Fable 5.1 is for demanding reasoning and long-horizon agentic work, and for cases where your evals on Opus 5.5 at higher effort still fall short. Claude Sonnet 5 is described as the best combination of speed and intelligence. Claude Haiku 4.5 is described as the fastest model with near-frontier intelligence.
| Model | Claude API ID | Latency | Price (input / output per MTok) | Context window | Max output | Thinking | Default effort |
|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | claude-fable-5-1 | Slower | $10 / $50 | 1M tokens | 128K tokens | Adaptive (always on) | high |
| Claude Opus 5.5 | claude-opus-5-5 | Moderate | $4 / $20 | 1M tokens | 128K tokens | Adaptive (always on) | medium |
| Claude Sonnet 5 | claude-sonnet-5 | Fast | $2 / $10 | 1M tokens | 128K tokens | Adaptive | high |
| Claude Haiku 4.5 | claude-haiku-4-5-20251001 | Fastest | $1 / $5 | 200K tokens | 64K tokens | Extended | Not supported |
Two rows besides price matter here. First, Haiku 4.5 has a 200K-token context window and 64K max output. The other three have 1M and 128K. A workload that needs a very large input cannot simply move to the cheapest tier. Second, Haiku 4.5 uses manual extended thinking instead of adaptive thinking, and it does not support the effort parameter. That removes a tuning lever the larger models have.
Some things are the same across all four. Every current model supports text and image input, text output, multilingual capabilities, vision and tool use, so modality rarely forces a choice. Access can, though. Claude Mythos 5.1 has the same capabilities as Fable 5.1 but is available to Project Glasswing participants only.
2.Four questions that decide the tier
Anthropic describes the gap between model classes as how hard a problem each can reliably carry, and what that capability costs in price and speed. That gives four questions to ask about a workload.
How hard is the task? If it usually takes a long time, has many steps, or has not been solved before, use a more capable class.
What latency does it need? For high-frequency, customer-facing work, Sonnet is often the best choice. Haiku models are built for high-frequency workloads where latency and cost matter. Needing speed does not always mean a smaller model, though. Claude Opus 5.5, Opus 5 and Opus 4.8 support fast mode (research preview), which delivers up to 2.5x higher output speed at premium pricing.
What are the access constraints? Mythos is limited to Project Glasswing, and some organizations don't make every model class available to every role.
What are the unit economics? High production volumes may suit lower classes, but only if evaluations show those tasks are completed well enough.
| When you need... | Consider starting with... | Example use cases |
|---|---|---|
| The highest available capability | Claude Fable 5.1 | Agent sessions that run for hours, multistep deep research |
| Complex agentic coding and enterprise work | Claude Opus 5.5 | Multihour autonomous coding agents, large-scale refactoring, computer use |
| Speed and capability for everyday coding, agent, and enterprise workloads | Claude Sonnet 5 | Code generation, data analysis, content creation, agentic tool use |
| The lowest latency and price, with extended thinking | Claude Haiku 4.5 | Real-time applications, high-volume intelligent processing, sub-agent tasks |
It isn't sure to be. The price per token misleads because it leaves out how many tokens a model needs to finish the job. Anthropic notes that more capable models often take fewer turns and less thinking time to get most tasks right. So their cost per task is often lower, especially at lower effort, even when each token costs more. The cost guide makes this the rule for choosing or switching models: compare cost per completed task, not cost per token.
A data engineering team needs to classify and summarize 40 million historical support tickets overnight. There is no interactive user waiting on responses, and the team wants to minimize total API spend while keeping quality reasonably high. Which combination of choices best reflects the intended cost trade-off for this workload?
Correct answer: A — Submit the requests through the Message Batches API, which processes large asynchronous volumes at roughly half the cost of standard synchronous API calls
- A. Correct. Batch processing lets you process large volumes of requests asynchronously for cost savings, with Batch API calls costing 50% less than standard API calls, which directly matches a non-interactive, cost-minimizing overnight workload.
- B. Incorrect. Prompt caching reduces cost and latency when repeated context is reused; disabling it does not itself reduce charges and would forgo available savings on any repeated instructions or context across the ticket batch.
- C. Incorrect. Batch API discounts apply broadly across models, not exclusively to the highest-tier model, and choosing Opus 4.8 for simple classification would add unnecessary per-token cost relative to a lighter model.
- D. Incorrect. Even without latency constraints, maximizing effort on every request increases token usage and cost without necessarily improving results on straightforward classification and summarization tasks, working against the stated goal of minimizing spend.
3.Two ways to start: cheapest-first or capability-first
The model selection docs describe two ways to start. They move in opposite directions.
Cheapest-first. Build with Claude Haiku 4.5, test your use case thoroughly, check whether performance meets your requirements, and upgrade only when a specific capability is missing. This keeps iteration fast and development cheap. It suits prototyping, applications with tight latency requirements, cost-sensitive deployments, and high-volume, straightforward tasks.
Capability-first. For complex tasks where intelligence matters most, build with Claude Opus 5.5 and optimize your prompts for it. Check performance, then cut cost over time by lowering effort or moving to a smaller model. This suits complex reasoning, scientific or mathematical work, tasks that need nuanced understanding, cases where accuracy matters more than cost, and advanced coding or high-autonomy agentic work. There is one step above Opus: if your evals at xhigh or max effort still fall short on demanding reasoning or long-horizon agentic work, move to Claude Fable 5.1.
Anthropic's own default leans capability-first: start with the most intelligent generally available model and use effort to tune performance and cost. Its reason is diagnostic. Starting with a smaller model can make it harder to tell model failures from setup failures. The same logic applies between Opus and Fable. If Opus already meets the quality bar, its speed and price may make it the better choice. Move to Fable only when evals show Opus struggling.
Cheapest-first. The workload is high-volume and straightforward with tight latency needs, which are the cases the docs list for starting with Claude Haiku 4.5. They should upgrade only if testing shows a specific capability gap. Haiku's 200K context window is not a problem for short messages.
A product manager asks an engineer to justify why the team began prototyping a new feature with Claude Haiku 4.5 instead of Claude Opus 4.8, even though Opus 4.8 has higher raw capability. Which of the following are valid justifications drawn from Anthropic's guidance on starting with a fast, cost-effective model? (Select all that apply)(Select 3)
Correct answers: A, B, C — Starting with a faster, more cost-effective model allows for quick iteration during initial prototyping and development; The approach is well suited to applications with tight latency requirements; The approach is well suited to cost-sensitive implementations and high-volume, straightforward tasks
- A. Correct. The documentation states this approach allows for quick iteration, lower development costs, and is often sufficient for many common applications, explicitly listing initial prototyping and development as a best fit.
- B. Correct. Applications with tight latency requirements are explicitly listed as a best fit for starting with a fast, cost-effective model like Haiku 4.5.
- C. Correct. Cost-sensitive implementations and high-volume, straightforward tasks are explicitly listed as best fits for this starting approach.
- D. Incorrect. There is no such access restriction; teams may start directly with Opus 4.8 under the alternative 'start with the most capable model' approach for complex reasoning tasks, with no requirement to benchmark on Haiku first.
- E. Incorrect. The guidance explicitly frames upgrading as necessary when performance does not meet requirements, meaning accuracy is not guaranteed to be equivalent; the recommendation is to test thoroughly and upgrade only if there are specific capability gaps.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The model with the lowest price per token is always the cheapest way to run a workload.Why is that wrong?
What counts is cost per completed task. More capable models often finish in fewer turns with less thinking, so they can cost less per task despite a higher token price.
Covered in Four questions that decide the tier
2.All current models share the same features, so any workload can drop to Haiku 4.5 to save money.Why is that wrong?
They share modalities and tool use, but Haiku 4.5 has a 200K context window against 1M for the others. It also uses extended thinking with no effort parameter.
Covered in The four current models and where they sit
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/models/overviewOfficial docs
“start with Claude Opus 5.5 for most workloads”
↩︎ The four current models and where they sit“All current models support text and image input, text output, multilingual capabilities, vision, and tool use.”
↩︎ The four current models and where they sit“Context window | 1M tokens | 1M tokens | 1M tokens | 200K tokens”
↩︎ Exam trap 2 - 2.https://claude.com/blog/claude-models-explained-choosing-the-best-model-for-your-use-caseSecondary source
“Every Claude model is trained to excel in areas like coding, agentic tasks, and knowledge work.”
↩︎ The four current models and where they sit“how hard a problem they can reliably carry, and what that capability costs in price and speed”
↩︎ Four questions that decide the tier“If the model is involved in high-frequency customer facing workloads, then Sonnet is often the best choice.”
↩︎ Four questions that decide the tier“Starting with a smaller model can also make it harder to distinguish between model failures and setup failures.”
↩︎ Two ways to start: cheapest-first or capability-first“Cost-per-task is often lower for more intelligent models, especially at lower effort levels, even if the price-per-token is higher.”
↩︎ Exam trap 1 - 3.https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligenceOfficial docs
“Compare on cost per completed task, not per token”
↩︎ Four questions that decide the tier“Compare on cost per completed task, not per token”
↩︎ Key concept - 4.
“Begin implementation with Claude Haiku 4.5.”
↩︎ Two ways to start: cheapest-first or capability-first“For complex tasks where intelligence and advanced capabilities are paramount, you may want to start capability-first”
↩︎ Two ways to start: cheapest-first or capability-first“If your evals at xhigh or max effort still fall short on demanding reasoning or long-horizon agentic work, move to Claude Fable 5.1.”
↩︎ Two ways to start: cheapest-first or capability-first