What you will be able to do
- Place Claude Fable 5.1, Opus 5.5, Sonnet 5 and Haiku 4.5 on the quality, latency and cost tradeoff, and name a typical use case for each
- Choose between the cost-first (Haiku) and capability-first (Opus) ways of starting a project
- Tell which models use adaptive thinking, which use extended thinking, and which let you turn thinking off
- Set up an evaluation that decides whether a model change is worth it
Key concept
Model tier vs. effort — The model you pick sets the capability ceiling, the speed and the price per token. Effort is a setting inside one model that trades intelligence against latency and cost. Changing effort is often the better move than changing models.
1.Four factors and four tiers
Every model choice trades quality against latency and cost. Anthropic asks you to weigh four factors: the capabilities the task needs, how fast the application must respond, your budget for development and for production, and effort. Effort is a parameter on several models that trades intelligence for latency and cost without changing the model.
The current models spread across that range. All of them handle text and image input, text output, multiple languages, vision and tool use. They differ in speed, price, thinking mode, context window and maximum output.
| Model (API ID) | Comparative latency | Price (input / output per MTok) | Thinking | Default effort | Context window / max output |
|---|---|---|---|---|---|
| Claude Fable 5.1 (claude-fable-5-1) | Slower | $10 / $50 | Adaptive (always on) | high | 1M / 128K |
| Claude Opus 5.5 (claude-opus-5-5) | Moderate | $4 / $20 | Adaptive (always on) | medium | 1M / 128K |
| Claude Sonnet 5 (claude-sonnet-5) | Fast | $2 / $10 | Adaptive | high | 1M / 128K |
| Claude Haiku 4.5 (claude-haiku-4-5-20251001) | Fastest | $1 / $5 | Extended | Not supported | 200K / 64K |
By default, start with Claude Opus 5.5. Move to Claude Fable 5.1 for demanding reasoning and long-horizon agentic work, or when Opus 5.5 still falls short on your evals at higher effort. The documented use cases for each tier:
- Fable 5.1: agent sessions that run for hours, and multistep deep research. - Opus 5.5: multihour autonomous coding agents, large-scale refactoring and computer use. - Sonnet 5: everyday code generation, data analysis and content creation. - Haiku 4.5: the lowest latency and price, for real-time applications, high-volume processing and sub-agent tasks.
For latency there is also fast mode. Claude Opus 5.5, Claude Opus 5 and Claude Opus 4.8 support it as a research preview, which gives up to 2.5x higher output speed at premium pricing.
Claude Haiku 4.5. The guidance lists it for the lowest latency and price, and for real-time, high-volume and cost-sensitive work. Opus 5.5 is the default for most workloads, but this task is simple, runs at high volume and needs low latency, which is exactly the profile where Haiku is the recommended starting point.
2.Cost-first or capability-first
The docs describe two ways to start. Which one fits depends on how hard the task is and on what a wrong answer costs.
| Approach | Steps | Best for |
|---|---|---|
| Cost-first with Claude Haiku 4.5 | Begin with Haiku 4.5, test thoroughly, evaluate, upgrade only if necessary for specific capability gaps | Initial prototyping, tight latency requirements, cost-sensitive implementations, high-volume straightforward tasks |
| Capability-first with Claude Opus 5.5 | Implement with Opus 5.5, optimize prompts, evaluate, then lower effort or downgrade models over time; move to Claude Fable 5.1 if xhigh or max effort still falls short | Complex reasoning, scientific or mathematical applications, nuanced understanding, accuracy over cost, advanced coding and high-autonomy agentic work |
The two approaches move in opposite directions. Cost-first starts cheap and upgrades only when a specific capability gap shows up. Capability-first starts strong and saves money later, first by lowering effort and then by moving to a smaller model. When accuracy matters more than cost, for example in scientific or mathematical work, the guidance points to capability-first.
A team is building a real-time customer support chat widget that must respond in under a second, at high volume, on a tight budget, but still needs solid reasoning quality. Which model should they start with?
Correct answer: A — Claude Haiku 4.5, for near-frontier reasoning at the fastest, most economical tier
- A. Correct. Claude Haiku 4.5 is positioned as the fastest, most economical model with near-frontier intelligence, matching real-time, high-volume, cost-sensitive chat requirements.
- B. Incorrect. Opus 4.8 targets complex agentic coding and enterprise work at moderate latency and higher cost, which is more than this low-latency, high-volume workload needs.
- C. Incorrect. Sonnet 5 is fast and capable, but it costs and latencies more than necessary for a simple, high-volume support widget compared to Haiku 4.5.
- D. Incorrect. Claude Fable 5 is built for long-running agents with large context needs and is priced and latency-tuned for a very different workload than a real-time chat widget.
Sources2
3.Adaptive vs. extended thinking by model
Thinking lets Claude work through a problem before it answers. It restates the question, tries approaches, checks intermediate results and drops paths that don't work. That helps on tasks where the quality of the answer depends on intermediate work. Thinking also costs money: reasoning tokens are billed as output tokens and count toward max_tokens.
The current models don't all handle thinking the same way:
- Adaptive thinking lets Claude decide when to think and how deeply. - Extended thinking is configured by hand with a fixed token budget. Haiku 4.5 is the current model that uses it. - On some models thinking is always on. On others it is on by default and can be turned off, or off until you ask for it.
| Model | Thinking behaviour | Can you disable it? |
|---|---|---|
| Claude Fable 5.1, Claude Opus 5.5 | Adaptive, always on | No: thinking: {type: "disabled"} is rejected |
| Claude Opus 5 | On by default | Yes at effort high or below; 400 error at xhigh or max |
| Claude Sonnet 5 | Adaptive, on by default | Yes, with thinking: {type: "disabled"} |
| Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Sonnet 4.6 | Off until you set thinking: {type: "adaptive"} | Off by default |
| Claude Haiku 4.5 | Extended: type: "enabled" with budget_tokens | Thinking is configured explicitly |
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-sonnet-5",
max_tokens=4096,
thinking={"type": "disabled"},
messages=[{"role": "user", "content": "Summarize this article in one sentence."}],
)The display setting also affects latency. On Opus 5.5, Sonnet 5 and Fable 5.1, display defaults to "omitted", so thinking blocks come back with an empty text field. When you stream, this gets the first text token out sooner, because the server skips streaming the thinking tokens. You pay for them either way.
Sources3
4.Letting evals make the decision
Tables and use-case lists only give you a starting point. To decide whether to upgrade or change models:
1. Build a benchmark for your own use case. 2. Run it with your real prompts and data. 3. Compare the models on accuracy, response quality and edge cases. 4. Weigh that performance against cost.
Without a good eval set, you can't tell whether a cheaper tier holds up or a pricier one is worth it.
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Every current Claude model uses adaptive thinking and accepts an effort level.Why is that wrong?
Claude Haiku 4.5 uses extended thinking, set with type "enabled" and budget_tokens, and does not support effort. Fable 5.1, Opus 5.5 and Sonnet 5 use adaptive thinking.
Covered in Adaptive vs. extended thinking by model
2.If thinking text is hidden (display "omitted"), the thinking is free.Why is that wrong?
Thinking tokens are billed as output tokens whether or not the text is returned. Omitting it speeds up the first streamed text token but does not reduce cost.
Covered in Adaptive vs. extended thinking by model
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/models/overviewOfficial docs
“If you're unsure which model to use, start with Claude Opus 5.5 for most workloads.”
↩︎ Four factors and four tiers“All current models support text and image input, text output, multilingual capabilities, vision, and tool use.”
↩︎ Four factors and four tiers“Thinking | Adaptive (always on) | Adaptive (always on) | Adaptive | Extended”
↩︎ Exam trap 1 - 2.
“support fast mode (research preview), which delivers up to 2.5x higher output speed at premium pricing”
↩︎ Four factors and four tiers“For many applications, starting with a faster, more cost-effective model like Claude Haiku 4.5 can be the optimal approach”
↩︎ Cost-first or capability-first“For complex tasks where intelligence and advanced capabilities are paramount, you may want to start capability-first”
↩︎ Cost-first or capability-first“Create benchmark tests specific to your use case - having a good evaluation set is the most important step in the process.”
↩︎ Letting evals make the decision“Tuning effort is often a better lever than switching models.”
↩︎ Key concept - 3.
“This is why thinking improves performance on complex tasks like math, coding, analysis, and long-running agentic work”
↩︎ Adaptive vs. extended thinking by model“configure it with type: "enabled" and a budget_tokens value instead”
↩︎ Adaptive vs. extended thinking by model“the tokens Claude spends reasoning are billed as output tokens, even when the thinking text isn't returned to you”
↩︎ Exam trap 2