CertSafari
    CLAUDE-CERTIFIED-DEVELOPER-FOUNDATIONS-CCDV-F · Lessons

    Domain 5 · Lesson 14/25

    Choosing a Claude Model: Fable, Opus, Sonnet or Haiku

    Model Selection and Tradeoffs

    7 min read
    4.2% of exam
    3 sources
    Published 29 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Place Claude Fable 5.1, Opus 5.5, Sonnet 5 and Haiku 4.5 on the quality, latency and cost tradeoff, and name a typical use case for each
    • Choose between the cost-first (Haiku) and capability-first (Opus) ways of starting a project
    • Tell which models use adaptive thinking, which use extended thinking, and which let you turn thinking off
    • Set up an evaluation that decides whether a model change is worth it

    Key concept

    Model tier vs. effort — The model you pick sets the capability ceiling, the speed and the price per token. Effort is a setting inside one model that trades intelligence against latency and cost. Changing effort is often the better move than changing models.

    1.Four factors and four tiers

    Every model choice trades quality against latency and cost. Anthropic asks you to weigh four factors: the capabilities the task needs, how fast the application must respond, your budget for development and for production, and effort. Effort is a parameter on several models that trades intelligence for latency and cost without changing the model.

    The current models spread across that range. All of them handle text and image input, text output, multiple languages, vision and tool use. They differ in speed, price, thinking mode, context window and maximum output.

    Current Claude models compared on the quality/latency/cost tradeoff
    Model (API ID)Comparative latencyPrice (input / output per MTok)ThinkingDefault effortContext window / max output
    Claude Fable 5.1 (claude-fable-5-1)Slower$10 / $50Adaptive (always on)high1M / 128K
    Claude Opus 5.5 (claude-opus-5-5)Moderate$4 / $20Adaptive (always on)medium1M / 128K
    Claude Sonnet 5 (claude-sonnet-5)Fast$2 / $10Adaptivehigh1M / 128K
    Claude Haiku 4.5 (claude-haiku-4-5-20251001)Fastest$1 / $5ExtendedNot supported200K / 64K

    By default, start with Claude Opus 5.5. Move to Claude Fable 5.1 for demanding reasoning and long-horizon agentic work, or when Opus 5.5 still falls short on your evals at higher effort. The documented use cases for each tier:

    - Fable 5.1: agent sessions that run for hours, and multistep deep research. - Opus 5.5: multihour autonomous coding agents, large-scale refactoring and computer use. - Sonnet 5: everyday code generation, data analysis and content creation. - Haiku 4.5: the lowest latency and price, for real-time applications, high-volume processing and sub-agent tasks.

    For latency there is also fast mode. Claude Opus 5.5, Claude Opus 5 and Claude Opus 4.8 support it as a research preview, which gives up to 2.5x higher output speed at premium pricing.

    Sources12

    2.Cost-first or capability-first

    The docs describe two ways to start. Which one fits depends on how hard the task is and on what a wrong answer costs.

    The two documented ways to start model selection
    ApproachStepsBest for
    Cost-first with Claude Haiku 4.5Begin with Haiku 4.5, test thoroughly, evaluate, upgrade only if necessary for specific capability gapsInitial prototyping, tight latency requirements, cost-sensitive implementations, high-volume straightforward tasks
    Capability-first with Claude Opus 5.5Implement with Opus 5.5, optimize prompts, evaluate, then lower effort or downgrade models over time; move to Claude Fable 5.1 if xhigh or max effort still falls shortComplex reasoning, scientific or mathematical applications, nuanced understanding, accuracy over cost, advanced coding and high-autonomy agentic work

    The two approaches move in opposite directions. Cost-first starts cheap and upgrades only when a specific capability gap shows up. Capability-first starts strong and saves money later, first by lowering effort and then by moving to a smaller model. When accuracy matters more than cost, for example in scientific or mathematical work, the guidance points to capability-first.

    A team is building a real-time customer support chat widget that must respond in under a second, at high volume, on a tight budget, but still needs solid reasoning quality. Which model should they start with?

    Sources2

    3.Adaptive vs. extended thinking by model

    Thinking lets Claude work through a problem before it answers. It restates the question, tries approaches, checks intermediate results and drops paths that don't work. That helps on tasks where the quality of the answer depends on intermediate work. Thinking also costs money: reasoning tokens are billed as output tokens and count toward max_tokens.

    The current models don't all handle thinking the same way:

    - Adaptive thinking lets Claude decide when to think and how deeply. - Extended thinking is configured by hand with a fixed token budget. Haiku 4.5 is the current model that uses it. - On some models thinking is always on. On others it is on by default and can be turned off, or off until you ask for it.

    Thinking configuration across models
    ModelThinking behaviourCan you disable it?
    Claude Fable 5.1, Claude Opus 5.5Adaptive, always onNo: thinking: {type: "disabled"} is rejected
    Claude Opus 5On by defaultYes at effort high or below; 400 error at xhigh or max
    Claude Sonnet 5Adaptive, on by defaultYes, with thinking: {type: "disabled"}
    Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Sonnet 4.6Off until you set thinking: {type: "adaptive"}Off by default
    Claude Haiku 4.5Extended: type: "enabled" with budget_tokensThinking is configured explicitly
    Turning thinking off on Claude Sonnet 5 for a simple taskpython
    client = anthropic.Anthropic()
    
    response = client.messages.create(
        model="claude-sonnet-5",
        max_tokens=4096,
        thinking={"type": "disabled"},
        messages=[{"role": "user", "content": "Summarize this article in one sentence."}],
    )

    The display setting also affects latency. On Opus 5.5, Sonnet 5 and Fable 5.1, display defaults to "omitted", so thinking blocks come back with an empty text field. When you stream, this gets the first text token out sooner, because the server skips streaming the thinking tokens. You pay for them either way.

    Sources3

    4.Letting evals make the decision

    Tables and use-case lists only give you a starting point. To decide whether to upgrade or change models:

    1. Build a benchmark for your own use case. 2. Run it with your real prompts and data. 3. Compare the models on accuracy, response quality and edge cases. 4. Weigh that performance against cost.

    Without a good eval set, you can't tell whether a cheaper tier holds up or a pricier one is worth it.

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Every current Claude model uses adaptive thinking and accepts an effort level.Why is that wrong?

      Claude Haiku 4.5 uses extended thinking, set with type "enabled" and budget_tokens, and does not support effort. Fable 5.1, Opus 5.5 and Sonnet 5 use adaptive thinking.

      Covered in Adaptive vs. extended thinking by model

    2. 2.If thinking text is hidden (display "omitted"), the thinking is free.Why is that wrong?

      Thinking tokens are billed as output tokens whether or not the text is returned. Omitting it speeds up the first streamed text token but does not reduce cost.

      Covered in Adaptive vs. extended thinking by model

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “If you're unsure which model to use, start with Claude Opus 5.5 for most workloads.”
      ↩︎ Four factors and four tiers
      “All current models support text and image input, text output, multilingual capabilities, vision, and tool use.”
      ↩︎ Four factors and four tiers
      “Thinking | Adaptive (always on) | Adaptive (always on) | Adaptive | Extended”
      ↩︎ Exam trap 1
    2. 2.
      “support fast mode (research preview), which delivers up to 2.5x higher output speed at premium pricing”
      ↩︎ Four factors and four tiers
      “For many applications, starting with a faster, more cost-effective model like Claude Haiku 4.5 can be the optimal approach”
      ↩︎ Cost-first or capability-first
      “For complex tasks where intelligence and advanced capabilities are paramount, you may want to start capability-first”
      ↩︎ Cost-first or capability-first
      “Create benchmark tests specific to your use case - having a good evaluation set is the most important step in the process.”
      ↩︎ Letting evals make the decision
      “Tuning effort is often a better lever than switching models.”
      ↩︎ Key concept
    3. 3.
      “This is why thinking improves performance on complex tasks like math, coding, analysis, and long-running agentic work”
      ↩︎ Adaptive vs. extended thinking by model
      “configure it with type: "enabled" and a budget_tokens value instead”
      ↩︎ Adaptive vs. extended thinking by model
      “the tokens Claude spends reasoning are billed as output tokens, even when the thinking text isn't returned to you”
      ↩︎ Exam trap 2

    Continue to page 2 of 2

    Claude Effort Levels, Cost Levers and Breaking Changes Across Model Releases