CertSafari
    CLAUDE-CERTIFIED-DEVELOPER-FOUNDATIONS-CCDV-F · Lessons

    Domain 5 · Lesson 14/25

    Claude Effort Levels, Cost Levers and Breaking Changes Across Model Releases

    Model Selection and Tradeoffs

    10 min read
    4.2% of exam
    4 sources
    Published 29 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Pick an effort level and know each model's default
    • Choose the right cost or quality lever for a situation, comparing models on cost per completed task
    • List the Claude Opus 5.5 request settings that now return a 400 error, and what replaces each
    • Plan a model migration that re-baselines effort, max_tokens, cost and latency

    1.Effort: the tradeoff inside one model

    Effort controls how many tokens Claude spends on a response. It applies to every output token: text, tool calls and thinking. So it works with or without thinking, and lower effort also means fewer and shorter tool calls. You set it with output_config.effort.

    Effort levels and their typical use
    LevelDescriptionTypical use case
    maxAbsolute maximum capability, no constraint on token spendingDeepest possible reasoning and most thorough analysis
    xhighExtended capability for long-horizon work (not on every model that supports max)Long-running agentic and coding tasks over 30 minutes
    highAs many tokens as the task needs; default on every effort-supporting model except Claude Opus 5.5Complex reasoning, difficult coding, agentic tasks
    mediumBalanced, moderate token savings; default on Claude Opus 5.5Agentic tasks balancing speed, cost and performance
    lowMost efficient, some capability reductionSimple tasks needing best speed and lowest cost, such as subagents

    Each model has its own recommended effort settings, and they override the general table above:

    - Fable 5.1 and Opus 5: start at high. Step up to xhigh or max for capability-sensitive work, and step down wherever your evals show quality holds. - Opus 5.5: medium is the default. Thinking is always on there, so effort is the main control over how much the model reasons and what a request costs. - Opus 4.8 and 4.7: start at xhigh for coding and agentic work, and use high as the minimum for other intelligence-sensitive work. The API default is still high, so to get xhigh you have to set it explicitly.

    At xhigh or max, set a large max_tokens. It is a hard cap on thinking plus response text, and 64k is a reasonable place to start.

    An engineering org wants to run an autonomous coding agent that will work independently for several hours making large-scale refactors across a big enterprise codebase. Which combination of model and effort setting best fits this workload?

    Sources12

    2.Choosing the right cost lever

    The cost guidance splits levers into two groups:

    - Free wins cut spend without lowering quality: prompt caching, trimming unneeded tokens, a prompt audit, batch processing at 50% off, and spend limits. - Tradeoffs give up some intelligence to save cost: model choice, effort, output caps and task budgets, an elapsed-time clock, and multi-model architectures.

    Model selection is in the second group. That is why it should follow evidence, not a price list.

    Which lever to use in which situation
    Your situationDo this
    Costs are too high; quality is fineSweep effort down on your current model
    You are not on the latest modelUpgrade; each newer model solved at least as many tasks, usually for less per solved task
    You are choosing or switching modelsCompare on cost per completed task, not per token
    Quality isn't good enoughIf you lowered effort, restore it; otherwise try the next tier up at low effort
    You can check outputs (tests, a verifier)Run everything at low effort and re-run failures at high
    A lower-cost model stalls only on hard decisionsAdd a frontier advisor
    The work exceeds one context windowDelegate partitions to cheaper workers

    Two rows need a closer look. A model with a higher price per token can still cost less overall if it finishes more tasks, so compare models on cost per completed task. And when quality is too low, the first fix is to undo an effort cut, not to jump straight to the most expensive tier.

    Multi-model setups pair a cheaper model with a frontier model, so most tokens are billed at the lower rate. There are two common patterns:

    - An executor handles the work and passes hard decisions to an advisor. - An orchestrator hands bulk work out to cheaper workers.

    Anthropic labels its measured results directional, so measure on your own workload.

    Sources32

    3.Breaking changes when moving to a newer model

    A newer model is not always a drop-in replacement. Claude Opus 5.5 is cheaper than Claude Opus 5 ($4 / $20 against $5 / $25 per million input/output tokens), but it rejects several request patterns that earlier models accepted. Each rejected setting returns a 400 error.

    What a request to claude-opus-5-5 must change
    SettingEarlier behaviourOn Claude Opus 5.5
    thinking: {"type": "disabled"} or budget_tokensAccepted on Claude Opus 5 / Claude Sonnet 5 (disabled) or older models (budget)Rejected; adaptive thinking is always on, use effort instead
    Effort defaulthigh on Claude Opus 5medium; set effort explicitly
    tool_choice any or toolCould force a tool callRejected; use auto plus strict tool use or structured outputs
    temperature, top_p, top_kTunableAny non-default value is rejected; guide with prompting
    Prefilled assistant turnAllowedRejected; use structured outputs or system prompt instructions
    computer_20251124Computer use toolRejected on the Claude API and Google Cloud; declare computer_toolset_20260801
    Context-window beta headerNeeded for 1M on older modelsNo effect; 1M is the default
    A request that satisfies every Claude Opus 5.5 requirement, reading text blocks by typepython
    client = anthropic.Anthropic()
    
    response = client.messages.create(
        model="claude-opus-5-5",
        max_tokens=4096,
        messages=[
            {
                "role": "user",
                "content": "Analyze the trade-offs between microservices and monolithic architectures",
            }
        ],
        output_config={"effort": "medium"},
    )
    
    for block in response.content:
        if block.type == "text":
            print(block.text)

    Response parsing can break too:

    - Reading by position. Opus 5.5 always thinks, so a response can start with thinking blocks. Code that reads content[0].text, or treats the first streamed block as text, breaks. Read blocks by their type field instead, as the loop above does. - Tool-use loops. Pass thinking blocks back complete and unmodified, including the empty ones. Edited, reordered or dropped blocks get a 400 error. - Streamed reasoning. Display defaults to "omitted", so a product that shows reasoning to users will look like it is stuck. - Routers and fallbacks. If a conversation can move from Opus 5.5 to another model, expect that model to run without Opus 5.5's thinking blocks. Fable 5.1 and Mythos 5.1 on the Claude API are the exception.

    Some changes alter behaviour without any error. Claude Opus 4.7 follows effort levels more strictly than 4.6, especially at low and medium, where it limits its work to what was asked. If reasoning looks shallow, the docs say to raise effort rather than work around it in the prompt.

    A team upgrades an integration from claude-opus-4-6 to claude-opus-4-8 without changing any code. Their existing calls set thinking: {type: "enabled", budget_tokens: 10000} and now fail with a 400 error. What is the correct fix?

    Sources41

    4.Re-baselining after a switch

    Fixing the request is only half the migration. Settings tuned on the old model don't carry over:

    - Effort. Run a new effort sweep on your own evals instead of reusing the old value. - Cost and latency. Measure both again at the effort level you choose. - max_tokens. Workloads that used to run without thinking now think on every request. Those thinking tokens are billed as output and count against max_tokens, so raise it (start at 64k for xhigh or max). Lower effort if you want less thinking. - Capacity. Plan it separately if you rely on Priority Tier, which Opus 5.5 does not support.

    The model table also lists retirement dates, such as "not sooner than October 15, 2026" for Haiku 4.5. That lets you schedule a migration before a model is retired.

    Sources4

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Leaving effort unset on Claude Opus 5.5 gives the same reasoning depth as it did on Claude Opus 5.Why is that wrong?

      Opus 5.5 defaults to medium, while Opus 5 and earlier Opus models default to high. The same request runs one level lower unless you set effort explicitly.

      Covered in Effort: the tradeoff inside one model

    2. 2.On Claude Opus 5.5 you can still disable thinking at low effort to save tokens, as on Claude Opus 5.Why is that wrong?

      Opus 5.5 rejects thinking: {"type": "disabled"} at every effort level. Lowering effort is the only way to reduce thinking.

      Covered in Breaking changes when moving to a newer model

    3. 3.The model with the lowest price per token is the cheapest choice.Why is that wrong?

      Model comparisons should use cost per completed task. A model that solves more tasks can cost less overall even if its per-token price is higher.

      Covered in Choosing the right cost lever

    Practise it for real

    Watch Claude Opus 5.5's breaking changes happen in a real request

    1. 1.Send the documented claude-opus-5-5 request with output_config={"effort": "medium"} and print only the blocks whose type is "text".

      Why: Shows the recommended request shape and reading blocks by type.

      You should see: A text answer is printed, and response.content also holds a thinking block ahead of it.

    2. 2.Print the type and thinking field of response.content[0].

      Why: Shows why content[0].text is unsafe on this model.

      You should see: A thinking block whose thinking field is empty, because display defaults to "omitted".

    3. 3.Add thinking={"type": "disabled"} to the same request and send it again.

      Why: Confirms that thinking cannot be turned off on Opus 5.5.

      You should see: A 400 error.

    4. 4.Remove the thinking field, and send the request again at effort low and at effort high, comparing output token usage.

      Why: Effort, not a thinking switch, controls reasoning depth and cost here.

      You should see: The high-effort request uses more output tokens than the low-effort one.

    Stuck? Get a nudge

    If step 3 succeeds, check that the model string really is claude-opus-5-5 and not claude-sonnet-5.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “You can trade off between response thoroughness and token efficiency with a single model.”
      ↩︎ Effort: the tradeoff inside one model
      “Because effort applies to every output token, it works whether or not thinking is enabled.”
      ↩︎ Effort: the tradeoff inside one model
      “Claude Opus 4.7 also respects effort levels more strictly than Claude Opus 4.6”
      ↩︎ Breaking changes when moving to a newer model
      “so a request that omits effort runs one level lower than it did on Claude Opus 5”
      ↩︎ Exam trap 1
      “Requests that set thinking: {"type": "disabled"} return a 400 error at every effort level.”
      ↩︎ Exam trap 2
    2. 2.
      “the xhigh effort level, between high and max, is the best setting for most coding and agentic use cases”
      ↩︎ Effort: the tradeoff inside one model
      “Multi-model strategies pair a lower-cost model with a frontier model so that most tokens are billed at the lower rate.”
      ↩︎ Choosing the right cost lever
    3. 3.
      “If you lowered effort, restore it; otherwise try the next tier up at low effort”
      ↩︎ Choosing the right cost lever
      “Compare on cost per completed task, not per token”
      ↩︎ Exam trap 3
    4. 4.
      “Omit temperature, top_p, and top_k, or leave them at their defaults: any other value is rejected.”
      ↩︎ Breaking changes when moving to a newer model
      “Select content blocks by their type field instead”
      ↩︎ Breaking changes when moving to a newer model
      “Run a fresh effort sweep on your own evals rather than carrying over a setting tuned for an earlier model.”
      ↩︎ Re-baselining after a switch
      “Re-baseline cost and latency at your chosen effort level.”
      ↩︎ Re-baselining after a switch

    Ready to test yourself?

    Practise the 20 questions on this subdomain.