CertSafari
    CCAR-P · Lessons

    Domain 2 · Lesson 7/38

    Tuning Claude Model Choice: Effort, Evals and Multi-Model Designs

    Select appropriate Claude models based on trade-offs

    7 min read
    2.6% of exam
    5 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Tell cost levers that don't affect quality apart from levers that trade intelligence for cost
    • Use the effort parameter before switching models, and know each model's default effort
    • Base a model decision on custom evals instead of saturated benchmarks
    • Recognise when an advisor or orchestrator multi-model design pays off

    1.Where model choice sits among cost levers

    Choosing a model is one of several ways to control what a workload costs, and it is not the first one to use. Anthropic's cost guide sorts the levers into two groups. Free wins cut spend without touching quality: prompt caching, token hygiene, auditing the prompt against the model you actually run, batch processing at 50% off for work that can wait up to 24 hours, and workspace spend limits as a backstop. Tradeoffs give up some intelligence to save cost: model choice, effort, output caps and task budgets, an elapsed-time clock, and multi-model architectures.

    This order matters in practice. Moving to a smaller model gives up quality. Caching and trimming tokens do not, and in Anthropic's measurements prompt caching was the largest lever by a wide margin. So when a bill is too high, use the free wins before giving up intelligence.

    Staying on an older model is not a free win either. In the cost guide's measurements, each newer model solved at least as many tasks as the one before it, usually for less per solved task. List prices point the same way. Claude Opus 4.1 is priced at $15 / $75 per MTok, while Claude Opus 5.5 is $4 / $20.

    Sources12

    2.Effort: tuning inside a single model

    Several Claude models support an effort parameter, which trades intelligence for latency and cost within one model. The selection docs say tuning effort is often a better lever than switching models. Default effort differs by model. Fable 5.1 and Sonnet 5 default to high, Opus 5.5 defaults to medium, and Haiku 4.5 does not support effort. For Fable 5.1 and Opus 5, the docs say to start at the default (high) and adjust based on your evals. For Opus 5.5, start at medium and adjust the same way. For Opus 4.8 and Opus 4.7, xhigh (between high and max) is the best setting for most coding and agentic work. On Opus 5.5, adaptive thinking is always on and cannot be turned off, so effort is how you control thinking depth.

    Effort also makes a larger model cheaper to run. Anthropic notes that higher-class models at lower effort can sometimes be more efficient than smaller models.

    Situation-to-action rules for effort and model tier (cost guide)
    Your situationDo this
    Costs are too high; quality is fineSweep effort down on your current model
    Quality isn't good enoughIf you lowered effort, restore it; otherwise try the next tier up at low effort
    You can check outputs (tests, a verifier)Run everything at low effort and re-run failures at high
    You are choosing or switching modelsCompare on cost per completed task, not per token

    Sources341

    3.Deciding with evals, not impressions

    Both starting strategies and every effort adjustment rely on the same thing: an evaluation set built for your use case. The selection docs call it the most important step in the process. Test with your actual prompts and data. Compare models on accuracy, response quality and edge-case handling, then weigh performance against cost.

    Public benchmarks are useful as a rough guide, but they have a known limit at the top tier. Powerful models such as Opus and Fable can solve almost every question on a test, which is called saturation. At that point a benchmark can't tell you which one your workload needs. Anthropic recommends testing on real workloads or your own evaluations instead: a curated set of problems drawn from production, including hard tasks where your current tools fall short, with success criteria your team defines.

    The same caution applies to Anthropic's own cost figures. The guide calls its results internal and directional, not guarantees, and tells you to measure on your own workload.

    Sources35

    4.Using two models instead of choosing one

    You don't always have to pick a single model. Multi-model strategies pair a lower-cost model with a frontier model, so most tokens are billed at the lower rate. There are two common patterns.

    Executor with an advisor. A cheaper model does the work and escalates hard decisions to a frontier model that checks its plan and reviews its output. Use this when the lower-cost model stalls only on hard decisions. The cost guide says it pays off when the advisor is priced well above the executor and is actually consulted. So first price the advisor's model on its own at low effort, and measure how often it gets consulted. Anthropic reports that on SWE-bench Pro, Sonnet 5 with a Fable 5 advisor came within 10% of Fable 5's score at 63% of the price of running Fable 5 for the whole task.

    Orchestrator with workers. A frontier model hands bulk work to lower-cost workers. Use this when the work is too big for one context window: split it into parts and give each part to a cheaper worker. This is the sub-agent role the tier descriptions assign to Sonnet (high-volume sub-agents in multi-agent setups) and Haiku 4.5 (sub-agent tasks).

    The cost guide calls these multi-model levers narrower than caching. They pay off in these two specific shapes, not as a default.

    A developer relations team is building an internal tool that programmatically decides which model to route a request to based on declared capabilities rather than hardcoding a model name. They want to know, before sending a request, the maximum input tokens, maximum output tokens, and capability flags of each candidate model. Which platform mechanism should they use to retrieve this information programmatically?

    Sources315

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.When a Claude workload costs too much, the first fix is to move to a smaller model.Why is that wrong?

      If quality is fine, lower the effort on the current model first. Anthropic says tuning effort is often a better lever than switching models.

      Covered in Effort: tuning inside a single model

    2. 2.Staying on an older Claude model is a safe way to save money.Why is that wrong?

      In Anthropic's measurements, newer models solved at least as many tasks, usually for less per solved task, so the cost guide says to upgrade.

      Covered in Where model choice sits among cost levers

    3. 3.Public benchmark scores are enough to choose between Opus and Fable.Why is that wrong?

      Top models saturate public benchmarks. Use custom evals drawn from your own production workload to separate them.

      Covered in Deciding with evals, not impressions

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Free wins cut spend without touching quality”
      ↩︎ Where model choice sits among cost levers
      “Tradeoffs exchange cost for intelligence: model choice, effort, output caps and task budgets, an elapsed-time clock, and multi-model architectures.”
      ↩︎ Where model choice sits among cost levers
      “Run everything at low effort and re-run failures at high”
      ↩︎ Effort: tuning inside a single model
      “It pays off when priced well above the executor and actually consulted”
      ↩︎ Using two models instead of choosing one
      “Delegate partitions to cheaper workers”
      ↩︎ Using two models instead of choosing one
      “each newer model solved at least as many tasks as the one before it, usually for less per solved task”
      ↩︎ Exam trap 2
    2. 3.
      “Tuning effort is often a better lever than switching models.”
      ↩︎ Effort: tuning inside a single model
      “having a good evaluation set is the most important step in the process”
      ↩︎ Deciding with evals, not impressions
      “Multi-model strategies pair a lower-cost model with a frontier model so that most tokens are billed at the lower rate.”
      ↩︎ Using two models instead of choosing one
      “Tuning effort is often a better lever than switching models.”
      ↩︎ Exam trap 1
    3. 4.
    4. 5.
      “which can solve almost all of the questions on the test (often referred to as saturation)”
      ↩︎ Deciding with evals, not impressions
      “The advisor strategy allows faster, lower-cost worker models to call more intelligent models to check their plan and evaluate their work”
      ↩︎ Using two models instead of choosing one
      “which can solve almost all of the questions on the test (often referred to as saturation)”
      ↩︎ Exam trap 3

    Ready to test yourself?

    Practise the 12 questions on this subdomain.