What you will be able to do
- Tell cost levers that don't affect quality apart from levers that trade intelligence for cost
- Use the effort parameter before switching models, and know each model's default effort
- Base a model decision on custom evals instead of saturated benchmarks
- Recognise when an advisor or orchestrator multi-model design pays off
1.Where model choice sits among cost levers
Choosing a model is one of several ways to control what a workload costs, and it is not the first one to use. Anthropic's cost guide sorts the levers into two groups. Free wins cut spend without touching quality: prompt caching, token hygiene, auditing the prompt against the model you actually run, batch processing at 50% off for work that can wait up to 24 hours, and workspace spend limits as a backstop. Tradeoffs give up some intelligence to save cost: model choice, effort, output caps and task budgets, an elapsed-time clock, and multi-model architectures.
This order matters in practice. Moving to a smaller model gives up quality. Caching and trimming tokens do not, and in Anthropic's measurements prompt caching was the largest lever by a wide margin. So when a bill is too high, use the free wins before giving up intelligence.
Staying on an older model is not a free win either. In the cost guide's measurements, each newer model solved at least as many tasks as the one before it, usually for less per solved task. List prices point the same way. Claude Opus 4.1 is priced at $15 / $75 per MTok, while Claude Opus 5.5 is $4 / $20.
2.Effort: tuning inside a single model
Several Claude models support an effort parameter, which trades intelligence for latency and cost within one model. The selection docs say tuning effort is often a better lever than switching models. Default effort differs by model. Fable 5.1 and Sonnet 5 default to high, Opus 5.5 defaults to medium, and Haiku 4.5 does not support effort. For Fable 5.1 and Opus 5, the docs say to start at the default (high) and adjust based on your evals. For Opus 5.5, start at medium and adjust the same way. For Opus 4.8 and Opus 4.7, xhigh (between high and max) is the best setting for most coding and agentic work. On Opus 5.5, adaptive thinking is always on and cannot be turned off, so effort is how you control thinking depth.
Effort also makes a larger model cheaper to run. Anthropic notes that higher-class models at lower effort can sometimes be more efficient than smaller models.
| Your situation | Do this |
|---|---|
| Costs are too high; quality is fine | Sweep effort down on your current model |
| Quality isn't good enough | If you lowered effort, restore it; otherwise try the next tier up at low effort |
| You can check outputs (tests, a verifier) | Run everything at low effort and re-run failures at high |
| You are choosing or switching models | Compare on cost per completed task, not per token |
Sweep effort down on the current model. Quality is fine and only cost is the problem, and effort is often a better lever than switching models. If the outputs can be checked, another option is to run at low effort and re-run only the failures at high. On the coding benchmark Anthropic measured, that kept the pass rate at about half the cost.
3.Deciding with evals, not impressions
Both starting strategies and every effort adjustment rely on the same thing: an evaluation set built for your use case. The selection docs call it the most important step in the process. Test with your actual prompts and data. Compare models on accuracy, response quality and edge-case handling, then weigh performance against cost.
Public benchmarks are useful as a rough guide, but they have a known limit at the top tier. Powerful models such as Opus and Fable can solve almost every question on a test, which is called saturation. At that point a benchmark can't tell you which one your workload needs. Anthropic recommends testing on real workloads or your own evaluations instead: a curated set of problems drawn from production, including hard tasks where your current tools fall short, with success criteria your team defines.
The same caution applies to Anthropic's own cost figures. The guide calls its results internal and directional, not guarantees, and tells you to measure on your own workload.
4.Using two models instead of choosing one
You don't always have to pick a single model. Multi-model strategies pair a lower-cost model with a frontier model, so most tokens are billed at the lower rate. There are two common patterns.
Executor with an advisor. A cheaper model does the work and escalates hard decisions to a frontier model that checks its plan and reviews its output. Use this when the lower-cost model stalls only on hard decisions. The cost guide says it pays off when the advisor is priced well above the executor and is actually consulted. So first price the advisor's model on its own at low effort, and measure how often it gets consulted. Anthropic reports that on SWE-bench Pro, Sonnet 5 with a Fable 5 advisor came within 10% of Fable 5's score at 63% of the price of running Fable 5 for the whole task.
Orchestrator with workers. A frontier model hands bulk work to lower-cost workers. Use this when the work is too big for one context window: split it into parts and give each part to a cheaper worker. This is the sub-agent role the tier descriptions assign to Sonnet (high-volume sub-agents in multi-agent setups) and Haiku 4.5 (sub-agent tasks).
The cost guide calls these multi-model levers narrower than caching. They pay off in these two specific shapes, not as a default.
A developer relations team is building an internal tool that programmatically decides which model to route a request to based on declared capabilities rather than hardcoding a model name. They want to know, before sending a request, the maximum input tokens, maximum output tokens, and capability flags of each candidate model. Which platform mechanism should they use to retrieve this information programmatically?
Correct answer: A — Query the Models API, which returns max_input_tokens, max_tokens, and a capabilities object for every available model
- A. Correct. The documentation states you can query model capabilities and token limits programmatically with the Models API, and that the response includes max_input_tokens, max_tokens, and a capabilities object for every available model.
- B. Incorrect. Capability metadata such as token limits is not embedded in or parseable from a model's system prompt; it is exposed through the dedicated Models API.
- C. Incorrect. Inferring capabilities from response latency is not a documented or reliable mechanism; effort affects latency and cost but does not reveal structured capability metadata like token limits.
- D. Incorrect. The Admin API's usage and cost tooling reports billing and usage data, not structured per-model capability metadata like context window or output token limits; that is the Models API's role.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.When a Claude workload costs too much, the first fix is to move to a smaller model.Why is that wrong?
If quality is fine, lower the effort on the current model first. Anthropic says tuning effort is often a better lever than switching models.
Covered in Effort: tuning inside a single model
2.Staying on an older Claude model is a safe way to save money.Why is that wrong?
In Anthropic's measurements, newer models solved at least as many tasks, usually for less per solved task, so the cost guide says to upgrade.
Covered in Where model choice sits among cost levers
3.Public benchmark scores are enough to choose between Opus and Fable.Why is that wrong?
Top models saturate public benchmarks. Use custom evals drawn from your own production workload to separate them.
Covered in Deciding with evals, not impressions
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligenceOfficial docs
“Free wins cut spend without touching quality”
↩︎ Where model choice sits among cost levers“Tradeoffs exchange cost for intelligence: model choice, effort, output caps and task budgets, an elapsed-time clock, and multi-model architectures.”
↩︎ Where model choice sits among cost levers“Run everything at low effort and re-run failures at high”
↩︎ Effort: tuning inside a single model“It pays off when priced well above the executor and actually consulted”
↩︎ Using two models instead of choosing one“Delegate partitions to cheaper workers”
↩︎ Using two models instead of choosing one“each newer model solved at least as many tasks as the one before it, usually for less per solved task”
↩︎ Exam trap 2 - 2.
“Claude Opus 4.1 | $15 / MTok | $75 / MTok”
↩︎ Where model choice sits among cost levers - 3.
“Tuning effort is often a better lever than switching models.”
↩︎ Effort: tuning inside a single model“having a good evaluation set is the most important step in the process”
↩︎ Deciding with evals, not impressions“Multi-model strategies pair a lower-cost model with a frontier model so that most tokens are billed at the lower rate.”
↩︎ Using two models instead of choosing one“Tuning effort is often a better lever than switching models.”
↩︎ Exam trap 1 - 4.
“Control thinking depth with the effort parameter.”
↩︎ Effort: tuning inside a single model - 5.https://claude.com/blog/claude-models-explained-choosing-the-best-model-for-your-use-caseSecondary source
“which can solve almost all of the questions on the test (often referred to as saturation)”
↩︎ Deciding with evals, not impressions“The advisor strategy allows faster, lower-cost worker models to call more intelligent models to check their plan and evaluate their work”
↩︎ Using two models instead of choosing one“which can solve almost all of the questions on the test (often referred to as saturation)”
↩︎ Exam trap 3