CertSafari
    CLAUDE-CERTIFIED-ASSOCIATE-FOUNDATIONS-CCAO-F-VAR5 · Lessons

    Domain 3 · Lesson 13/30

    Choosing a starting Claude model, then testing and tuning it

    Align model selection with task requirements (cost, speed, quality)

    7 min read
    3% of exam
    4 sources
    Published 28 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Choose between the cost-first and capability-first starting approaches for a new task
    • Test a model choice against your own prompts, data and success criteria
    • Compare tiers on cost per completed task rather than price per token
    • Diagnose a weak output before reaching for a higher tier, and know when pairing tiers makes sense

    1.Two ways to make the first choice

    A model choice is a starting hypothesis that you then test. Anthropic documents two directions for making it, and the right one depends on the task.

    The two starting approaches in Anthropic's model-selection guidance
    ApproachSequenceBest suited to
    Cost-first: start fast and cheapBegin with Haiku, test the use case, evaluate against requirements, upgrade only for specific capability gapsPrototyping, tight latency requirements, cost-sensitive implementations, high-volume straightforward tasks
    Capability-first: start strongImplement with the strongest starting point, optimise prompts, evaluate, then lower effort or downgrade over timeComplex reasoning, scientific or mathematical work, nuanced understanding, accuracy over cost

    Both approaches share the same core loop: run the task, check the result against your requirements, and move a step only when the evidence says to. The cost-first path steps up when it finds a capability gap. The capability-first path steps down once it knows the task well enough to save money.

    Anthropic's own default, set out on its blog, leans capability-first. Start with the most intelligent generally available model and use effort to dial in performance and cost. Then test lower tiers as latency- or cost-sensitive use cases come up. One reason it gives is diagnostic: starting with a smaller model can make it harder to tell a model failure from a setup failure. Both directions are valid. The task's difficulty and its volume decide which one to use.

    Sources12

    2.Test on your own work, not on reputation

    Whichever way you start, the decision to change models should come from evidence. The guidance is explicit that a good evaluation set comes first. Build tests specific to your use case, run them with your actual prompts and data, compare models on accuracy, response quality and edge cases, then weigh performance against cost.

    Speed and cost belong in those success criteria alongside quality. Anthropic's advice on writing evaluations asks what response time is acceptable and what the budget is, noting that cost depends on the size of the model and on how often it's used. Public benchmarks give a direction but become less useful at the top: the most powerful models can solve almost every question on a standard test. At that point, a curated set of problems drawn from your own production work, with success criteria your team defines, is what separates the options.

    A marketing team wants to move its weekly performance report from the deepest-reasoning model to a lighter one to save allowance. What should they do first?

    Sources132

    3.Cost per task, effort, and the vague first draft

    The obvious way to compare tiers is by price per token, and it can mislead. Anthropic's cost guidance says to compare on cost per *completed* task. A more capable model often gets the task right in fewer turns with less thinking time, so its cost per task can be lower even when its price per token is higher.

    Before switching tier, check the lever inside the model. Effort trades intelligence for latency and cost within a single model, and the guidance says tuning it is often a better lever than switching models. The cost playbook turns this into simple rules. If costs are too high but quality is fine, lower effort on your current model. If quality isn't good enough and you had lowered effort, restore it. Only after that should you try the next tier up at low effort.

    Sources421

    4.Using two tiers instead of one

    The choice doesn't have to be one model for everything. Multi-model strategies pair a lower-cost model with a frontier model so that most of the work is billed at the lower rate. There are two common patterns. In the first, an executor escalates hard decisions to an advisor model. In the second, an orchestrator delegates bulk work to lower-cost workers. The cost playbook ties each to a symptom. Add an advisor when a cheaper model stalls only on hard decisions. Use cheaper workers when the work exceeds one context window.

    Anthropic reports that Sonnet 5 with a Fable 5 advisor came within 10% of Fable 5's SWE-bench Pro score at 63% of the price of running Fable 5 for the whole task. It also says its measured results are directional and not guarantees, so measure on your own workload.

    An education lead is choosing a model tier to help draft feedback that will contribute to students' marks. What should she do first?

    Sources142

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.The cheapest model per token is always the cheapest way to get the work done.Why is that wrong?

      What matters is cost per completed task. A more capable model can finish in fewer turns and so cost less per task despite a higher token price.

      Covered in Cost per task, effort, and the vague first draft

    2. 2.A weak or vague output means you should move straight to a deeper-reasoning tier.Why is that wrong?

      Check the setup and the effort level first: restore lowered effort, and only then try the next tier up at low effort.

      Covered in Cost per task, effort, and the vague first draft

    3. 3.Public benchmark rankings are enough to decide which tier a task needs.Why is that wrong?

      Build task-specific tests and run them with your actual prompts and data. A good evaluation set is the most important step.

      Covered in Test on your own work, not on reputation

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Upgrade only if necessary for specific capability gaps.”
      ↩︎ Two ways to make the first choice
      “Consider increasing efficiency by lowering effort or downgrading models over time with greater workflow optimization.”
      ↩︎ Two ways to make the first choice
      “Create benchmark tests specific to your use case - having a good evaluation set is the most important step in the process.”
      ↩︎ Test on your own work, not on reputation
      “Test with your actual prompts and data.”
      ↩︎ Test on your own work, not on reputation
      “Tuning effort is often a better lever than switching models.”
      ↩︎ Cost per task, effort, and the vague first draft
      “Multi-model strategies pair a lower-cost model with a frontier model so that most tokens are billed at the lower rate.”
      ↩︎ Using two tiers instead of one
      “having a good evaluation set is the most important step in the process.”
      ↩︎ Exam trap 3
    2. 2.
      “start with the most intelligent generally available model and use effort level to dial in performance and cost.”
      ↩︎ Two ways to make the first choice
      “Starting with a smaller model can also make it harder to distinguish between model failures and setup failures.”
      ↩︎ Two ways to make the first choice
      “which can solve almost all of the questions on the test (often referred to as saturation).”
      ↩︎ Test on your own work, not on reputation
      “This is because more capable models often take fewer turns and less thinking time to get most tasks right.”
      ↩︎ Cost per task, effort, and the vague first draft
      “is within 10% of Fable 5’s score at 63% of the price of using Fable 5 for the whole task.”
      ↩︎ Using two tiers instead of one
    3. 3.
      “Consider factors like the cost for each API call, the size of the model, and the frequency of usage.”
      ↩︎ Test on your own work, not on reputation
    4. 4.
      “Compare on cost per completed task, not per token”
      ↩︎ Cost per task, effort, and the vague first draft
      “Sweep effort down on your current model”
      ↩︎ Cost per task, effort, and the vague first draft
      “If you lowered effort, restore it; otherwise try the next tier up at low effort”
      ↩︎ Cost per task, effort, and the vague first draft
      “A lower-cost model stalls only on hard decisions | Add a frontier advisor.”
      ↩︎ Using two tiers instead of one
      “directional, not guarantees, so measure on your own workload”
      ↩︎ Using two tiers instead of one
      “Compare on cost per completed task, not per token”
      ↩︎ Exam trap 1
      “If you lowered effort, restore it; otherwise try the next tier up at low effort”
      ↩︎ Exam trap 2