What you will be able to do
- Frame a Claude solution's value as cost per completed outcome rather than token consumption
- Map each design decision — model, effort, caching, batching, fast mode — to the pillar it serves
- Pick and defend cost, performance-SLA, productivity and efficiency metrics for a proposed solution
- State honestly which pillars the vendor documentation quantifies and which it only frames
Key concept
Cost per completed outcome — The unit of value for a Claude solution is the finished business task, not the token. You compare designs by what each completed outcome costs and what that outcome is worth to the business, so a pricier model or a higher effort level can be the cheaper design.
1.The frame: business value is measured per outcome, not per token
Every pillar in this objective — efficiency, transformation, productivity, cost, performance SLAs — is a claim about the business, not about the model. So the first move when aligning a solution to value is to fix the unit of measurement. Anthropic's guidance is to measure AI's cost-per-outcome rather than token consumption, and the platform guidance says the same thing in engineering terms: compare candidate designs on cost per completed task, not per token.
That reframing changes which design wins. A cheaper model that burns tokens on retries and needs human correction can produce a more expensive finished task; a frontier model on routine document processing pays for judgement the task never uses. The pillar question is therefore always "is this work hard, or merely large?" — hard work that requires judgment and reasoning justifies a capable model, high volumes of straightforward work do not.
| Model family | The work it is offered for |
|---|---|
| Fable | the hardest problems |
| Opus | long-horizon work and coding |
| Sonnet | everyday work and analysis |
| Haiku | high-volume and routine tasks |
"What does one completed task cost, and what is that completed task worth to us?" Cost per token is not the metric; cost per completed task is, and the value side is business-specific — no vendor can measure for you what the work would have cost without AI.
2.The cost pillar: separate the free wins from the trades
Cost becomes a design constraint at the moment a prototype turns into production. The useful discipline for a solution design is to sort every cost lever into two buckets, because only one of them touches the quality your business case depends on: free wins lower spend without lowering output quality, while tradeoffs exchange cost for intelligence.
Prompt caching, token hygiene, batch processing and a prompt audit sit in the first bucket; model choice, effort, output caps, an elapsed-time clock and multi-model architectures sit in the second. Ordering matters for the business case: in Anthropic's measurements prompt caching was the largest lever by a wide margin, so a design that argues about model tiers before turning on caching is optimising the smaller number. Two caveats keep the first bucket honest — batching trades latency for its discount, and context editing cost more than it saved in the run measured.
| Model | Input | Output |
|---|---|---|
| Claude Fable 5.1 | $10 / MTok | $50 / MTok |
| Claude Opus 5.5 | $4 / MTok | $20 / MTok |
| Claude Sonnet 5 | $2 / MTok | $10 / MTok |
| Claude Haiku 4.5 | $1 / MTok | $5 / MTok |
The documented remedy when costs are too high but quality is fine is also not a model swap: sweep effort down on your current model. Effort is the lever that trades intelligence for latency and cost inside one model, which makes it the cheapest thing to try before renegotiating the architecture.
One caution before any of these numbers enters a business case: the measured factors quoted above are Anthropic-internal results, and the documentation calls them directional, not guarantees, so measure on your own workload with the four-step method. The pillar you commit to in a design review is the one you measured on your own traffic.
A customer-support chatbot must answer high volumes of routine account questions in under a second per response, and the product team's success metric is purely response latency at a sustainable per-ticket cost. Which model choice best aligns the architecture with this performance SLA and cost pillar?
Correct answer: A — Start with Claude Haiku 4.5, the fastest current model with near-frontier intelligence at the lowest per-token price
- A. Correct. Claude Haiku 4.5 is described as the fastest model with near-frontier intelligence at the most economical price point, making it the recommended starting point for latency-sensitive, high-volume, cost-sensitive deployments like a routine-question support bot.
- B. Fast mode on Opus 4.8 does increase output speed, but it does so at premium pricing on an already more expensive model, working against the sustainable per-ticket cost requirement in the scenario.
- C. Adaptive thinking on Sonnet 5 cannot simply be turned off to remove latency; the scenario needs a model chosen for inherent speed and cost efficiency, not a configuration change on a moderately priced model.
- D. Claude Fable 5 targets long-running agentic intelligence with slower comparative latency and higher pricing, which is misaligned with a sub-second response requirement and a low per-ticket cost target.
3.The performance SLA pillar: latency is bought, not wished for
A performance SLA turns latency into a contractual number, so the design has to name which lever buys it and what that lever costs. Three appear in the documentation. Effort is the first: it trades intelligence for latency and cost within a single model, and lower effort also means fewer and terser tool calls, which shortens agentic turns as well as single responses.
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=4096,
messages=[
{
"role": "user",
"content": "Analyze the trade-offs between microservices and monolithic architectures",
}
],
output_config={"effort": "medium"},
)The second is fast mode, and it is the cleanest example of paying money for an SLA with no quality story attached: it runs the same model with a faster inference configuration, with no change to intelligence or capabilities, at premium pricing. Read the SLA it actually improves before promising it — speed benefits are focused on output tokens per second (OTPS), not time to first token (TTFT), so a design whose SLA is measured on first-byte latency buys nothing here.
| Model | Fast mode input | Fast mode output |
|---|---|---|
| Claude Opus 5.5 | $8 USD / MTok | $40 USD / MTok |
| Claude Opus 5 / Claude Opus 4.8 | $10 USD / MTok | $50 USD / MTok |
The third lever is telling the model that time matters and showing it the elapsed time. This one moves cost and SLA together but is the only latency lever with a measured quality cost: runs took 33% to 69% less time at a 28% to 54% lower cost per task, with scores up to 1.9 points lower. That number is what a wall-clock SLA costs you in quality, and it belongs in the design document. In the other direction, batch processing trades latency for its discount, so anything under a tight SLA cannot be moved there.
SLAs also exist above the API. In a support-routing solution, the business-side latency target is an assignment time: best-in-class systems often achieve average assignment times of under 5 minutes, and near-instantaneous routing is possible with LLM implementations. That is the number a value-pillar argument commits to, not tokens per second.
A legal operations team currently has paralegals manually reading and tagging thousands of contracts per month. Leadership frames the initiative not as automating the existing manual process but as fundamentally redesigning the contract intake workflow so that paralegals shift to reviewing model-flagged exceptions and negotiating terms. Which business value pillar does this framing primarily represent?
Correct answer: A — Transformation, because the initiative restructures the underlying workflow and staff roles, not merely speeding up the same process
- A. Correct. Redesigning the intake workflow and shifting staff to exception review and negotiation is a structural change to how work is done, which is the hallmark of the transformation pillar rather than simply accelerating an unchanged process.
- B. Productivity gains describe doing the same tasks faster or with less effort within the existing workflow; this scenario explicitly changes the workflow and role structure, which goes beyond a productivity framing.
- C. Cost reduction focuses on lowering spend for a given output, but the scenario is framed around role and workflow redesign, not a stated cost target, so cost is not the primary pillar being described.
- D. Performance SLA pillars center on measurable service commitments such as latency or turnaround guarantees; the scenario describes a role and workflow change with no SLA metric mentioned.
4.The productivity pillar: output per worker, measured before and after
Productivity is the pillar most often asserted and least often instrumented. The documented way to state it is as a per-person throughput delta with a target attached. In the ticket-routing guide the claim is that improved routing should increase productivity, and the measurement is spelled out: tracking tickets resolved per agent per day or hour, aiming for a 10–20% improvement after implementing a new routing system.
Notice the shape, because it generalises to any solution you design: an existing operational counter, a per-unit denominator, and a percentage improvement the business will accept as success. Related routing targets in the same guide — routing accuracy of 90–95%, a rerouting rate below 10%, an escalation rate below 20% — are quality measures that feed the productivity number rather than replacing it.
Because the pillar needs a counted unit and a target. The documented form is tickets resolved per agent per day or hour, with a 10–20% improvement aimed for after the change; sentiment is measured separately, as CSAT.
Sources6
5.The efficiency pillar: less work per unit of outcome
Efficiency differs from cost: cost is what you pay, efficiency is how much work the system needs in order to produce one outcome. Two documented patterns express it. The first removes outcomes from the expensive path entirely — a deflection rate measures the share resolved by self-service before routing, and the guide aims for a deflection rate of 20–30%, with top performers achieving rates of 40% or higher. The downstream effect is stated as a cost figure: many organizations aim to reduce cost per ticket by 10–15% after implementing an improved routing system.
The second pattern makes the model path itself efficient by spending capability only where it is needed. Where the outputs can be checked with tests or a verifier, the documented rule is to run everything at low effort and re-run failures at high; on the coding benchmark measured, the pass rate held at about half the cost. The same logic drives the two multi-model shapes: an executor that escalates hard decisions to an advisor, and an orchestrator that delegates bulk work to lower-cost workers, so that most tokens are billed at the lower rate.
An efficiency claim needs its precondition named. The advisor shape pays off only when the advisor is priced well above the executor and actually consulted, which is why the documentation tells you to price the advisor's model alone at low effort and measure the consult rate first. The orchestrator shape is for work that exceeds one context window. Neither is a general-purpose saving.
6.The transformation pillar: work that would not have happened otherwise
The sources do not offer a transformation framework as such; they express transformation through one question and one arithmetic. The question is what the work would have cost without AI, whether in resources, time, or never attempting the project at all. That last clause is the transformation pillar in a single phrase — value created where there was previously no project, not a percentage saved on an existing one. The documentation is explicit that this side is yours to size: the answer is specific to your business and needs, and no vendor can measure it for you.
The arithmetic makes that sizing concrete at the level of a repeatable workflow. Each skill represents a repeatable workflow, so its cost can be weighed directly against what that workflow is worth. You assign a value per run to each skill — an estimate of the employee time it replaces — and compute the net value each skill has generated from the cost per use and total uses reported in analytics. Even rough estimates carry the argument: a call-prep skill costing $0.90 per run against $20 of value returns 20x on every use.
That is the whole of what these sources support on transformation, and it is worth saying plainly rather than padding: the documented treatment is a cost-per-outcome comparison against the no-AI baseline, plus per-workflow ROI from usage analytics. Anything broader — process redesign, new product lines — is not asserted here.
A solutions architect is scoping a new Claude-based document summarization feature for an enterprise sponsor. The sponsor asks for a business case that ties the proposed architecture to measurable outcomes before funding is approved. Which set of pillars should the architect use to structure that business case?
Correct answer: A — Efficiency, transformation, productivity, cost, and performance SLAs, each mapped to a specific metric the sponsor already tracks
- A. Correct. Anthropic's guidance frames solution value in terms of the business pillars of efficiency, transformation, productivity, cost, and performance SLAs; mapping each to a metric the sponsor already tracks is what makes a business case measurable and fundable.
- B. Context window size, throughput, and rate limits are technical model specifications, not business outcomes, and a funding sponsor needs the case tied to value pillars rather than raw model capability numbers.
- C. Counting prompt templates or integrated tools measures implementation activity, not business value delivered, so it does not give the sponsor a basis for approving funding against outcomes.
- D. Post-launch developer sentiment is a useful secondary signal but is collected after shipping and does not itself constitute the pre-approval business case the sponsor is requesting.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A cheaper model always improves the cost pillar, so pushing complex reasoning down to Haiku is a saving.Why is that wrong?
Underpowered models burn tokens on retries and need human correction, which raises the cost of the finished task even though the per-token price fell.
Covered in The frame: business value is measured per outcome, not per token
2.Model selection is the biggest cost decision in a Claude architecture, so the design review should start there.Why is that wrong?
In Anthropic's measurements caching was the largest lever and cache reads are routinely the largest single component of task cost, outweighing most model-choice decisions.
Covered in The cost pillar: separate the free wins from the trades
3.Fast mode is a higher-capability configuration, so buying it improves quality as well as latency.Why is that wrong?
Fast mode is the same model weights and behaviour with a faster inference configuration; only output speed changes, and you pay a premium for it.
Covered in The performance SLA pillar: latency is bought, not wished for
4.Batch processing is a free win, so any workload can be moved there to cut the bill in half.Why is that wrong?
The discount is paid for in latency — batch work can wait up to 24 hours — so a workload under a latency SLA cannot use it.
Covered in The performance SLA pillar: latency is bought, not wished for
5.Adding a frontier advisor above a cheaper executor is a general efficiency win.Why is that wrong?
It only pays when the advisor's model is priced well above the executor's and is actually consulted, which you establish by pricing it alone and measuring the consult rate.
Covered in The efficiency pillar: less work per unit of outcome
6.A published Anthropic benchmark result is enough evidence for the productivity and cost pillars in your own business case.Why is that wrong?
The measurements are Anthropic-internal and directional rather than guarantees; the documentation tells you to measure the same levers on your own workload.
Covered in The cost pillar: separate the free wins from the trades
Practise it for real
Turn the cost and SLA pillars into two numbers for one real request, by running it at different effort levels and comparing cost per completed task rather than cost per token.
1.Send your representative request with output_config={"effort": "medium"} on claude-opus-5-5, exactly as in the sample above, and record output tokens, wall-clock time and whether the answer was acceptable.
Why: Medium is the default on Claude Opus 5.5, so this is your baseline rather than an intervention.
You should see: One usable answer plus a baseline row: tokens, seconds, pass or fail.
2.Re-send the identical request with effort set to low.
Why: Low is the most efficient level, giving significant token savings with some capability reduction — this measures the size of both on your workload.
You should see: Fewer output tokens and a faster response; the answer may or may not still pass.
3.For any low-effort attempt that failed your check, re-run only that attempt at high effort.
Why: This is the documented run-cheap-then-retry pattern: run everything at low effort and re-run failures at high.
You should see: The failures recover, and only the failing subset pays the higher token bill.
4.Compute total spend across the whole set divided by the number of acceptable answers, at each configuration, using the per-MTok prices in the pricing table above.
Why: This is the pillar metric: compare on cost per completed task, not per token.
You should see: A single figure per configuration, and often a low-plus-retry figure below the all-medium one.
5.If a latency SLA applies, repeat the winning configuration with speed="fast" and the fast-mode beta header, and price it at the fast-mode rates.
Why: Fast mode buys output tokens per second at a premium with no capability change, so the only question it raises is whether the SLA is worth the multiplier.
You should see: Higher output tokens per second, an unchanged pass rate, and a cost per completed task you can set against the SLA's business value.
Stuck? Get a nudge
If a stop_reason of max_tokens appears, that is a budget problem, not an effort problem — raise max_tokens before drawing any conclusion about the effort level.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“measure AI’s cost-per-outcome instead of token consumption as the primary metric of value”
↩︎ The frame: business value is measured per outcome, not per token“hard and requires judgment and reasoning, or is it just large”
↩︎ The frame: business value is measured per outcome, not per token“What would this work have cost without AI, whether in resources, time, or never attempting the project at all?”
↩︎ The transformation pillar: work that would not have happened otherwise“Assigning a less expensive model complex reasoning often makes the finished task more expensive”
↩︎ Exam trap 1 - 2.https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligenceOfficial docs
“Compare on cost per completed task, not per token”
↩︎ The frame: business value is measured per outcome, not per token“When a workload moves from prototype to production, cost becomes a first-class design constraint.”
↩︎ The cost pillar: separate the free wins from the trades“Free wins cut spend without touching quality”
↩︎ The cost pillar: separate the free wins from the trades“Tradeoffs exchange cost for intelligence: model choice, effort, output caps and task budgets”
↩︎ The cost pillar: separate the free wins from the trades“prompt caching was the largest lever by a wide margin”
↩︎ The cost pillar: separate the free wins from the trades“Sweep effort down on your current model”
↩︎ The cost pillar: separate the free wins from the trades“runs took 33% to 69% less time at a 28% to 54% lower cost per task, with scores up to 1.9 points lower”
↩︎ The performance SLA pillar: latency is bought, not wished for“batch processing trades latency for its discount”
↩︎ The performance SLA pillar: latency is bought, not wished for“Run everything at low effort and re-run failures at high; on the coding benchmark measured, the pass rate held at about half the cost”
↩︎ The efficiency pillar: less work per unit of outcome“The work exceeds one context window”
↩︎ The efficiency pillar: less work per unit of outcome“Compare on cost per completed task, not per token”
↩︎ Key concept“making caching worth more than most model-choice decisions”
↩︎ Exam trap 2“batch processing trades latency for its discount”
↩︎ Exam trap 4“It pays off when priced well above the executor and actually consulted”
↩︎ Exam trap 5“directional, not guarantees, so measure on your own workload with the four-step method”
↩︎ Exam trap 6 - 3.
“an effort parameter that trades intelligence for latency and cost within a single model”
↩︎ The cost pillar: separate the free wins from the trades“an executor that escalates hard decisions to an advisor, and an orchestrator that delegates bulk work to lower-cost workers”
↩︎ The efficiency pillar: less work per unit of outcome - 4.
“Lower effort also means fewer and terser tool calls.”
↩︎ The performance SLA pillar: latency is bought, not wished for - 5.
“Fast mode runs the same model with a faster inference configuration. There is no change to intelligence or capabilities.”
↩︎ The performance SLA pillar: latency is bought, not wished for“Speed benefits are focused on output tokens per second (OTPS), not time to first token (TTFT)”
↩︎ The performance SLA pillar: latency is bought, not wished for“Fast mode is priced at a multiplier on standard rates across the full context window”
↩︎ The performance SLA pillar: latency is bought, not wished for“Same model weights and behavior (not a different model)”
↩︎ Exam trap 3 - 6.
“Best-in-class systems often achieve average assignment times of under 5 minutes”
↩︎ The performance SLA pillar: latency is bought, not wished for“Improved routing should increase productivity.”
↩︎ The productivity pillar: output per worker, measured before and after“tracking tickets resolved per agent per day or hour, aiming for a 10–20% improvement after implementing a new routing system”
↩︎ The productivity pillar: output per worker, measured before and after“Industry benchmarks often aim for 90–95% accuracy”
↩︎ The productivity pillar: output per worker, measured before and after“Aim for a deflection rate of 20–30%, with top performers achieving rates of 40% or higher.”
↩︎ The efficiency pillar: less work per unit of outcome“many organizations aim to reduce cost per ticket by 10–15% after implementing an improved routing system”
↩︎ The efficiency pillar: less work per unit of outcome - 7.
“Each skill represents a repeatable workflow”
↩︎ The transformation pillar: work that would not have happened otherwise“a call-prep skill costing $0.90 per run against $20 of value returns 20x on every use”
↩︎ The transformation pillar: work that would not have happened otherwise“Assign a value per run to each skill”
↩︎ The transformation pillar: work that would not have happened otherwise