What you will be able to do
- Explain why input, output, cache-write and cache-read tokens have to be priced separately in a cost model
- Use the token counting endpoint to measure a request's input tokens before sending it, and name the inputs it cannot count
- Re-baseline token counts when moving to a model with a different tokenizer
- Build a per-request cost estimate from real usage data and pick the cost levers that do not reduce quality
Key concept
Per-token-class pricing — Claude bills per million tokens, and each class of token has its own rate: input, output, cache writes (5-minute or 1-hour) and cache hits. Every budgeting, tracking or caching decision comes down to moving tokens from an expensive class into a cheaper one.
1.Every token class has its own price
Claude is billed per million tokens (MTok), and tokens are not all priced the same. For each model the pricing page lists five rates: input tokens, output tokens, 5-minute cache writes, 1-hour cache writes, and cache hits and refreshes. So a cost model that multiplies total tokens by one blended rate is wrong before it starts. You need to know how many tokens of each class a request uses.
| Model | Input | Output | Output relative to input |
|---|---|---|---|
| Claude Fable 5.1 | $10 / MTok | $50 / MTok | 5x |
| Claude Opus 5.5 | $4 / MTok | $20 / MTok | 5x |
| Claude Sonnet 5 | $2 / MTok | $10 / MTok | 5x |
| Claude Haiku 4.5 | $1 / MTok | $5 / MTok | 5x |
On all four current models an output token costs five times as much as an input token. That tells you where to look first. A workload that writes long answers is driven by output cost. An agent that re-reads a large context on every turn is driven by input cost. Also, a newer model is not automatically a more expensive one. Among the additional models on the page, Claude Opus 4.1 is listed at $15 input and $75 output, while Claude Opus 5.5 is $4 and $20. The model ID in your configuration is a cost decision in its own right.
Sources1
2.Count tokens before you send
You can't budget what you haven't measured. The token counting endpoint tells you how many input tokens a request contains without running the model. It takes the same structured input as creating a message, including system prompts, tools, images and PDFs, and returns the total input tokens. The documentation lists three uses: managing rate limits and costs, making model routing decisions, and fitting prompts to a target length.
client = anthropic.Anthropic()
response = client.messages.count_tokens(
model="claude-opus-5-5",
system="You are a scientist",
messages=[{"role": "user", "content": "Hello, Claude"}],
)
print(response.json())This request returns 14 input tokens. The documentation's other examples show how quickly the fixed parts of a request add up. A short weather question sent with one tool definition counts 403 tokens. A base64 image plus "Describe this image" counts 1,028. A base64 PDF with a summarise instruction counts 2,188. Tool definitions and attachments are billed on every request that carries them, so count them explicitly. Don't guess from the length of the user's message.
The endpoint has gaps. It returns an invalid_request_error for server tools (web search, web fetch, code execution, tool search), for the MCP connector, and for image or document blocks that use a url or file source. Send images and PDFs as base64 if you want them counted. For requests that use server tools or MCP, the only measurement is the usage object on the real Messages API response. Counting is free, but it is rate-limited by usage tier: 5,000 requests per minute on Start, 10,000 on Build and 20,000 on Scale. It is also an estimate that ignores caching. You can include cache_control blocks in a count request, but nothing is cached, and the count does not show what caching would save.
Not necessarily. Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5 and Claude Mythos 5 use the tokenizer introduced with Claude Opus 4.7. The same prompt counts roughly 30 percent higher on that tokenizer than on models before Opus 4.7, and the exact difference depends on the content. The endpoint counts with the tokenizer of whichever model you pass. So before a migration, count the same request once with the current model and once with the target model, and compare the two input_tokens values. Historical dashboards built on the old tokenizer will under-project spend on the new one.
A document-analysis pipeline caches a 40,000-token reference manual using an ephemeral cache_control breakpoint with the default TTL. Analysts submit follow-up questions against the manual in bursts, but gaps between bursts are frequently 8 to 10 minutes. The team observes that cache reads are not occurring for most follow-up bursts, so they keep paying the full cache-write price again. What change best fits this usage pattern?
Correct answer: A — Set the cache_control ttl to "1h" so the manual stays cached across the longer gaps between analyst bursts, avoiding repeated 5-minute cache writes
- A. Correct. The default ephemeral cache TTL is 5 minutes, which is shorter than the observed 8-10 minute gaps, so the cache expires between bursts. The 1-hour TTL option (2x base input price for the write) covers that gap and lets later bursts hit the cache instead of re-writing it.
- B. A fixed 5-minute pre-warm cadence only refreshes the cache if requests actually land within the 5-minute TTL window; it does not address gaps of 8-10 minutes without essentially running a background job indefinitely, and it isn't the direct mechanism Anthropic provides for extending cache lifetime.
- C. Splitting into four breakpoints changes granularity, not TTL; each 10,000-token chunk would still expire after 5 minutes of inactivity under the default TTL, so the same problem recurs.
- D. There is no context-window setting that retains prompt content in server-side memory between sessions; the context window defines the maximum tokens per request, and caching is controlled through cache_control, not context window size.
Sources2
3.Track real usage and build the cost model
Counting gives you estimates before you send. Tracking records what you were actually billed after the call. Every Messages API response includes a usage object, and the documentation's caching examples print it with response.usage.model_dump_json(). This is the authoritative record. It covers the requests the counting endpoint rejects, and it splits cache writes out by duration, for example as cache_creation.ephemeral_5m_input_tokens. Log it for every request. Your cost model should be built from these logged numbers, not from assumptions.
The model itself is plain arithmetic. For each request, multiply each token class by its rate and add the results. Then multiply by request volume. Cached input has to be split out, because cache writes and cache reads have different rates from normal input. That is covered where caching is taught.
Input: 10,000 × $2 / 1,000,000 = $0.02. Output: 1,000 × $10 / 1,000,000 = $0.01. Total: $0.03 per request. Input is two-thirds of the bill, even though output tokens are five times dearer per token, so the fixed prompt is where to cut first.
Anthropic's cost guide divides the levers into two groups. The first group are free wins, which cut spend without touching quality: prompt caching, token hygiene, a prompt audit against the model you actually run, batch processing, and workspace spend limits as a backstop. Batch processing costs 50% less, in return for waiting up to 24 hours for results. The second group are tradeoffs that buy lower cost with less intelligence: model choice, effort, output caps and task budgets. When you compare models, compare on cost per completed task. A cheaper per-token model that needs more attempts can cost more overall.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The token counting endpoint shows how much a request will cost after prompt caching, because it accepts cache_control blocks.Why is that wrong?
Counting accepts cache_control but ignores caching. It returns a raw input-token estimate, and caching only happens when a message is actually created.
Covered in Count tokens before you send
2.Token counts measured on the old model still work for projecting cost after moving to a newer model.Why is that wrong?
Tokenizers differ between models, so the same prompt can count noticeably higher. Recount with the target model, because the endpoint uses the tokenizer of the model you pass.
Covered in Count tokens before you send
3.The model with the lowest per-token price is always the cheapest choice.Why is that wrong?
What you pay for is completed work. A model that needs more attempts or more tokens per task can cost more overall, so compare models per completed task.
Covered in Track real usage and build the cost model
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“This page provides detailed pricing information for Anthropic's models and features. All prices are in USD.”
↩︎ Every token class has its own price“This section covers partner-operated cloud platforms, where the cloud provider invoices you.”
↩︎ Every token class has its own price - 2.
“Token counting lets you determine the number of tokens in a message before you send it to Claude.”
↩︎ Count tokens before you send“Token counting is free to use but subject to requests per minute rate limits based on your usage tier.”
↩︎ Count tokens before you send“prompt caching only occurs during actual message creation.”
↩︎ Count tokens before you send“roughly 30 percent higher than on models before Claude Opus 4.7 (the exact increase depends on the content)”
↩︎ Count tokens before you send“For requests that use server tools or MCP servers, the Messages API response reports the tokens used in its usage object.”
↩︎ Track real usage and build the cost model“token counting provides an estimate without using caching logic.”
↩︎ Exam trap 1“The token counting endpoint counts under the tokenizer of the model you pass.”
↩︎ Exam trap 2 - 3.https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligenceOfficial docs
“batch processing at 50% off for work that can wait up to 24 hours”
↩︎ Track real usage and build the cost model“Compare on cost per completed task, not per token”
↩︎ Exam trap 3
Also cited
“Prompt caching introduces a new pricing structure. The following table shows the price per million tokens for each supported model”
↩︎ Key concept