CertSafari
    CLAUDE-CERTIFIED-DEVELOPER-FOUNDATIONS-CCDV-F · Lessons

    Domain 5 · Lesson 15/25

    Claude Token Pricing, Counting and Cost Modeling

    Cost and Token Management

    8 min read
    4.2% of exam
    4 sources
    Published 29 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Explain why input, output, cache-write and cache-read tokens have to be priced separately in a cost model
    • Use the token counting endpoint to measure a request's input tokens before sending it, and name the inputs it cannot count
    • Re-baseline token counts when moving to a model with a different tokenizer
    • Build a per-request cost estimate from real usage data and pick the cost levers that do not reduce quality

    Key concept

    Per-token-class pricing — Claude bills per million tokens, and each class of token has its own rate: input, output, cache writes (5-minute or 1-hour) and cache hits. Every budgeting, tracking or caching decision comes down to moving tokens from an expensive class into a cheaper one.

    1.Every token class has its own price

    Claude is billed per million tokens (MTok), and tokens are not all priced the same. For each model the pricing page lists five rates: input tokens, output tokens, 5-minute cache writes, 1-hour cache writes, and cache hits and refreshes. So a cost model that multiplies total tokens by one blended rate is wrong before it starts. You need to know how many tokens of each class a request uses.

    Base input and output prices for the current Claude models (USD per million tokens)
    ModelInputOutputOutput relative to input
    Claude Fable 5.1$10 / MTok$50 / MTok5x
    Claude Opus 5.5$4 / MTok$20 / MTok5x
    Claude Sonnet 5$2 / MTok$10 / MTok5x
    Claude Haiku 4.5$1 / MTok$5 / MTok5x

    On all four current models an output token costs five times as much as an input token. That tells you where to look first. A workload that writes long answers is driven by output cost. An agent that re-reads a large context on every turn is driven by input cost. Also, a newer model is not automatically a more expensive one. Among the additional models on the page, Claude Opus 4.1 is listed at $15 input and $75 output, while Claude Opus 5.5 is $4 and $20. The model ID in your configuration is a cost decision in its own right.

    Sources1

    2.Count tokens before you send

    You can't budget what you haven't measured. The token counting endpoint tells you how many input tokens a request contains without running the model. It takes the same structured input as creating a message, including system prompts, tools, images and PDFs, and returns the total input tokens. The documentation lists three uses: managing rate limits and costs, making model routing decisions, and fitting prompts to a target length.

    Counting the input tokens of a request with messages.count_tokenspython
    client = anthropic.Anthropic()
    
    response = client.messages.count_tokens(
        model="claude-opus-5-5",
        system="You are a scientist",
        messages=[{"role": "user", "content": "Hello, Claude"}],
    )
    
    print(response.json())

    This request returns 14 input tokens. The documentation's other examples show how quickly the fixed parts of a request add up. A short weather question sent with one tool definition counts 403 tokens. A base64 image plus "Describe this image" counts 1,028. A base64 PDF with a summarise instruction counts 2,188. Tool definitions and attachments are billed on every request that carries them, so count them explicitly. Don't guess from the length of the user's message.

    The endpoint has gaps. It returns an invalid_request_error for server tools (web search, web fetch, code execution, tool search), for the MCP connector, and for image or document blocks that use a url or file source. Send images and PDFs as base64 if you want them counted. For requests that use server tools or MCP, the only measurement is the usage object on the real Messages API response. Counting is free, but it is rate-limited by usage tier: 5,000 requests per minute on Start, 10,000 on Build and 20,000 on Scale. It is also an estimate that ignores caching. You can include cache_control blocks in a count request, but nothing is cached, and the count does not show what caching would save.

    Not necessarily. Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5 and Claude Mythos 5 use the tokenizer introduced with Claude Opus 4.7. The same prompt counts roughly 30 percent higher on that tokenizer than on models before Opus 4.7, and the exact difference depends on the content. The endpoint counts with the tokenizer of whichever model you pass. So before a migration, count the same request once with the current model and once with the target model, and compare the two input_tokens values. Historical dashboards built on the old tokenizer will under-project spend on the new one.

    A document-analysis pipeline caches a 40,000-token reference manual using an ephemeral cache_control breakpoint with the default TTL. Analysts submit follow-up questions against the manual in bursts, but gaps between bursts are frequently 8 to 10 minutes. The team observes that cache reads are not occurring for most follow-up bursts, so they keep paying the full cache-write price again. What change best fits this usage pattern?

    Sources2

    3.Track real usage and build the cost model

    Counting gives you estimates before you send. Tracking records what you were actually billed after the call. Every Messages API response includes a usage object, and the documentation's caching examples print it with response.usage.model_dump_json(). This is the authoritative record. It covers the requests the counting endpoint rejects, and it splits cache writes out by duration, for example as cache_creation.ephemeral_5m_input_tokens. Log it for every request. Your cost model should be built from these logged numbers, not from assumptions.

    The model itself is plain arithmetic. For each request, multiply each token class by its rate and add the results. Then multiply by request volume. Cached input has to be split out, because cache writes and cache reads have different rates from normal input. That is covered where caching is taught.

    Anthropic's cost guide divides the levers into two groups. The first group are free wins, which cut spend without touching quality: prompt caching, token hygiene, a prompt audit against the model you actually run, batch processing, and workspace spend limits as a backstop. Batch processing costs 50% less, in return for waiting up to 24 hours for results. The second group are tradeoffs that buy lower cost with less intelligence: model choice, effort, output caps and task budgets. When you compare models, compare on cost per completed task. A cheaper per-token model that needs more attempts can cost more overall.

    Sources23

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.The token counting endpoint shows how much a request will cost after prompt caching, because it accepts cache_control blocks.Why is that wrong?

      Counting accepts cache_control but ignores caching. It returns a raw input-token estimate, and caching only happens when a message is actually created.

      Covered in Count tokens before you send

    2. 2.Token counts measured on the old model still work for projecting cost after moving to a newer model.Why is that wrong?

      Tokenizers differ between models, so the same prompt can count noticeably higher. Recount with the target model, because the endpoint uses the tokenizer of the model you pass.

      Covered in Count tokens before you send

    3. 3.The model with the lowest per-token price is always the cheapest choice.Why is that wrong?

      What you pay for is completed work. A model that needs more attempts or more tokens per task can cost more overall, so compare models per completed task.

      Covered in Track real usage and build the cost model

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “This page provides detailed pricing information for Anthropic's models and features. All prices are in USD.”
      ↩︎ Every token class has its own price
      “This section covers partner-operated cloud platforms, where the cloud provider invoices you.”
      ↩︎ Every token class has its own price
    2. 2.
      “Token counting lets you determine the number of tokens in a message before you send it to Claude.”
      ↩︎ Count tokens before you send
      “Token counting is free to use but subject to requests per minute rate limits based on your usage tier.”
      ↩︎ Count tokens before you send
      “prompt caching only occurs during actual message creation.”
      ↩︎ Count tokens before you send
      “roughly 30 percent higher than on models before Claude Opus 4.7 (the exact increase depends on the content)”
      ↩︎ Count tokens before you send
      “For requests that use server tools or MCP servers, the Messages API response reports the tokens used in its usage object.”
      ↩︎ Track real usage and build the cost model
      “token counting provides an estimate without using caching logic.”
      ↩︎ Exam trap 1
      “The token counting endpoint counts under the tokenizer of the model you pass.”
      ↩︎ Exam trap 2
    3. 3.
      “batch processing at 50% off for work that can wait up to 24 hours”
      ↩︎ Track real usage and build the cost model
      “Compare on cost per completed task, not per token”
      ↩︎ Exam trap 3

    Also cited

    Continue to page 2 of 2

    Prompt Caching: Breakpoints, TTLs and Cache Pricing