CertSafari
    CCAR-P · Lessons

    Domain 2 · Lesson 10/38

    Claude Context Windows and Token Budgeting

    Optimize context windows and manage token usage

    8 min read
    2.6% of exam
    2 sources
    Published 27 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Explain why a larger context window does not by itself give better answers
    • List every part of a request that consumes context window tokens, including thinking and cached tokens
    • State the context window and output limits of current Claude models
    • Estimate a request's size before sending it with the token counting API

    Key concept

    Context rot — The context window is the model's working memory for a request. As more tokens fill it, accuracy and recall drop, so choosing what goes into the window matters as much as how big the window is.

    1.The context window is working memory, not a bucket to fill

    The context window is everything Claude can refer to while it generates a response, and that includes the response itself. It is separate from the data the model was trained on. Think of it as working memory for one request.

    In a normal API conversation nothing drops out on its own. Each user message and each assistant reply stays in the window, so every new turn sends the full history plus the new message. The output of one turn becomes input for the next. Over a long conversation or agent run, the window only gets fuller.

    It is tempting to treat a large window as permission to put everything in it. The documentation says otherwise: more context is not automatically better. Accuracy and recall degrade as token count grows, so a well-chosen 50k-token prompt can beat a careless 500k-token one. Keeping the window efficient is partly about cost, but it is also about answer quality.

    Sources1

    2.What counts toward the window

    To plan a token budget you need to know what is in it. Every part of the request counts: the system prompt, every entry in messages (tool results, images and documents included), and your tool definitions. The output Claude generates on the turn counts too, including extended thinking.

    Two cases are easy to get wrong. First, tool definitions use context even when Claude never calls the tool. A long tool catalogue therefore costs space on every turn. The docs point to the tool search tool as a way to defer tool definitions. Second, prompt caching does not shrink the window. With caching on, the usage field splits input into input_tokens, cache_read_input_tokens and cache_creation_input_tokens, and all three count toward the window.

    Where the tokens in one request come from
    ComponentCounts toward the window?
    System promptYes
    Messages, including tool results, images and documentsYes
    Tool definitionsYes
    Output for this turn, including extended thinkingYes
    Cache reads and cache writesYes. All three input counts count

    Sources1

    3.Window sizes, output caps and thinking tokens

    Current flagship models, including Claude Fable 5.1, Claude Opus 5.5 and Claude Sonnet 5, have a 1M-token context window. On these models 1M is the default: you need no beta header, and long-context requests are billed at standard pricing. Other models, including Claude Sonnet 4.5, have 200k. A single request to a 1M model can generate up to 128k output tokens through max_tokens. Large media can hit a different limit first. A request can include up to 600 images or PDF pages (100 on 200k models), so a document-heavy request may reach the request size limit before the token limit.

    Context window by model family (from the context windows doc)
    ModelsContext windowNotes
    Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5, Claude Sonnet 4.6 and others listed1M tokensDefault, no beta header; up to 128k output tokens per request
    Other models, including Claude Sonnet 4.5200k tokensUp to 100 images or PDF pages per request

    Thinking needs its own entry in the budget. Thinking tokens come out of max_tokens, are billed as output, and count toward rate limits. With adaptive thinking, Claude decides how much to think, so thinking usage varies from request to request.

    What happens to earlier thinking depends on the model. On Claude Opus 4.5 and later Opus models, Claude Sonnet 4.6 and later Sonnet models, and the Fable and Mythos models, the API keeps previous thinking blocks, and they take up space like any other input. On earlier Opus and Sonnet models and on all Haiku models, the API strips previous thinking blocks for you. There is one case where you must send thinking back: during a tool-use cycle, return the complete, unmodified thinking block, including its signature, together with the tool result.

    Sources1

    4.How the model tracks its own budget

    Some models track how much space is left without any help from you. Claude Sonnet 5, Claude Sonnet 4.6, Claude Sonnet 4.5 and Claude Haiku 4.5 have context awareness. The API adds the total budget to the system prompt and, after each tool call, reports how many tokens are used and how many remain. That lets the model pace a long task against the space that remains instead of guessing. The feature is automatic; you never send these tags yourself.

    Claude Opus 4.7 and later Opus models, and the Fable and Mythos models, do not get these injected tags. On those models you can set an explicit budget with task budgets, which are in beta.

    The update the API injects after each tool call on context-aware modelsxml
    <system_warning>Token usage: 35000/200000; 165000 remaining</system_warning>

    Sources1

    5.Measure before you send: the token counting API

    The usage field tells you what a request cost after it ran. To budget beforehand, use the token counting endpoint. It takes the same inputs as a Messages request (system prompt, tools, images, PDFs) and returns the total input tokens. The docs list three uses: managing rate limits and costs up front, routing a request to a model based on its size, and trimming a prompt to a target length.

    Counting a request's input tokens before sending itpython
    response = client.messages.count_tokens(
        model="claude-opus-5-5",
        system="You are a scientist",
        messages=[{"role": "user", "content": "Hello, Claude"}],
    )
    
    print(response.json())

    Keep a few details in mind. The endpoint counts with the tokenizer of the model you pass. Claude Opus 4.7 introduced a new tokenizer, which the Fable and Mythos models share, and it counts the same prompt roughly 30 percent higher than earlier models did. So before you migrate, count the same request under both models and compare. The endpoint rejects server tools, the MCP connector, and images or documents passed by url or file source; send media as base64. Counting is free but rate-limited by usage tier. It is also only an estimate: it does not apply caching logic.

    A team is building a customer support agent that sends a 6,000-token static system prompt containing product policies and tone guidelines, followed by a short per-request user message that always differs. Both latency and cost matter, and the team wants to maximize the prompt cache hit rate across thousands of daily requests. Where should they place the cache_control breakpoint?

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A 1M-token window means you should include every document that might be relevant.Why is that wrong?

      Accuracy and recall degrade as the token count grows, so what you put in the window matters as much as how much room there is.

      Covered in The context window is working memory, not a bucket to fill

    2. 2.Cached tokens are free space: caching a large prefix leaves more of the window for the conversation.Why is that wrong?

      Caching changes what you pay for those tokens. Cache reads and cache writes still count toward the context window.

      Covered in What counts toward the window

    3. 3.A token count from one Claude model carries over unchanged to any other Claude model.Why is that wrong?

      The endpoint counts with the tokenizer of the model you pass, and the newer tokenizer produces counts roughly 30 percent higher than earlier models.

      Covered in Measure before you send: the token counting API

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “This makes curating what's in context just as important as how much space is available.”
      ↩︎ The context window is working memory, not a bucket to fill
      “Everything in the request counts toward the context window: the system prompt, every message in messages (including tool results, images, and documents)”
      ↩︎ What counts toward the window
      “The output Claude generates for the turn, including its extended thinking, counts too.”
      ↩︎ What counts toward the window
      “If you use prompt caching, the input count is split across input_tokens, cache_read_input_tokens, and cache_creation_input_tokens, and all three count toward the window.”
      ↩︎ What counts toward the window
      “A single request to any of them can generate up to 128k output tokens (max_tokens).”
      ↩︎ Window sizes, output caps and thinking tokens
      “Other Claude models, including Claude Sonnet 4.5, have a 200k-token context window.”
      ↩︎ Window sizes, output caps and thinking tokens
      “Thinking tokens are a subset of your max_tokens parameter, are billed as output tokens, and count toward rate limits.”
      ↩︎ Window sizes, output caps and thinking tokens
      “You must return the thinking block with the corresponding tool results. This is the only case where you have to return thinking blocks.”
      ↩︎ Window sizes, output caps and thinking tokens
      “these models track their remaining context window (their "token budget") throughout a conversation.”
      ↩︎ How the model tracks its own budget
      “you can give the model an explicit budget with task budgets, which are in beta.”
      ↩︎ How the model tracks its own budget
      “As token count grows, accuracy and recall degrade, a phenomenon known as context rot.”
      ↩︎ Key concept
      “A larger context window allows the model to handle more complex and lengthy prompts, but more context isn't automatically better.”
      ↩︎ Exam trap 1
      “Cached prompt prefixes still occupy the context window: prompt caching changes what you pay for those tokens, not whether they count.”
      ↩︎ Exam trap 2
    2. 2.
      “Token counting lets you determine the number of tokens in a message before you send it to Claude.”
      ↩︎ Measure before you send: the token counting API
      “roughly 30 percent higher than on models before Claude Opus 4.7”
      ↩︎ Measure before you send: the token counting API
      “token counting provides an estimate without using caching logic”
      ↩︎ Measure before you send: the token counting API
      “The token counting endpoint counts under the tokenizer of the model you pass.”
      ↩︎ Exam trap 3

    Continue to page 2 of 2

    Prompt Caching, Context Editing and Compaction in Claude