CertSafari
    CLAUDE-CERTIFIED-DEVELOPER-FOUNDATIONS-CCDV-F · Lessons

    Domain 6 · Lesson 16/25

    Claude Context Window Management: What Counts and Why Context Rot Matters

    Context Engineering

    11 min read
    3.67% of exam
    6 sources
    Published 29 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Explain context engineering and why a larger context window does not automatically give better answers
    • List every part of a request that uses up the context window, including thinking blocks and cached prefixes
    • Describe how context awareness tells a model how much of its window remains, and which models receive it
    • Apply context drift prevention in a long-running agent by clearing stale tool results or compacting older turns
    • Keep a multi-step agentic workflow's intermediate work out of the main context with programmatic tool calling and the memory tool

    Key concept

    Context as a finite budget — The context window is the model's working memory. Recall and accuracy get worse as it fills, so the job is to choose what goes into it, not only to fit things in.

    1.From prompt engineering to context engineering

    Prompt engineering is about how you write and organise instructions. Context engineering asks a wider question: what is the whole set of tokens the model sees on a given turn? That set includes system instructions, tools, MCP, external data and message history, and Anthropic describes context engineering as the natural next step after prompt engineering. The difference matters most for agents. An agent running in a loop keeps producing material that might be relevant to its next turn, so choosing what to include is not a one-off task. It happens every time you decide what to pass to the model.

    The context window is everything the model can refer to when it writes a response, including the response itself. It is separate from the data the model was trained on and works as the model's working memory. Windows go up to 1M tokens, depending on the model, but more context is not automatically better. As the token count grows, accuracy and recall get worse. The documentation calls this context rot.

    Anthropic's engineering blog gives the reason. In a transformer, every token can attend to every other token, which gives n² pairwise relationships for n tokens. As context gets longer, the model's ability to track those relationships is spread thinner. Models also learn their attention patterns mostly from shorter sequences. The result is a gradual decline, not a sudden drop: models stay highly capable at long lengths but may be less precise at retrieving information and reasoning across long distances. That is why you should spend context deliberately, as a budget, instead of treating it as a bucket to fill.

    Prevention of context drift and bloat: in a long agent loop, the window slowly fills with material that mattered once and is now stale, such as old file contents, finished search results and past reasoning. The context editing documentation puts it plainly: context is a finite resource with diminishing returns, and irrelevant content degrades model focus. The model drifts away from the current task because the stale material is still competing for its attention. Preventing that drift means taking stale content out as the conversation grows, not waiting for the hard limit.

    The documentation offers two main ways to prune. Tool result clearing (clear_tool_uses_20250919, a context editing strategy) clears the oldest tool results once input passes a threshold you set. The API replaces each cleared result with placeholder text so Claude knows something was removed. You can keep the most recent tool uses, exclude specific tools, and use clear_at_least so that each clearing is worth the prompt-cache invalidation it causes. Thinking block clearing (clear_thinking_20251015) does the same job for old thinking blocks. Compaction takes a different approach: it replaces older turns with a summary Claude writes on the server, which keeps the active context small because response quality degrades as a conversation grows. Server-side compaction is in beta for Claude 4.6 and later models.

    Context editing happens on the server, before the prompt reaches Claude. Your application keeps the full, unmodified conversation history and does not need to sync with the edited version. That is why pruning is safe to apply: the model sees less, but you still have the complete record.

    Sources1234

    2.What uses up the context window

    To manage a window, you first need to know what uses it up. The short answer is every part of the request, plus what Claude generates in reply. Each response reports what the request consumed in its usage field. To estimate a request's size before you send it, use the token counting API.

    What counts toward the context window
    ComponentCounts toward the window?Nuance
    System promptYesSent on every request
    Messages, including tool results, images and documentsYesUp to 600 images or PDF pages per request (100 on 200k-token models), so you can hit request size limits before the token limit
    Tool definitionsYesLarge toolsets use context before any work starts
    Output for the turn, including extended thinkingYesThinking tokens are part of max_tokens and are billed as output
    Previous thinking blocksDepends on the modelKept by default on Opus 4.5+ and Sonnet 4.6+; stripped automatically on earlier Opus/Sonnet and all Haiku models
    Cached prefix (cache_read_input_tokens, cache_creation_input_tokens)YesCaching changes what you pay, not whether the tokens count

    Window size depends on the model. Claude Fable 5.1, Claude Opus 5.5, Claude Opus 4.6 and later, Claude Sonnet 5, Claude Sonnet 4.6 and the Mythos models have a 1M-token window by default, with no beta header, and each request can generate up to 128k output tokens. Other models, including Claude Sonnet 4.5, have 200k.

    Thinking is where developers most often get the count wrong. On models that keep earlier thinking blocks, those blocks count as ordinary input on every later request. On models that strip them, the API removes them for you when you pass them back, so you don't need to strip them yourself. One rule applies to every model. During a tool use cycle, you must send the thinking block back with the tool result, complete and unmodified, including its signature. If you change the block, the API returns an error.

    Before a long session compacts, a team wants to archive the full, unsummarized transcript for audit purposes. Which mechanism should they use?

    Sources1

    3.How a model knows how much room is left

    Claude Sonnet 5, Claude Sonnet 4.6, Claude Sonnet 4.5 and Claude Haiku 4.5 have context awareness. They keep track of how much of their window remains during a conversation, so they can plan long tasks around the real space left instead of guessing. It happens automatically: the API adds the tags itself, and you never send them. The system prompt tells the model its total budget, for example <budget:token_budget>200000</budget:token_budget>, and after each tool call the model gets an update like this:

    The update the API injects after each tool call on models with context awarenessxml
    <system_warning>Token usage: 35000/200000; 165000 remaining</system_warning>

    Image tokens are included in these budgets. Not every model gets them. Claude Opus 4.7 and later Opus models, and the Fable and Mythos 5-series models, receive no injected tags. For those models you can set an explicit budget with task budgets, which are in beta.

    Sources1

    4.Context isolation in multi-step agentic workflows

    Context isolation in multi-step agentic workflows means keeping each step's working material out of the main conversation, so the context the model reasons over holds only what the current step needs. Pruning removes stale content after it has arrived. Isolation stops much of it from arriving in the first place.

    Programmatic tool calling isolates a chain of tool calls. Claude writes a single script, Anthropic's code execution sandbox runs it, and the script calls all the tools in the chain. Instead of five tool_use and tool_result roundtrips, the conversation gets one result, and the intermediate results never enter the conversation history. Consider it when you see repeated chains of small tool calls that could run as one batch.

    The memory tool isolates knowledge across steps and sessions. Claude records what it learns in files under /memories and reads them back when it needs them, instead of loading everything into the prompt at the start. The documentation calls this just-in-time context retrieval, and it keeps the active context focused on the current task. The tool runs on the client: Claude only requests file operations, your application carries them out against storage you control, and a later session picks up the same memory when your handler serves the same store. Tool search applies the same idea to tool definitions, keeping them out of the context window until Claude asks for them.

    Where each technique keeps material out of the main context
    TechniqueWhat stays out of the conversationWhen it fits
    Programmatic tool callingIntermediate tool_result blocks from a chain of callsChains of tool calls that can execute as a single script
    Memory tool (memory_20250818)Knowledge not needed for the current step, stored under /memoriesLong-running work that spans multiple agent sessions
    Tool searchTool definitions Claude hasn't asked for yetLarge toolsets (20+ tools) where most tools aren't needed every turn

    Sources56

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A 1M-token window means you should fill it, because more context gives better answers.Why is that wrong?

      Accuracy and recall get worse as the token count grows (context rot), so extra material that isn't relevant makes answers worse.

      Covered in From prompt engineering to context engineering

    2. 2.Prompt caching frees up room in the context window because cached tokens are read from the cache.Why is that wrong?

      Cached prefixes still take up space in the window. Caching lowers the cost of those tokens, but they still count.

      Covered in What uses up the context window

    3. 3.Every current Claude model automatically receives token-budget updates after each tool call.Why is that wrong?

      Only some models get the injected tags. Opus 4.7 and later, and the Fable and Mythos 5-series models, do not; you give them a budget with task budgets.

      Covered in How a model knows how much room is left

    4. 4.Tool result clearing deletes old results from your client's history, so your application must resync its copy of the conversation.Why is that wrong?

      Context editing is applied on the server before the prompt reaches Claude. Your application keeps the full history and does not need to sync with the edited version.

      Covered in From prompt engineering to context engineering

    5. 5.With programmatic tool calling, every intermediate tool result is still added to the conversation history.Why is that wrong?

      The script runs in the code execution sandbox, so the intermediate results stay there and only the outcome returns to the conversation.

      Covered in Context isolation in multi-step agentic workflows

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “more context isn't automatically better. As token count grows, accuracy and recall degrade, a phenomenon known as context rot.”
      ↩︎ From prompt engineering to context engineering
      “It is available in beta for Claude 4.6 and later models and Claude Mythos Preview.”
      ↩︎ From prompt engineering to context engineering
      “the system prompt, every message in messages (including tool results, images, and documents), and your tool definitions”
      ↩︎ What uses up the context window
      “The output Claude generates for the turn, including its extended thinking, counts too.”
      ↩︎ What uses up the context window
      “the API keeps previous thinking blocks by default, and they count toward the context window like any other input tokens”
      ↩︎ What uses up the context window
      “You must return the thinking block with the corresponding tool results.”
      ↩︎ What uses up the context window
      “Other Claude models, including Claude Sonnet 4.5, have a 200k-token context window.”
      ↩︎ What uses up the context window
      “This lets the model manage long-running tasks against the space that remains rather than guess how many tokens are left.”
      ↩︎ How a model knows how much room is left
      “To estimate a request before you send it, use the token counting API.”
      ↩︎ How a model knows how much room is left
      “If your conversations regularly approach context window limits, use server-side compaction.”
      ↩︎ How a model knows how much room is left
      “This makes curating what's in context just as important as how much space is available.”
      ↩︎ Key concept
      “more context isn't automatically better. As token count grows, accuracy and recall degrade, a phenomenon known as context rot.”
      ↩︎ Exam trap 1
      “Cached prompt prefixes still occupy the context window: prompt caching changes what you pay for those tokens, not whether they count.”
      ↩︎ Exam trap 2
      “don't receive these injected tags. On these models, you can give the model an explicit budget with task budgets”
      ↩︎ Exam trap 3
    2. 2.
      “Context engineering refers to the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference”
      ↩︎ From prompt engineering to context engineering
      “This results in n² pairwise relationships for n tokens.”
      ↩︎ From prompt engineering to context engineering
    3. 3.
      “context is a finite resource with diminishing returns, and irrelevant content degrades model focus”
      ↩︎ From prompt engineering to context engineering
      “The API replaces each cleared result with placeholder text indicating to Claude that it was removed.”
      ↩︎ From prompt engineering to context engineering
      “Your client application maintains the full, unmodified conversation history.”
      ↩︎ From prompt engineering to context engineering
      “Your client application maintains the full, unmodified conversation history. You do not need to sync your client state with the edited version.”
      ↩︎ Exam trap 4
    4. 4.
      “it keeps the active context small, because response quality degrades as a conversation grows.”
      ↩︎ From prompt engineering to context engineering
    5. 5.
      “The intermediate results never enter the conversation history.”
      ↩︎ Context isolation in multi-step agentic workflows
      “Tool search keeps tool definitions out of the context window until Claude asks for them.”
      ↩︎ Context isolation in multi-step agentic workflows
      “The intermediate results never enter the conversation history.”
      ↩︎ Exam trap 5
    6. 6.
      “This keeps the active context focused on the current task, which matters for long-running sessions that would otherwise overwhelm the context window.”
      ↩︎ Context isolation in multi-step agentic workflows
      “Maintain project context across multiple agent sessions”
      ↩︎ Context isolation in multi-step agentic workflows
      “The memory tool operates client-side: Claude requests file operations, and your application executes them.”
      ↩︎ Context isolation in multi-step agentic workflows

    Continue to page 2 of 2

    Pruning, Compaction and Context Isolation in Long Claude Agent Runs