CertSafari
    CLAUDE-CERTIFIED-DEVELOPER-FOUNDATIONS-CCDV-F · Lessons

    Domain 6 · Lesson 16/25

    Pruning, Compaction and Context Isolation in Long Claude Agent Runs

    Context Engineering

    10 min read
    3.67% of exam
    5 sources
    Published 29 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Configure tool result clearing and thinking block clearing with context editing
    • Weigh each clearing strategy against its effect on prompt caching
    • Choose between on-demand, threshold and client-side compaction
    • Keep intermediate output out of the main conversation with programmatic tool calling, tool search and the memory tool

    1.Pruning stale tool output with context editing

    In a long agent run, tool results are usually what fills the context fastest. A file read or a search result is useful when Claude first processes it and just takes up space afterwards, and a long loop can pile up hundreds of them. The context editing documentation treats clearing them as a way to decide what Claude sees, not only a way to save money: context is finite, extra tokens help less and less, and irrelevant content degrades model focus.

    The server-side strategy clear_tool_uses_20250919 handles this. When the context passes a threshold you set, the API clears the oldest tool results first, in chronological order. It replaces each one with placeholder text so Claude knows something was removed. By default only the results are cleared. Set clear_tool_inputs to true to clear the tool call parameters as well. Requests use the context-management-2025-06-27 beta header, and context editing is available on all supported Claude models.

    Tool result clearing with a custom threshold, retention count, minimum clear size and an excluded toolpython
        context_management={
            "edits": [
                {
                    "type": "clear_tool_uses_20250919",
                    # Trigger clearing when threshold is exceeded
                    "trigger": {"type": "input_tokens", "value": 30000},
                    # Number of tool uses to keep after clearing
                    "keep": {"type": "tool_uses", "value": 3},
                    # Optional: Clear at least this many tokens
                    "clear_at_least": {"type": "input_tokens", "value": 5000},
                    # Exclude these tools from being cleared
                    "exclude_tools": ["web_search"],
                }
            ]
        },
    Tool result clearing parameters and what each one controls
    ParameterWhat it controls
    triggerThe input-token threshold at which clearing starts
    keepHow many recent tool uses stay after clearing
    clear_at_leastThe minimum number of tokens removed each time clearing runs
    exclude_toolsTools whose results are never cleared
    clear_tool_inputsWhen true, clears the tool call parameters as well as the results

    Editing happens on the server before the prompt reaches Claude. Your application keeps its own full copy of the history and carries on managing it as normal.

    Sources1

    2.Clearing thinking blocks, and the cost to the prompt cache

    The second server-side strategy, clear_thinking_20251015, manages thinking blocks when extended thinking is on. Its keep parameter sets how many recent thinking turns survive. Keeping more preserves the chain of reasoning; keeping fewer frees more space.

    Thinking block clearing that keeps the last two thinking turnspython
        context_management={
            "edits": [
                {
                    "type": "clear_thinking_20251015",
                    "keep": {"type": "thinking_turns", "value": 2},
                }
            ]
        },

    Both strategies interact with prompt caching. Clearing tool results invalidates the cached prompt prefix. You pay a cache write each time, although later requests can reuse the newly cached prefix, so clear enough tokens at once to make that cost worthwhile. That is what clear_at_least is for. Thinking blocks that stay in context keep the cache intact; clearing them invalidates the cache from the point where the clearing happens. On Claude Fable 5.1 and Claude Opus 5.5, server-side context management never invalidates thinking blocks. Edits you make on the client side to earlier turns can, however, invalidate the thinking blocks in every later assistant turn.

    Sources1

    3.Compaction: summarising older turns so the conversation can continue

    Context editing removes content according to rules. Compaction replaces the older turns with a summary that Claude writes on the server, so you don't write any summarization code. It keeps a long conversation or agent task inside the window and keeps the active context small. The context windows documentation calls server-side compaction the primary strategy for long-running conversations and agentic workflows. It is in beta for Claude 4.6 and later models and Claude Mythos Preview.

    The three ways to compact a conversation
    ApproachWho decides whenRuns in the backgroundChoose it when
    Compaction on demand (beta header compact-2026-09-04)You, by sending a requestYesYour application needs to control when compaction happens, can't pause while a summary is written, or must keep recent turns and their thinking
    Compaction at a token thresholdThe API, when input tokens reach the trigger you setNo: it runs inside the request that reaches the thresholdYou want the API to manage context inside ordinary requests
    Your own summarizer (compact on the client)YouYes, in your own codeYou already run your own summarizer and it replaces the whole history with the summary

    Use on-demand compaction wherever it is available. You can adjust it in three ways. You can keep the most recent turns word for word so the summary covers only the older ones. You can run compaction in the background while work continues. And you can write your own summarization prompt if the default summary drops something a later turn needs. When a conversation that already starts with a compaction block grows long again, you compact again. There is also client-side SDK compaction in the TypeScript and Ruby SDKs when you use tool_runner. It generates a summary and replaces the full history, but server-side compaction is generally preferred.

    An engineer is building a research agent that needs to read dozens of files across a large repository to answer one question, without letting all that file content pile up in the main conversation's context. Which design best achieves this?

    Sources231

    4.Isolating work so intermediate output never enters the context

    Pruning and compaction clean up context after it has grown. Isolation stops bulky intermediate material from reaching the main conversation at all. The sources for this lesson do not describe how to configure subagents, so this section covers only the isolation mechanisms they document.

    With programmatic tool calling, a chain of tool calls becomes one script that Claude writes and Anthropic's code execution sandbox runs. The intermediate results never enter the conversation history. Tool search keeps tool definitions out of the window until Claude asks for them. It suits toolsets of 20 or more tools and costs one extra turn of latency. The memory tool supports just-in-time retrieval: the agent writes what it learns to files under /memories and reads them back when it needs them, which keeps the active context focused on the current task. The memory tool runs on the client side, so your handler carries out each operation and must reject any path outside /memories. Configuring it takes a single tools entry:

    The complete tools entry for the memory tooljson
    {"type": "memory_20250818", "name": "memory"}
    Four approaches to tool-driven context pressure, and what each one reduces
    ApproachWhat it reducesWhen it fits
    Tool searchTool definitions loaded upfrontLarge toolsets (20+ tools) where most tools aren't needed every turn
    Programmatic tool callingtool_result roundtripsChains of tool calls that can execute as a single script
    Prompt cachingToken cost of repeated tool definitionsStable toolsets across many requests
    Context editingOld tool_result blocks in historyLong conversations where early results are no longer relevant

    You can use these approaches together. The suggested starting point for a high-volume agent is: cache tool definitions from day one (cache writes cost 25% more than base input, which pays back on the second request that hits the cache); add tool search once you have more than about 20 tools; add context editing once conversations run long enough that early results stop mattering; and consider programmatic tool calling when you see repeated chains of small calls.

    A code review workflow runs a style-checker subagent, a security-scanner subagent, and a test-coverage subagent for the same pull request. The team wants the review to finish as fast as possible while still keeping each subagent's exploration out of the main conversation. What should they do?

    Sources45

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.After context editing clears tool results, your application must rewrite its local history to match the edited version.Why is that wrong?

      Editing happens on the server. Your application keeps the full, unmodified history and doesn't need to sync it.

      Covered in Pruning stale tool output with context editing

    2. 2.Tool result clearing has no effect on prompt caching, so clearing small amounts often costs nothing.Why is that wrong?

      Each clearing invalidates the cached prefix and costs a cache write, so use clear_at_least to make every clearing worth the cost.

      Covered in Clearing thinking blocks, and the cost to the prompt cache

    3. 3.Compaction and context editing are the same feature, and both summarise old tool results.Why is that wrong?

      Compaction replaces older turns with a summary. Context editing clears old tool results or thinking blocks according to rules, without summarising them.

      Covered in Compaction: summarising older turns so the conversation can continue

    Practise it for real

    Add tool result clearing to a tool-using Messages API request and tune it

    1. 1.Call client.beta.messages.create with betas=["context-management-2025-06-27"], a web_search_20250305 tool, and context_management={"edits": [{"type": "clear_tool_uses_20250919"}]}.

      Why: Giving only the strategy type enables clearing with every other option at its default.

      You should see: A normal response whose usage field reports what the request consumed.

    2. 2.Add "trigger": {"type": "input_tokens", "value": 30000} and "keep": {"type": "tool_uses", "value": 3} to the edit.

      Why: These set when clearing starts and how many recent tool uses survive it.

      You should see: Once input passes 30,000 tokens, older tool results are replaced with placeholder text and the three most recent tool uses stay.

    3. 3.Add "clear_at_least": {"type": "input_tokens", "value": 5000}.

      Why: Every clearing invalidates the cached prefix, so each one should free a worthwhile amount of space.

      You should see: Each clearing removes at least 5,000 tokens.

    4. 4.Add "exclude_tools": ["web_search"].

      Why: Some tool results should stay in context for the whole run.

      You should see: web_search results are never cleared.

    5. 5.Check your local messages list after a long run.

      Why: Editing happens on the server before the prompt reaches Claude.

      You should see: Your local history still contains every original tool result.

    Stuck? Get a nudge

    If clearing never seems to happen, compare the input count in the response's usage field with your trigger value.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Older tool results (like file contents or search results) are no longer needed once Claude has processed them.”
      ↩︎ Pruning stale tool output with context editing
      “The API replaces each cleared result with placeholder text indicating to Claude that it was removed.”
      ↩︎ Pruning stale tool output with context editing
      “You can optionally clear both tool results and tool calls (the tool use parameters) by setting clear_tool_inputs to true.”
      ↩︎ Pruning stale tool output with context editing
      “Context editing is applied server-side before the prompt reaches Claude.”
      ↩︎ Pruning stale tool output with context editing
      “Configure the keep parameter based on whether you want to prioritize cache performance or context window availability.”
      ↩︎ Clearing thinking blocks, and the cost to the prompt cache
      “Use the clear_at_least parameter to ensure a minimum number of tokens is cleared each time.”
      ↩︎ Clearing thinking blocks, and the cost to the prompt cache
      “On Claude Fable 5.1 and Claude Opus 5.5, server-side context management never invalidates thinking blocks.”
      ↩︎ Clearing thinking blocks, and the cost to the prompt cache
      “An SDK-based alternative for summary-based context management (server-side compaction is generally preferred)”
      ↩︎ Compaction: summarising older turns so the conversation can continue
      “Your client application maintains the full, unmodified conversation history. You do not need to sync your client state with the edited version.”
      ↩︎ Exam trap 1
      “Invalidates cached prompt prefixes when content is cleared.”
      ↩︎ Exam trap 2
    2. 2.
      “Compaction replaces the older turns of a conversation with a summary that Claude writes on the server”
      ↩︎ Compaction: summarising older turns so the conversation can continue
      “Use on-demand compaction wherever it is available.”
      ↩︎ Compaction: summarising older turns so the conversation can continue
      “Write your own summarization prompt: when the default summary drops something a later turn needs.”
      ↩︎ Compaction: summarising older turns so the conversation can continue
      “To clear old tool results or old thinking blocks by rule instead of summarizing them, see Context editing.”
      ↩︎ Exam trap 3
    3. 3.
      “For long-running conversations and agentic workflows, server-side compaction is the primary strategy for context management.”
      ↩︎ Compaction: summarising older turns so the conversation can continue
    4. 4.
      “The intermediate results never enter the conversation history.”
      ↩︎ Isolating work so intermediate output never enters the context
      “Tool search keeps tool definitions out of the context window until Claude asks for them.”
      ↩︎ Isolating work so intermediate output never enters the context
    5. 5.
      “Rather than loading all relevant information up front, an agent records what it learns in memory files and reads them back on demand.”
      ↩︎ Isolating work so intermediate output never enters the context

    Ready to test yourself?

    Practise the 20 questions on this subdomain.