CertSafari
    CCAR-P · Lessons

    Domain 2 · Lesson 10/38

    Prompt Caching, Context Editing and Compaction in Claude

    Optimize context windows and manage token usage

    9 min read
    2.6% of exam
    4 sources
    Published 27 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Enable automatic and explicit prompt caching and choose between the 5-minute and 1-hour TTL
    • Work out what a cache hit saves on a given model
    • Configure tool result clearing and thinking block clearing, and predict how each affects the cache
    • Choose between on-demand, threshold and client-side compaction

    1.Prompt caching: pay less for a prefix you repeat

    Many requests begin with the same content: a long system prompt, a set of examples, a reference document, or the conversation so far. Prompt caching lets Claude resume from a stored prefix instead of processing it again, which cuts both cost and processing time. The documentation names four good fits: prompts with many examples, large background context, repetitive tasks with the same instructions, and long multi-turn conversations.

    There are two ways to turn it on. Automatic caching is a single cache_control field at the top level of the request. The API places the breakpoint on the last cacheable block and moves it forward as the conversation grows, so each turn reads the earlier prefix from cache and writes only the new part. Explicit breakpoints put cache_control on individual blocks when you want exact control. The two can be combined: the automatic breakpoint takes one of the four breakpoint slots.

    An explicit breakpoint on the system prompt plus automatic caching for the conversationjson
    {
      "model": "claude-opus-5-5",
      "max_tokens": 1024,
      "cache_control": { "type": "ephemeral" },
      "system": [
        {
          "type": "text",
          "text": "You are a helpful assistant.",
          "cache_control": { "type": "ephemeral" }
        }
      ],
      "messages": [{ "role": "user", "content": "What are the key terms?" }]
    }

    Some combinations fail. If the last block already has an explicit cache_control with a different TTL, the API returns a 400 error. It also returns a 400 if four explicit breakpoints already fill every slot.

    Lifetime. The default cache lives for five minutes, and each use refreshes it at no extra cost. You can set a 1-hour TTL, which charges cache writes at 2x the base input price. The clock starts when the request that reads or writes the cache *starts*. So if a response takes four minutes to stream, the follow-up request has roughly one minute left to hit the cache.

    Prompt caching prices per million tokens for the current headline models
    ModelBase input5m writes1h writesHits and refreshes
    Claude Opus 5.5$4 / MTok$5 / MTok$8 / MTok$0.20 / MTok (0.05x)
    Claude Sonnet 5$2 / MTok$2.50 / MTok$4 / MTok$0.20 / MTok
    Claude Haiku 4.5$1 / MTok$1.25 / MTok$2 / MTok$0.10 / MTok
    Claude Fable 5.1$10 / MTok$12.50 / MTok$20 / MTok$0.25 / MTok (0.025x)

    A developer calling Claude Opus 4.8 sends a request where the input tokens alone fit comfortably within the model's context window, but input tokens plus the requested max_tokens together exceed it. What should they expect from the API?

    Sources1

    2.Context editing: clear what Claude no longer needs

    Caching makes a long context cheaper but leaves it just as long. To actually shrink the window, remove content. Context editing does this on the server before the prompt reaches Claude. The docs frame it as curation: irrelevant content hurts model focus as well as your bill.

    Tool result clearing (clear_tool_uses_20250919) targets agent loops. After Claude has processed a file read or a search result, the result is usually no longer needed. Once input passes your trigger, the API clears the oldest tool results first and replaces each with placeholder text saying it was removed. keep keeps the most recent tool uses, exclude_tools protects specific tools, and clear_tool_inputs also clears the tool-call parameters.

    Tool result clearing with a trigger, a keep window, a minimum clear and an exclusionpython
        context_management={
            "edits": [
                {
                    "type": "clear_tool_uses_20250919",
                    # Trigger clearing when threshold is exceeded
                    "trigger": {"type": "input_tokens", "value": 30000},
                    # Number of tool uses to keep after clearing
                    "keep": {"type": "tool_uses", "value": 3},
                    # Optional: Clear at least this many tokens
                    "clear_at_least": {"type": "input_tokens", "value": 5000},
                    # Exclude these tools from being cleared
                    "exclude_tools": ["web_search"],
                }
            ]
        },
    )

    Thinking block clearing (clear_thinking_20251015) handles the reasoning that extended thinking leaves behind. Its keep setting, for example {"type": "thinking_turns", "value": 2}, sets how much recent thinking survives.

    Both strategies interact with caching, and this is where the budgeting trade-off lies. Clearing tool results invalidates the cached prefix, so each clear costs a new cache write. Use clear_at_least so each invalidation removes enough tokens to pay for itself. With thinking, keeping blocks preserves the cache and clearing them breaks it at the point of clearing, so keep is a choice between cache performance and free space in the window. In both cases your client keeps the full, unedited history; you do not need to sync it with the edited version the server sends to Claude.

    Sources2

    3.Compaction: summarise instead of deleting

    Clearing removes content by rule. Compaction replaces the older turns with a summary that Claude writes on the server, so a conversation can go past the window limit without you writing any summarisation code. For long conversations and agent workflows, the context windows doc calls server-side compaction the primary strategy. It is in beta for Claude 4.6 and later models. There are three ways to run it:

    Choosing a compaction approach
    ApproachWho decides whenChoose it when
    Compaction on demandYou, by sending a requestYou must control timing, can't pause while a summary is written, or must keep recent turns and their thinking
    Compaction at a token thresholdThe API, when input tokens reach your triggerYou want the API to manage context inside ordinary requests
    Your own summarizer (client-side)YouYou already run a summarizer that replaces the whole history

    The docs recommend on-demand compaction wherever it is available. It uses the compact-2026-09-04 beta header and a top-level compaction parameter. You send back the returned block first in messages, in place of the turns it summarises. On-demand compaction can keep recent turns word for word and can run in the background. If the default summary drops something a later turn needs, you can supply your own summarisation prompt. Threshold compaction runs inside the request that crosses the trigger, so it cannot run in the background.

    A production agent uses 45 different tools across long-running conversations. The team identifies three separate cost sources: tool schema definitions consuming context on every request even when most tools go unused that turn, many small sequential tool calls each leaving a tool_result in history, and the same stable tool definitions being resent and re-priced identically on every request. Which changes directly address these three sources?(Select 3)

    Sources34

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.The five-minute cache timer starts when the previous response finishes streaming.Why is that wrong?

      The lifetime is measured from the start of the request that wrote or read the cache, so time spent generating a long response uses it up.

      Covered in Prompt caching: pay less for a prefix you repeat

    2. 2.Tool result clearing and prompt caching work together with no cost, so clearing a little on every turn is harmless.Why is that wrong?

      Clearing tool results invalidates the cached prefix, so each clear triggers a new cache write. Use clear_at_least so each clear frees enough tokens to justify that cost.

      Covered in Context editing: clear what Claude no longer needs

    Practise it for real

    Measure a request before sending it, then confirm prompt caching reduces the cost of a repeated prefix.

    1. 1.Call client.messages.count_tokens with your model, a long system prompt and one user message.

      Why: Measuring first tells you how big the prefix is before you pay for it.

      You should see: A JSON response with a single input_tokens value.

    2. 2.Send the same request with client.messages.create and a top-level cache_control={"type": "ephemeral"}, then print response.usage.

      Why: Automatic caching places the breakpoint on the last cacheable block and writes the prefix to the cache.

      You should see: A non-zero cache_creation_input_tokens on this first call.

    3. 3.Within five minutes of the first request starting, send the same prefix with a new user message and print usage again.

      Why: A request that reuses the cached prefix within the TTL reads it instead of reprocessing it.

      You should see: A non-zero cache_read_input_tokens. The sum of the three input fields is still what the request takes up in the window.

    Stuck? Get a nudge

    If both calls show zero cache tokens, the prefix is probably below the minimum token threshold for caching. Make the system prompt longer and try again.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Add a single cache_control field at the top level of your request.”
      ↩︎ Prompt caching: pay less for a prefix you repeat
      “When used together, the automatic cache breakpoint uses one of the 4 available breakpoint slots.”
      ↩︎ Prompt caching: pay less for a prefix you repeat
      “By default, the cache has a 5-minute lifetime. The cache is refreshed for no additional cost each time the cached content is used.”
      ↩︎ Prompt caching: pay less for a prefix you repeat
      “You can specify a 1-hour TTL at 2x the base input token price”
      ↩︎ Prompt caching: pay less for a prefix you repeat
      “All other models use the standard 0.1x multiplier.”
      ↩︎ Prompt caching: pay less for a prefix you repeat
      “The lifetime is measured from the start of the request that writes or reads the cache entry, not from the end of its response.”
      ↩︎ Exam trap 1
    2. 2.
      “context is a finite resource with diminishing returns, and irrelevant content degrades model focus”
      ↩︎ Context editing: clear what Claude no longer needs
      “The API replaces each cleared result with placeholder text indicating to Claude that it was removed.”
      ↩︎ Context editing: clear what Claude no longer needs
      “Configure the keep parameter based on whether you want to prioritize cache performance or context window availability.”
      ↩︎ Context editing: clear what Claude no longer needs
      “Your client application maintains the full, unmodified conversation history.”
      ↩︎ Context editing: clear what Claude no longer needs
      “Use the clear_at_least parameter to ensure a minimum number of tokens is cleared each time.”
      ↩︎ Exam trap 2
    3. 3.
      “Compaction replaces the older turns of a conversation with a summary that Claude writes on the server”
      ↩︎ Compaction: summarise instead of deleting
      “Use on-demand compaction wherever it is available.”
      ↩︎ Compaction: summarise instead of deleting
    4. 4.
      “server-side compaction is the primary strategy for context management”
      ↩︎ Compaction: summarise instead of deleting

    Ready to test yourself?

    Practise the 12 questions on this subdomain.