What you will be able to do
- Enable automatic and explicit prompt caching and choose between the 5-minute and 1-hour TTL
- Work out what a cache hit saves on a given model
- Configure tool result clearing and thinking block clearing, and predict how each affects the cache
- Choose between on-demand, threshold and client-side compaction
1.Prompt caching: pay less for a prefix you repeat
Many requests begin with the same content: a long system prompt, a set of examples, a reference document, or the conversation so far. Prompt caching lets Claude resume from a stored prefix instead of processing it again, which cuts both cost and processing time. The documentation names four good fits: prompts with many examples, large background context, repetitive tasks with the same instructions, and long multi-turn conversations.
There are two ways to turn it on. Automatic caching is a single cache_control field at the top level of the request. The API places the breakpoint on the last cacheable block and moves it forward as the conversation grows, so each turn reads the earlier prefix from cache and writes only the new part. Explicit breakpoints put cache_control on individual blocks when you want exact control. The two can be combined: the automatic breakpoint takes one of the four breakpoint slots.
{
"model": "claude-opus-5-5",
"max_tokens": 1024,
"cache_control": { "type": "ephemeral" },
"system": [
{
"type": "text",
"text": "You are a helpful assistant.",
"cache_control": { "type": "ephemeral" }
}
],
"messages": [{ "role": "user", "content": "What are the key terms?" }]
}Some combinations fail. If the last block already has an explicit cache_control with a different TTL, the API returns a 400 error. It also returns a 400 if four explicit breakpoints already fill every slot.
Lifetime. The default cache lives for five minutes, and each use refreshes it at no extra cost. You can set a 1-hour TTL, which charges cache writes at 2x the base input price. The clock starts when the request that reads or writes the cache *starts*. So if a response takes four minutes to stream, the follow-up request has roughly one minute left to hit the cache.
| Model | Base input | 5m writes | 1h writes | Hits and refreshes |
|---|---|---|---|---|
| Claude Opus 5.5 | $4 / MTok | $5 / MTok | $8 / MTok | $0.20 / MTok (0.05x) |
| Claude Sonnet 5 | $2 / MTok | $2.50 / MTok | $4 / MTok | $0.20 / MTok |
| Claude Haiku 4.5 | $1 / MTok | $1.25 / MTok | $2 / MTok | $0.10 / MTok |
| Claude Fable 5.1 | $10 / MTok | $12.50 / MTok | $20 / MTok | $0.25 / MTok (0.025x) |
A developer calling Claude Opus 4.8 sends a request where the input tokens alone fit comfortably within the model's context window, but input tokens plus the requested max_tokens together exceed it. What should they expect from the API?
Correct answer: A — The API accepts the request, and if generation reaches the context window limit it stops early with stop_reason set to model_context_window_exceeded
- A. Correct. On Claude 4.5 models and newer, a request where input plus max_tokens exceeds the window is still accepted. If generation actually reaches the limit, the response stops with stop_reason model_context_window_exceeded rather than failing upfront.
- B. Incorrect. The 400 invalid_request_error for a prompt being too long is reserved for cases where the input alone already exceeds the context window, not when only input-plus-max_tokens would exceed it.
- C. Incorrect. The Claude API does not perform rolling first-in-first-out trimming of conversation history on your behalf; that behavior is associated with some chat interfaces, not standard API requests.
- D. Incorrect. Context window size is fixed per model and is not dynamically expanded for an individual request based on the max_tokens value requested.
Sources1
2.Context editing: clear what Claude no longer needs
Caching makes a long context cheaper but leaves it just as long. To actually shrink the window, remove content. Context editing does this on the server before the prompt reaches Claude. The docs frame it as curation: irrelevant content hurts model focus as well as your bill.
Tool result clearing (clear_tool_uses_20250919) targets agent loops. After Claude has processed a file read or a search result, the result is usually no longer needed. Once input passes your trigger, the API clears the oldest tool results first and replaces each with placeholder text saying it was removed. keep keeps the most recent tool uses, exclude_tools protects specific tools, and clear_tool_inputs also clears the tool-call parameters.
context_management={
"edits": [
{
"type": "clear_tool_uses_20250919",
# Trigger clearing when threshold is exceeded
"trigger": {"type": "input_tokens", "value": 30000},
# Number of tool uses to keep after clearing
"keep": {"type": "tool_uses", "value": 3},
# Optional: Clear at least this many tokens
"clear_at_least": {"type": "input_tokens", "value": 5000},
# Exclude these tools from being cleared
"exclude_tools": ["web_search"],
}
]
},
)Thinking block clearing (clear_thinking_20251015) handles the reasoning that extended thinking leaves behind. Its keep setting, for example {"type": "thinking_turns", "value": 2}, sets how much recent thinking survives.
Both strategies interact with caching, and this is where the budgeting trade-off lies. Clearing tool results invalidates the cached prefix, so each clear costs a new cache write. Use clear_at_least so each invalidation removes enough tokens to pay for itself. With thinking, keeping blocks preserves the cache and clearing them breaks it at the point of clearing, so keep is a choice between cache performance and free space in the window. In both cases your client keeps the full, unedited history; you do not need to sync it with the edited version the server sends to Claude.
Each clear invalidates the cached prefix, so every turn pays for a new cache write while freeing very little space. Set clear_at_least so that each clear removes enough tokens to justify rebuilding the cache.
Sources2
3.Compaction: summarise instead of deleting
Clearing removes content by rule. Compaction replaces the older turns with a summary that Claude writes on the server, so a conversation can go past the window limit without you writing any summarisation code. For long conversations and agent workflows, the context windows doc calls server-side compaction the primary strategy. It is in beta for Claude 4.6 and later models. There are three ways to run it:
| Approach | Who decides when | Choose it when |
|---|---|---|
| Compaction on demand | You, by sending a request | You must control timing, can't pause while a summary is written, or must keep recent turns and their thinking |
| Compaction at a token threshold | The API, when input tokens reach your trigger | You want the API to manage context inside ordinary requests |
| Your own summarizer (client-side) | You | You already run a summarizer that replaces the whole history |
The docs recommend on-demand compaction wherever it is available. It uses the compact-2026-09-04 beta header and a top-level compaction parameter. You send back the returned block first in messages, in place of the turns it summarises. On-demand compaction can keep recent turns word for word and can run in the background. If the default summary drops something a later turn needs, you can supply your own summarisation prompt. Threshold compaction runs inside the request that crosses the trigger, so it cannot run in the background.
A production agent uses 45 different tools across long-running conversations. The team identifies three separate cost sources: tool schema definitions consuming context on every request even when most tools go unused that turn, many small sequential tool calls each leaving a tool_result in history, and the same stable tool definitions being resent and re-priced identically on every request. Which changes directly address these three sources?(Select 3)
Correct answers: A, B, C — Enable the tool search tool so schemas load on demand instead of being sent upfront; Use programmatic tool calling so chains of tool calls run as a single script instead of many roundtrips; Apply prompt caching to the stable tool definitions so repeated requests pay the discounted cache-read rate
- A. Correct. Tool search keeps tool definitions out of the context window until Claude requests them, directly reducing the baseline cost of unused schemas loaded upfront.
- B. Correct. Programmatic tool calling collapses chains of tool calls into a single executed script, so intermediate tool_result blocks never enter the conversation history.
- C. Correct. Prompt caching does not shrink the token count of tool definitions but cuts what repeated, stable definitions cost on subsequent requests via the discounted cache-read rate.
- D. Incorrect. max_tokens controls the ceiling on generated output length and has no effect on how much context tool definitions or tool_result history consume.
- E. Incorrect. A larger context window gives more headroom but does not reduce the underlying token cost of the toolset or address any of the three described cost sources.
- F. Incorrect. Shortening tool names yields negligible savings and is not a documented strategy for managing tool-related context cost compared to the other approaches.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The five-minute cache timer starts when the previous response finishes streaming.Why is that wrong?
The lifetime is measured from the start of the request that wrote or read the cache, so time spent generating a long response uses it up.
2.Tool result clearing and prompt caching work together with no cost, so clearing a little on every turn is harmless.Why is that wrong?
Clearing tool results invalidates the cached prefix, so each clear triggers a new cache write. Use clear_at_least so each clear frees enough tokens to justify that cost.
Covered in Context editing: clear what Claude no longer needs
Practise it for real
Measure a request before sending it, then confirm prompt caching reduces the cost of a repeated prefix.
1.Call client.messages.count_tokens with your model, a long system prompt and one user message.
Why: Measuring first tells you how big the prefix is before you pay for it.
You should see: A JSON response with a single input_tokens value.
2.Send the same request with client.messages.create and a top-level cache_control={"type": "ephemeral"}, then print response.usage.
Why: Automatic caching places the breakpoint on the last cacheable block and writes the prefix to the cache.
You should see: A non-zero cache_creation_input_tokens on this first call.
3.Within five minutes of the first request starting, send the same prefix with a new user message and print usage again.
Why: A request that reuses the cached prefix within the TTL reads it instead of reprocessing it.
You should see: A non-zero cache_read_input_tokens. The sum of the three input fields is still what the request takes up in the window.
Stuck? Get a nudge
If both calls show zero cache tokens, the prefix is probably below the minimum token threshold for caching. Make the system prompt longer and try again.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Add a single cache_control field at the top level of your request.”
↩︎ Prompt caching: pay less for a prefix you repeat“When used together, the automatic cache breakpoint uses one of the 4 available breakpoint slots.”
↩︎ Prompt caching: pay less for a prefix you repeat“By default, the cache has a 5-minute lifetime. The cache is refreshed for no additional cost each time the cached content is used.”
↩︎ Prompt caching: pay less for a prefix you repeat“You can specify a 1-hour TTL at 2x the base input token price”
↩︎ Prompt caching: pay less for a prefix you repeat“All other models use the standard 0.1x multiplier.”
↩︎ Prompt caching: pay less for a prefix you repeat“The lifetime is measured from the start of the request that writes or reads the cache entry, not from the end of its response.”
↩︎ Exam trap 1 - 2.
“context is a finite resource with diminishing returns, and irrelevant content degrades model focus”
↩︎ Context editing: clear what Claude no longer needs“The API replaces each cleared result with placeholder text indicating to Claude that it was removed.”
↩︎ Context editing: clear what Claude no longer needs“Configure the keep parameter based on whether you want to prioritize cache performance or context window availability.”
↩︎ Context editing: clear what Claude no longer needs“Your client application maintains the full, unmodified conversation history.”
↩︎ Context editing: clear what Claude no longer needs“Use the clear_at_least parameter to ensure a minimum number of tokens is cleared each time.”
↩︎ Exam trap 2 - 3.
“Compaction replaces the older turns of a conversation with a summary that Claude writes on the server”
↩︎ Compaction: summarise instead of deleting“Use on-demand compaction wherever it is available.”
↩︎ Compaction: summarise instead of deleting - 4.
“server-side compaction is the primary strategy for context management”
↩︎ Compaction: summarise instead of deleting