What you will be able to do
- Explain how prompt caching reuses a prompt prefix up to a cache breakpoint
- Choose between automatic caching and explicit cache breakpoints, and combine them within the four-slot limit
- Reason about cache lifetime, the 1-hour TTL option and cache-hit pricing
- Predict which request changes invalidate the tools, system or messages cache
Key concept
Prefix caching — Prompt caching stores the processed prefix of a request, up to a cache breakpoint. Later requests that begin with exactly the same prefix read it from cache instead of processing it again. Reuse therefore depends on keeping the stable parts of a prompt at the front and unchanged.
1.What gets reused, and when
The cheapest prompt to reuse is one you never process twice. With prompt caching enabled, the API checks whether the start of your request, up to a cache breakpoint, is already cached from a recent request. If it is, that cached version is used, which cuts processing time and cost. If it isn't, the API processes the whole prompt and caches the prefix once the response begins.
The documentation lists four situations where this pays off: prompts with many examples, large amounts of context or background information, repetitive tasks with consistent instructions, and long multi-turn conversations. All four share one property: a large block of content that stays the same from one request to the next. Caching is how you turn a reusable prompt template into a reusable *processed* prompt.
Sources1
2.Automatic caching: one field, a moving breakpoint
You can enable caching in two ways. The simpler one is automatic caching: add a single cache_control field at the top level of the request. The system places the breakpoint on the last cacheable block and moves it forward as the conversation grows. The documentation recommends this for multi-turn conversations.
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=1024,
cache_control={"type": "ephemeral"},
system="You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.",
messages=[
{
"role": "user",
"content": "Analyze the major themes in 'Pride and Prejudice'.",
}
],
)
print(response.usage.model_dump_json())| Request | Read from cache | Written to cache |
|---|---|---|
| Request 1 | Nothing | System + User(1) + Asst(1) + User(2) |
| Request 2 | System through User(2) | Asst(2) + User(3) |
| Request 3 | System through User(3) | Asst(3) + User(4) |
Automatic caching isn't a separate cache. It uses the same infrastructure as explicit breakpoints, so the same rules apply to it: pricing, minimum token thresholds, context ordering, and a 20-block lookback window. These sources name the lookback window but don't explain how it works. What they do give is the remedy for long sessions: combine the automatic breakpoint with explicit ones on content that never changes, which the next section covers.
A long agent session accumulates conversation turns using only automatic request-level caching. By turn 15, a cache entry exists at content block 12. By turn 25, the conversation has grown to 35 content blocks, and the current breakpoint sits on the last block. Cache reads stop matching even though the early conversation content hasn't changed. What is the best explanation and fix?
Correct answer: A — The cache lookup only scans 20 blocks backward from the breakpoint, so the block-12 entry falls outside that window; adding an explicit breakpoint further back re-establishes a match.
- A. Correct. The system checks up to 20 blocks backward from a breakpoint for a prior cache write. With the breakpoint on block 35 and the last written entry at block 12, the gap is 23 blocks, outside the lookback window, so it misses. Adding an intermediate explicit breakpoint keeps the gap within 20 blocks and restores hits.
- B. Incorrect. Automatic caching moves its breakpoint forward as a conversation grows and is designed to handle multi-turn conversations, not just a single most recent turn.
- C. Incorrect. There is no fixed block count at which cache entries are deleted outright; the miss described here is caused by the 20-block lookback window relative to breakpoint placement, not entry deletion.
- D. Incorrect. Breakpoints should be placed on the last stable block, not the first, since caching keys off the prefix ending at the breakpoint moving forward as new stable content accumulates.
Sources1
3.Explicit breakpoints, and combining the two
Explicit breakpoints put cache_control directly on individual content blocks, so you decide exactly what gets cached. This is how you cache a prompt in modules. A common pattern is an explicit breakpoint on the system prompt, which is stable and shared by every request, while the top-level automatic breakpoint follows the growing conversation.
{
"model": "claude-opus-5-5",
"max_tokens": 1024,
"cache_control": { "type": "ephemeral" },
"system": [
{
"type": "text",
"text": "You are a helpful assistant.",
"cache_control": { "type": "ephemeral" }
}
],
"messages": [{ "role": "user", "content": "What are the key terms?" }]
}A request can have at most four breakpoints, and the automatic breakpoint takes one of those slots. That produces a few edge cases worth knowing:
| Situation | Result |
|---|---|
| Last block already has explicit cache_control with the same TTL | Automatic caching is a no-op |
| Last block has explicit cache_control with a different TTL | 400 error |
| 4 explicit block-level breakpoints already exist | 400 error: no slot left for automatic caching |
| Last block is not an eligible breakpoint target | Walks backward to the nearest eligible block; if none, caching is skipped |
Tool definitions can be cached too. To cache the whole tool-definitions prefix, put cache_control on the last tool in the tools array. For an mcp_toolset, where you don't control tool order, put it on the toolset entry itself.
4.Lifetime and price
The cache lasts 5 minutes by default. Each time the cached content is used, the lifetime is refreshed at no extra cost. The clock starts when the request that writes or reads the entry *starts*, not when its response ends.
About 1 minute. The 5-minute lifetime started with the previous request, and the 4 minutes spent generating the response count against it. If your gaps are longer, you can set a 1-hour TTL with { "cache_control": { "type": "ephemeral", "ttl": "1h" } }. It costs 2x the base input token price to write.
| Model | Base input | 5m writes | 1h writes | Hits and refreshes |
|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $12.50 | $20 | $0.25 |
| Claude Opus 5.5 | $4 | $5 | $8 | $0.20 |
| Claude Sonnet 5 | $2 | $2.50 | $4 | $0.20 |
| Claude Haiku 4.5 | $1 | $1.25 | $2 | $0.10 |
Cache writes cost a little more than base input. Reads cost far less: 0.05x base input on Opus 5.5, 0.025x on Fable 5.1 and Mythos 5.1, and the standard 0.1x on every other model. You pay a small premium once so that every later hit is cheap.
Sources1
5.What breaks the cache
Because caching is prefix-based, the order of a request matters: tools come first, then system, then messages. A change at one level invalidates that level and every level after it, but none before it.
| Change | Invalidates |
|---|---|
| Modifying tool definitions | Entire cache (tools, system, messages) |
| Toggling web search or citations | System and messages caches |
| Changing tool_choice | Messages cache |
| Changing disable_parallel_tool_use | Messages cache |
| Toggling images present/absent | Messages cache |
| Changing thinking parameters | Messages cache always; tool and system caches too on models that render the thinking configuration ahead of them |
An agentic workflow sends requests with a large tools array (15 tool definitions, cached), a cached system prompt, and growing conversation history. For one request, the developer adds tool_choice: {"type": "tool", "name": "search"} to force a specific tool call, while leaving the tools array, system prompt, and message content unchanged. What happens to prompt caching for this request?
Correct answer: A — Only the message-level cache is invalidated, so tokens up through the system prompt are still served from cache while the messages must be reprocessed.
- A. Correct. Cache invalidation cascades downward through tools, then system, then messages. A tool_choice change only invalidates the message-level cache; the tools and system caches, which sit above it in the hierarchy, remain valid and are still served from cache.
- B. Incorrect. Changing tool_choice is not equivalent to editing the tool definitions themselves; it only affects the message-level cache, not the tools-array cache.
- C. Incorrect. tool_choice does affect caching: it is one of the request properties whose change invalidates the message-level cache, so it is not entirely outside the hierarchy.
- D. Incorrect. tool_choice does not invalidate the tools-array cache; the tool definitions are unchanged. It invalidates the message-level cache instead, while the tools and system caches stay valid.
Adding tools on demand doesn't have to break the cache. Deferred tools aren't part of the system-prompt prefix. When the model discovers one through tool search, its definition is appended to the conversation as a tool_reference block and the cached prefix stays untouched.
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Automatic caching and explicit cache_control breakpoints are mutually exclusive, so you have to pick one.Why is that wrong?
You can use both in the same request. The automatic breakpoint simply takes one of the four breakpoint slots.
Covered in Explicit breakpoints, and combining the two
2.The 5-minute cache lifetime starts counting when the previous response finishes streaming.Why is that wrong?
The lifetime starts when the writing or reading request starts, so time spent generating the response uses it up.
Covered in Lifetime and price
3.Changing tool_choice changes the tool configuration, so the whole cache is invalidated.Why is that wrong?
Only modifying the tool definitions themselves invalidates the entire cache. Changing tool_choice invalidates only the messages cache, so the cached tools and system prefix still hit.
Covered in What breaks the cache
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The system checks if a prompt prefix, up to a specified cache breakpoint, is already cached from a recent query.”
↩︎ What gets reused, and when“The cache breakpoint automatically moves to the last cacheable block in each request”
↩︎ Automatic caching: one field, a moving breakpoint“Pricing, minimum token thresholds, context ordering requirements, and the 20-block lookback window all apply the same as with explicit breakpoints.”
↩︎ Automatic caching: one field, a moving breakpoint“use an explicit breakpoint to cache your system prompt, while automatic caching handles the conversation”
↩︎ Explicit breakpoints, and combining the two“If the last block has an explicit cache_control with a different TTL, the API returns a 400 error.”
↩︎ Explicit breakpoints, and combining the two“By default, the cache has a 5-minute lifetime. The cache is refreshed for no additional cost each time the cached content is used.”
↩︎ Lifetime and price“You can specify a 1-hour TTL at 2x the base input token price”
↩︎ Lifetime and price“Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.”
↩︎ Lifetime and price“Prompt caching optimizes your API usage by allowing resuming from specific prefixes in your prompts.”
↩︎ Key concept“Automatic caching is compatible with explicit cache breakpoints. When used together, the automatic cache breakpoint uses one of the 4 available breakpoint slots.”
↩︎ Exam trap 1“The lifetime is measured from the start of the request that writes or reads the cache entry, not from the end of its response.”
↩︎ Exam trap 2 - 2.https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-use-with-prompt-cachingOfficial docs
“Place cache_control: {"type": "ephemeral"} on the last tool in your tools array.”
↩︎ Explicit breakpoints, and combining the two“The cache follows a prefix hierarchy (tools → system → messages), so a change at one level invalidates that level and everything after it”
↩︎ What breaks the cache“Deferred tools are not included in the system-prompt prefix.”
↩︎ What breaks the cache“The cache follows a prefix hierarchy (tools → system → messages), so a change at one level invalidates that level and everything after it”
↩︎ Exam trap 3