What you will be able to do
- Explain why prompt caching is the largest cost lever for multi-turn and agentic workloads
- Choose between automatic caching and explicit cache breakpoints, and place breakpoints within the four-slot limit
- Price cache writes and reads per model and choose between the 5-minute and 1-hour TTL from real pause patterns
- Name the request changes that invalidate the cache, and use the cache-read share to check that caching is working
1.Why caching is the biggest lever
Claude has no memory between requests. Every call sends the whole prompt again. In a single-shot app that is harmless. In an agent it compounds: every turn resends the system prompt, the tool definitions and every earlier turn. The first turn of a 40-turn task is sent 40 times, so the cost of a task grows roughly with the square of the number of turns.
Prompt caching doesn't reduce what you send. It reduces what you pay for the repeated part. When a request has caching enabled, the system checks whether the prompt prefix up to a cache breakpoint was cached by a recent request. If it was, the cached version is used. If not, the full prompt is processed and the prefix is cached once the response starts. The cached prefix is then billed at the cache-read rate, and each turn pays the higher write rate only for its new content. In Anthropic's measurements caching was the largest lever by a wide margin. It cut agent-loop cost by a factor of 2.7 to 5.3, and cut a small triage agent's bill by 83%.
2.Automatic caching and explicit breakpoints
There are two ways to turn caching on. Automatic caching is a single top-level cache_control field. The system puts the breakpoint on the last cacheable block and moves it forward as the conversation grows. On each new request, everything up to the previous turn is read from cache and only the newest assistant and user turns are written. You never edit a marker. Explicit breakpoints put cache_control on individual content blocks, which gives you exact control over what is cached.
You can combine the two. A request has four breakpoint slots in total, and automatic caching uses one of them. If four explicit block-level breakpoints already exist, adding automatic caching returns a 400 error. So does an explicit marker on the last block with a different TTL. The same pricing, minimum token thresholds, ordering requirements and 20-block lookback window apply to both methods. These sources say that minimum thresholds exist but give no per-model numbers.
{
"model": "claude-opus-5-5",
"max_tokens": 1024,
"cache_control": { "type": "ephemeral" },
"system": [
{
"type": "text",
"text": "You are a helpful assistant.",
"cache_control": { "type": "ephemeral" }
}
],
"messages": [{ "role": "user", "content": "What are the key terms?" }]
}Tool definitions are a common thing to checkpoint. Put cache_control on the last tool in the tools array, and the whole tool-definitions prefix is cached, from the first tool up to that marker. With an mcp_toolset you don't control the order of the tools, so put the marker on the toolset entry and the API applies it to the last expanded tool. Tools loaded with defer_loading through tool search are added inline to the conversation rather than to the prefix, so discovering tools mid-session doesn't break the cache.
A finance team wants a monthly cost model for an internal Claude-powered summarization service. They know the average input tokens per call, average output tokens per call, and call volume, but the engineering team also plans to introduce prompt caching for a shared instruction block. Which combination of usage fields should the cost model consume from the API responses to produce an accurate blended cost once caching is live?
Correct answer: A — input_tokens, cache_creation_input_tokens, cache_read_input_tokens, and output_tokens, applying each field's own per-token rate since they are billed at different prices
- A. Correct. With caching enabled, the usage object separately reports input_tokens (uncached, post-breakpoint), cache_creation_input_tokens (billed at the write multiplier), and cache_read_input_tokens (billed at roughly 0.1x base price), each with distinct pricing; a cost model must apply each field's rate separately, plus output_tokens at the output rate, to be accurate.
- B. Once caching is active, input_tokens alone excludes the cache_creation and cache_read fields, which carry different prices than standard input tokens; summing only input_tokens and output_tokens would misstate the true blended cost.
- C. Cached reads are discounted, not free; cache_read_input_tokens is still billed, just at a reduced rate (roughly 10% of base input price), so omitting it understates cost.
- D. cache_creation_input_tokens and cache_read_input_tokens are billing-relevant usage fields with their own prices, not merely diagnostic; excluding them produces an inaccurate cost model once caching is live.
3.Cache pricing and choosing a TTL
| Model | Base input | 5m writes | 1h writes | Hits and refreshes | Hit multiplier |
|---|---|---|---|---|---|
| Claude Fable 5.1 | $10 / MTok | $12.50 / MTok | $20 / MTok | $0.25 / MTok | 0.025x |
| Claude Opus 5.5 | $4 / MTok | $5 / MTok | $8 / MTok | $0.20 / MTok | 0.05x |
| Claude Sonnet 5 | $2 / MTok | $2.50 / MTok | $4 / MTok | $0.20 / MTok | 0.1x |
| Claude Haiku 4.5 | $1 / MTok | $1.25 / MTok | $2 / MTok | $0.10 / MTok | 0.1x |
A 5-minute write costs 1.25x the base input price and a 1-hour write costs 2x, on every model. Reads are where models differ. Most use the standard 0.1x, Claude Opus 5.5 uses 0.05x and Claude Fable 5.1 uses 0.025x. Here is the cost model from real numbers. A 100,000-token prefix on Claude Sonnet 5 costs $0.20 as plain input, $0.25 to write to the 5-minute cache, and $0.02 on each hit after that. One reuse is enough to pay back the write premium.
The default lifetime is 5 minutes, and every hit refreshes it at no extra cost. The lifetime is counted from the start of the request that wrote or read the entry, not from the end of its response. If a response takes 4 minutes to stream, the next request has to start within about 1 minute of it finishing to reuse the prefix.
It depends on how long the gaps are. A cache miss on either duration bills the whole prefix at the write price, so the 1-hour TTL pays off when more than about 1 gap in 20 falls between 5 minutes and an hour and gaps longer than an hour are rare. A gap longer than an hour expires both durations, and the 1-hour setting then pays its higher write price to rebuild the prefix. If about 60% or more of your pauses over 5 minutes also run past an hour, stay on the default. If turns arrive seconds apart, the 5-minute default cost about 15% less in Anthropic's measurements. On Claude Fable 5.1, reads are so cheap that keeping the 5-minute cache warm with keep-alive requests beat the 1-hour TTL whenever pauses lasted minutes.
A latency-sensitive voice assistant relies on a large cached system prompt to keep response times low. Product wants to guarantee the cache is always warm the instant a user starts speaking, even after periods with no traffic, without waiting for an actual user request to trigger the cache write. Which technique addresses this directly?
Correct answer: A — Issue a periodic request with max_tokens: 0 that includes the cached prefix and its cache_control breakpoint, refreshing the cache before it expires so it is warm when real traffic arrives
- A. Correct. Pre-warming with a max_tokens: 0 request that includes the cached prefix and its breakpoint writes or refreshes the cache entry ahead of real traffic, which is the documented technique for keeping latency-sensitive, large cached prompts ready without waiting on an actual user request.
- B. temperature controls output randomness/sampling and has no relationship to cache retention duration.
- C. Extended thinking affects reasoning behavior and token accounting for thinking blocks; it does not create a persistent, non-expiring cache, and thinking blocks themselves cannot be marked with cache_control.
- D. Streaming affects how response content is delivered incrementally; it does not extend or otherwise modify cache TTL behavior.
4.Keep the cache hitting
A cache only saves money while the prefix stays byte-for-byte stable. The prefix is ordered tools, then system, then messages. A change at one level invalidates that level and everything after it, so the higher up the change, the more you lose.
| Change | Invalidates |
|---|---|
| Modifying tool definitions | Entire cache (tools, system, messages) |
| Toggling web search or citations | System and messages caches |
| Changing tool_choice | Messages cache |
| Toggling images present/absent | Messages cache |
| Changing thinking parameters | Messages cache always; tool and system caches too on models that render the thinking configuration ahead of them |
Use the usage data you already log to check the cache's health. Over a full day of real traffic, agent loops read a median 84% of their input from cache, and the top tenth read 94% or more. If your share is below about 80%, something is breaking the cache. Expect one surprise in the usage data too. When a cached request uses a server tool such as web search, the API adds its own breakpoint on the tool result. That breakpoint always uses the 5-minute TTL, so 5-minute writes show up even if every marker you set is 1-hour.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Whenever users pause between turns, the 1-hour TTL is the cheaper choice.Why is that wrong?
Gaps over an hour expire both TTLs, and each of those misses then rebuilds the prefix at the 1-hour setting's higher write price. When long gaps are common, the default is cheaper.
Covered in Cache pricing and choosing a TTL
2.The 5-minute cache window starts when the previous response finishes streaming.Why is that wrong?
The window starts when the request that wrote or read the entry begins, so a long generation uses up most of it.
Covered in Cache pricing and choosing a TTL
3.Editing one tool definition only invalidates the cached tools; the cached system prompt and conversation are unaffected.Why is that wrong?
Tools come first in the prefix, so changing them invalidates the whole cache: tools, system and messages.
Covered in Keep the cache hitting
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligenceOfficial docs
“A 40-turn task sends its first turn 40 times, so task cost grows with roughly the square of turn count.”
↩︎ Why caching is the biggest lever“prompt caching was the largest lever by a wide margin”
↩︎ Why caching is the biggest lever“A miss on either duration bills the whole prefix at the write price instead of the read price”
↩︎ Cache pricing and choosing a TTL“agent loops read a median 84% of their input from the cache”
↩︎ Keep the cache hitting“Below about 80%, look for something breaking the cache”
↩︎ Keep the cache hitting“A gap over an hour expires both durations, and the 1-hour setting then re-writes the prefix at its higher write price”
↩︎ Exam trap 1 - 2.
“Otherwise, it processes the full prompt and caches the prefix once the response begins.”
↩︎ Why caching is the biggest lever“The system automatically applies the cache breakpoint to the last cacheable block.”
↩︎ Automatic caching and explicit breakpoints“If 4 explicit block-level breakpoints already exist, the API returns a 400 error (no slots left for automatic caching).”
↩︎ Automatic caching and explicit breakpoints“Pricing, minimum token thresholds, context ordering requirements, and the 20-block lookback window all apply the same as with explicit breakpoints.”
↩︎ Automatic caching and explicit breakpoints“Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.”
↩︎ Cache pricing and choosing a TTL“By default, the cache has a 5-minute lifetime. The cache is refreshed for no additional cost each time the cached content is used.”
↩︎ Cache pricing and choosing a TTL“The lifetime is measured from the start of the request that writes or reads the cache entry, not from the end of its response.”
↩︎ Exam trap 2 - 3.https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-use-with-prompt-cachingOfficial docs
“This caches the entire tool-definitions prefix, from the first tool through the marked breakpoint”
↩︎ Automatic caching and explicit breakpoints“This automatic breakpoint always uses the default 5-minute TTL, independent of any TTL you set on your own cache_control markers.”
↩︎ Keep the cache hitting“a change at one level invalidates that level and everything after it”
↩︎ Exam trap 3