CertSafari
    CLAUDE-CERTIFIED-DEVELOPER-FOUNDATIONS-CCDV-F · Lessons

    Domain 5 · Lesson 15/25

    Prompt Caching: Breakpoints, TTLs and Cache Pricing

    Cost and Token Management

    9 min read
    4.2% of exam
    3 sources
    Published 29 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Explain why prompt caching is the largest cost lever for multi-turn and agentic workloads
    • Choose between automatic caching and explicit cache breakpoints, and place breakpoints within the four-slot limit
    • Price cache writes and reads per model and choose between the 5-minute and 1-hour TTL from real pause patterns
    • Name the request changes that invalidate the cache, and use the cache-read share to check that caching is working

    1.Why caching is the biggest lever

    Claude has no memory between requests. Every call sends the whole prompt again. In a single-shot app that is harmless. In an agent it compounds: every turn resends the system prompt, the tool definitions and every earlier turn. The first turn of a 40-turn task is sent 40 times, so the cost of a task grows roughly with the square of the number of turns.

    Prompt caching doesn't reduce what you send. It reduces what you pay for the repeated part. When a request has caching enabled, the system checks whether the prompt prefix up to a cache breakpoint was cached by a recent request. If it was, the cached version is used. If not, the full prompt is processed and the prefix is cached once the response starts. The cached prefix is then billed at the cache-read rate, and each turn pays the higher write rate only for its new content. In Anthropic's measurements caching was the largest lever by a wide margin. It cut agent-loop cost by a factor of 2.7 to 5.3, and cut a small triage agent's bill by 83%.

    Sources12

    2.Automatic caching and explicit breakpoints

    There are two ways to turn caching on. Automatic caching is a single top-level cache_control field. The system puts the breakpoint on the last cacheable block and moves it forward as the conversation grows. On each new request, everything up to the previous turn is read from cache and only the newest assistant and user turns are written. You never edit a marker. Explicit breakpoints put cache_control on individual content blocks, which gives you exact control over what is cached.

    You can combine the two. A request has four breakpoint slots in total, and automatic caching uses one of them. If four explicit block-level breakpoints already exist, adding automatic caching returns a 400 error. So does an explicit marker on the last block with a different TTL. The same pricing, minimum token thresholds, ordering requirements and 20-block lookback window apply to both methods. These sources say that minimum thresholds exist but give no per-model numbers.

    An explicit breakpoint caches the system prompt while automatic caching handles the conversationjson
    {
      "model": "claude-opus-5-5",
      "max_tokens": 1024,
      "cache_control": { "type": "ephemeral" },
      "system": [
        {
          "type": "text",
          "text": "You are a helpful assistant.",
          "cache_control": { "type": "ephemeral" }
        }
      ],
      "messages": [{ "role": "user", "content": "What are the key terms?" }]
    }

    Tool definitions are a common thing to checkpoint. Put cache_control on the last tool in the tools array, and the whole tool-definitions prefix is cached, from the first tool up to that marker. With an mcp_toolset you don't control the order of the tools, so put the marker on the toolset entry and the API applies it to the last expanded tool. Tools loaded with defer_loading through tool search are added inline to the conversation rather than to the prefix, so discovering tools mid-session doesn't break the cache.

    A finance team wants a monthly cost model for an internal Claude-powered summarization service. They know the average input tokens per call, average output tokens per call, and call volume, but the engineering team also plans to introduce prompt caching for a shared instruction block. Which combination of usage fields should the cost model consume from the API responses to produce an accurate blended cost once caching is live?

    Sources23

    3.Cache pricing and choosing a TTL

    Prompt caching prices per million tokens, and the cache-hit price as a multiple of base input
    ModelBase input5m writes1h writesHits and refreshesHit multiplier
    Claude Fable 5.1$10 / MTok$12.50 / MTok$20 / MTok$0.25 / MTok0.025x
    Claude Opus 5.5$4 / MTok$5 / MTok$8 / MTok$0.20 / MTok0.05x
    Claude Sonnet 5$2 / MTok$2.50 / MTok$4 / MTok$0.20 / MTok0.1x
    Claude Haiku 4.5$1 / MTok$1.25 / MTok$2 / MTok$0.10 / MTok0.1x

    A 5-minute write costs 1.25x the base input price and a 1-hour write costs 2x, on every model. Reads are where models differ. Most use the standard 0.1x, Claude Opus 5.5 uses 0.05x and Claude Fable 5.1 uses 0.025x. Here is the cost model from real numbers. A 100,000-token prefix on Claude Sonnet 5 costs $0.20 as plain input, $0.25 to write to the 5-minute cache, and $0.02 on each hit after that. One reuse is enough to pay back the write premium.

    The default lifetime is 5 minutes, and every hit refreshes it at no extra cost. The lifetime is counted from the start of the request that wrote or read the entry, not from the end of its response. If a response takes 4 minutes to stream, the next request has to start within about 1 minute of it finishing to reuse the prefix.

    It depends on how long the gaps are. A cache miss on either duration bills the whole prefix at the write price, so the 1-hour TTL pays off when more than about 1 gap in 20 falls between 5 minutes and an hour and gaps longer than an hour are rare. A gap longer than an hour expires both durations, and the 1-hour setting then pays its higher write price to rebuild the prefix. If about 60% or more of your pauses over 5 minutes also run past an hour, stay on the default. If turns arrive seconds apart, the 5-minute default cost about 15% less in Anthropic's measurements. On Claude Fable 5.1, reads are so cheap that keeping the 5-minute cache warm with keep-alive requests beat the 1-hour TTL whenever pauses lasted minutes.

    A latency-sensitive voice assistant relies on a large cached system prompt to keep response times low. Product wants to guarantee the cache is always warm the instant a user starts speaking, even after periods with no traffic, without waiting for an actual user request to trigger the cache write. Which technique addresses this directly?

    Sources21

    4.Keep the cache hitting

    A cache only saves money while the prefix stays byte-for-byte stable. The prefix is ordered tools, then system, then messages. A change at one level invalidates that level and everything after it, so the higher up the change, the more you lose.

    Which request changes invalidate which parts of the cache
    ChangeInvalidates
    Modifying tool definitionsEntire cache (tools, system, messages)
    Toggling web search or citationsSystem and messages caches
    Changing tool_choiceMessages cache
    Toggling images present/absentMessages cache
    Changing thinking parametersMessages cache always; tool and system caches too on models that render the thinking configuration ahead of them

    Use the usage data you already log to check the cache's health. Over a full day of real traffic, agent loops read a median 84% of their input from cache, and the top tenth read 94% or more. If your share is below about 80%, something is breaking the cache. Expect one surprise in the usage data too. When a cached request uses a server tool such as web search, the API adds its own breakpoint on the tool result. That breakpoint always uses the 5-minute TTL, so 5-minute writes show up even if every marker you set is 1-hour.

    Sources13

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Whenever users pause between turns, the 1-hour TTL is the cheaper choice.Why is that wrong?

      Gaps over an hour expire both TTLs, and each of those misses then rebuilds the prefix at the 1-hour setting's higher write price. When long gaps are common, the default is cheaper.

      Covered in Cache pricing and choosing a TTL

    2. 2.The 5-minute cache window starts when the previous response finishes streaming.Why is that wrong?

      The window starts when the request that wrote or read the entry begins, so a long generation uses up most of it.

      Covered in Cache pricing and choosing a TTL

    3. 3.Editing one tool definition only invalidates the cached tools; the cached system prompt and conversation are unaffected.Why is that wrong?

      Tools come first in the prefix, so changing them invalidates the whole cache: tools, system and messages.

      Covered in Keep the cache hitting

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “A 40-turn task sends its first turn 40 times, so task cost grows with roughly the square of turn count.”
      ↩︎ Why caching is the biggest lever
      “prompt caching was the largest lever by a wide margin”
      ↩︎ Why caching is the biggest lever
      “A miss on either duration bills the whole prefix at the write price instead of the read price”
      ↩︎ Cache pricing and choosing a TTL
      “agent loops read a median 84% of their input from the cache”
      ↩︎ Keep the cache hitting
      “Below about 80%, look for something breaking the cache”
      ↩︎ Keep the cache hitting
      “A gap over an hour expires both durations, and the 1-hour setting then re-writes the prefix at its higher write price”
      ↩︎ Exam trap 1
    2. 2.
      “Otherwise, it processes the full prompt and caches the prefix once the response begins.”
      ↩︎ Why caching is the biggest lever
      “The system automatically applies the cache breakpoint to the last cacheable block.”
      ↩︎ Automatic caching and explicit breakpoints
      “If 4 explicit block-level breakpoints already exist, the API returns a 400 error (no slots left for automatic caching).”
      ↩︎ Automatic caching and explicit breakpoints
      “Pricing, minimum token thresholds, context ordering requirements, and the 20-block lookback window all apply the same as with explicit breakpoints.”
      ↩︎ Automatic caching and explicit breakpoints
      “Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.”
      ↩︎ Cache pricing and choosing a TTL
      “By default, the cache has a 5-minute lifetime. The cache is refreshed for no additional cost each time the cached content is used.”
      ↩︎ Cache pricing and choosing a TTL
      “The lifetime is measured from the start of the request that writes or reads the cache entry, not from the end of its response.”
      ↩︎ Exam trap 2
    3. 3.
      “This caches the entire tool-definitions prefix, from the first tool through the marked breakpoint”
      ↩︎ Automatic caching and explicit breakpoints
      “This automatic breakpoint always uses the default 5-minute TTL, independent of any TTL you set on your own cache_control markers.”
      ↩︎ Keep the cache hitting
      “a change at one level invalidates that level and everything after it”
      ↩︎ Exam trap 3

    Ready to test yourself?

    Practise the 20 questions on this subdomain.