CertSafari
    CCAR-P · Lessons

    Domain 2 · Lesson 11/38

    Prompt Caching on Claude: Breakpoints, TTL and Cache Invalidation

    Implement prompt reuse strategies

    9 min read
    2.6% of exam
    2 sources
    Published 27 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Explain how prompt caching reuses a prompt prefix up to a cache breakpoint
    • Choose between automatic caching and explicit cache breakpoints, and combine them within the four-slot limit
    • Reason about cache lifetime, the 1-hour TTL option and cache-hit pricing
    • Predict which request changes invalidate the tools, system or messages cache

    Key concept

    Prefix caching — Prompt caching stores the processed prefix of a request, up to a cache breakpoint. Later requests that begin with exactly the same prefix read it from cache instead of processing it again. Reuse therefore depends on keeping the stable parts of a prompt at the front and unchanged.

    1.What gets reused, and when

    The cheapest prompt to reuse is one you never process twice. With prompt caching enabled, the API checks whether the start of your request, up to a cache breakpoint, is already cached from a recent request. If it is, that cached version is used, which cuts processing time and cost. If it isn't, the API processes the whole prompt and caches the prefix once the response begins.

    The documentation lists four situations where this pays off: prompts with many examples, large amounts of context or background information, repetitive tasks with consistent instructions, and long multi-turn conversations. All four share one property: a large block of content that stays the same from one request to the next. Caching is how you turn a reusable prompt template into a reusable *processed* prompt.

    Sources1

    2.Automatic caching: one field, a moving breakpoint

    You can enable caching in two ways. The simpler one is automatic caching: add a single cache_control field at the top level of the request. The system places the breakpoint on the last cacheable block and moves it forward as the conversation grows. The documentation recommends this for multi-turn conversations.

    Automatic caching: cache_control at the top level of the request, not on a content blockpython
    client = anthropic.Anthropic()
    
    response = client.messages.create(
        model="claude-opus-5-5",
        max_tokens=1024,
        cache_control={"type": "ephemeral"},
        system="You are an AI assistant tasked with analyzing literary works. Your goal is to provide insightful commentary on themes, characters, and writing style.",
        messages=[
            {
                "role": "user",
                "content": "Analyze the major themes in 'Pride and Prejudice'.",
            }
        ],
    )
    print(response.usage.model_dump_json())
    How the automatic breakpoint moves forward over a conversation
    RequestRead from cacheWritten to cache
    Request 1NothingSystem + User(1) + Asst(1) + User(2)
    Request 2System through User(2)Asst(2) + User(3)
    Request 3System through User(3)Asst(3) + User(4)

    Automatic caching isn't a separate cache. It uses the same infrastructure as explicit breakpoints, so the same rules apply to it: pricing, minimum token thresholds, context ordering, and a 20-block lookback window. These sources name the lookback window but don't explain how it works. What they do give is the remedy for long sessions: combine the automatic breakpoint with explicit ones on content that never changes, which the next section covers.

    A long agent session accumulates conversation turns using only automatic request-level caching. By turn 15, a cache entry exists at content block 12. By turn 25, the conversation has grown to 35 content blocks, and the current breakpoint sits on the last block. Cache reads stop matching even though the early conversation content hasn't changed. What is the best explanation and fix?

    Sources1

    3.Explicit breakpoints, and combining the two

    Explicit breakpoints put cache_control directly on individual content blocks, so you decide exactly what gets cached. This is how you cache a prompt in modules. A common pattern is an explicit breakpoint on the system prompt, which is stable and shared by every request, while the top-level automatic breakpoint follows the growing conversation.

    Explicit breakpoint on the system prompt, with automatic caching for the conversationjson
    {
      "model": "claude-opus-5-5",
      "max_tokens": 1024,
      "cache_control": { "type": "ephemeral" },
      "system": [
        {
          "type": "text",
          "text": "You are a helpful assistant.",
          "cache_control": { "type": "ephemeral" }
        }
      ],
      "messages": [{ "role": "user", "content": "What are the key terms?" }]
    }

    A request can have at most four breakpoints, and the automatic breakpoint takes one of those slots. That produces a few edge cases worth knowing:

    Combining automatic caching with explicit breakpoints: outcomes
    SituationResult
    Last block already has explicit cache_control with the same TTLAutomatic caching is a no-op
    Last block has explicit cache_control with a different TTL400 error
    4 explicit block-level breakpoints already exist400 error: no slot left for automatic caching
    Last block is not an eligible breakpoint targetWalks backward to the nearest eligible block; if none, caching is skipped

    Tool definitions can be cached too. To cache the whole tool-definitions prefix, put cache_control on the last tool in the tools array. For an mcp_toolset, where you don't control tool order, put it on the toolset entry itself.

    Sources12

    4.Lifetime and price

    The cache lasts 5 minutes by default. Each time the cached content is used, the lifetime is refreshed at no extra cost. The clock starts when the request that writes or reads the entry *starts*, not when its response ends.

    Prompt caching prices per million tokens (current models)
    ModelBase input5m writes1h writesHits and refreshes
    Claude Fable 5.1$10$12.50$20$0.25
    Claude Opus 5.5$4$5$8$0.20
    Claude Sonnet 5$2$2.50$4$0.20
    Claude Haiku 4.5$1$1.25$2$0.10

    Cache writes cost a little more than base input. Reads cost far less: 0.05x base input on Opus 5.5, 0.025x on Fable 5.1 and Mythos 5.1, and the standard 0.1x on every other model. You pay a small premium once so that every later hit is cheap.

    Sources1

    5.What breaks the cache

    Because caching is prefix-based, the order of a request matters: tools come first, then system, then messages. A change at one level invalidates that level and every level after it, but none before it.

    Changes and the cache levels they invalidate
    ChangeInvalidates
    Modifying tool definitionsEntire cache (tools, system, messages)
    Toggling web search or citationsSystem and messages caches
    Changing tool_choiceMessages cache
    Changing disable_parallel_tool_useMessages cache
    Toggling images present/absentMessages cache
    Changing thinking parametersMessages cache always; tool and system caches too on models that render the thinking configuration ahead of them

    An agentic workflow sends requests with a large tools array (15 tool definitions, cached), a cached system prompt, and growing conversation history. For one request, the developer adds tool_choice: {"type": "tool", "name": "search"} to force a specific tool call, while leaving the tools array, system prompt, and message content unchanged. What happens to prompt caching for this request?

    Adding tools on demand doesn't have to break the cache. Deferred tools aren't part of the system-prompt prefix. When the model discovers one through tool search, its definition is appended to the conversation as a tool_reference block and the cached prefix stays untouched.

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Automatic caching and explicit cache_control breakpoints are mutually exclusive, so you have to pick one.Why is that wrong?

      You can use both in the same request. The automatic breakpoint simply takes one of the four breakpoint slots.

      Covered in Explicit breakpoints, and combining the two

    2. 2.The 5-minute cache lifetime starts counting when the previous response finishes streaming.Why is that wrong?

      The lifetime starts when the writing or reading request starts, so time spent generating the response uses it up.

      Covered in Lifetime and price

    3. 3.Changing tool_choice changes the tool configuration, so the whole cache is invalidated.Why is that wrong?

      Only modifying the tool definitions themselves invalidates the entire cache. Changing tool_choice invalidates only the messages cache, so the cached tools and system prefix still hit.

      Covered in What breaks the cache

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “The system checks if a prompt prefix, up to a specified cache breakpoint, is already cached from a recent query.”
      ↩︎ What gets reused, and when
      “The cache breakpoint automatically moves to the last cacheable block in each request”
      ↩︎ Automatic caching: one field, a moving breakpoint
      “Pricing, minimum token thresholds, context ordering requirements, and the 20-block lookback window all apply the same as with explicit breakpoints.”
      ↩︎ Automatic caching: one field, a moving breakpoint
      “use an explicit breakpoint to cache your system prompt, while automatic caching handles the conversation”
      ↩︎ Explicit breakpoints, and combining the two
      “If the last block has an explicit cache_control with a different TTL, the API returns a 400 error.”
      ↩︎ Explicit breakpoints, and combining the two
      “By default, the cache has a 5-minute lifetime. The cache is refreshed for no additional cost each time the cached content is used.”
      ↩︎ Lifetime and price
      “You can specify a 1-hour TTL at 2x the base input token price”
      ↩︎ Lifetime and price
      “Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.”
      ↩︎ Lifetime and price
      “Prompt caching optimizes your API usage by allowing resuming from specific prefixes in your prompts.”
      ↩︎ Key concept
      “Automatic caching is compatible with explicit cache breakpoints. When used together, the automatic cache breakpoint uses one of the 4 available breakpoint slots.”
      ↩︎ Exam trap 1
      “The lifetime is measured from the start of the request that writes or reads the cache entry, not from the end of its response.”
      ↩︎ Exam trap 2
    2. 2.
      “Place cache_control: {"type": "ephemeral"} on the last tool in your tools array.”
      ↩︎ Explicit breakpoints, and combining the two
      “The cache follows a prefix hierarchy (tools → system → messages), so a change at one level invalidates that level and everything after it”
      ↩︎ What breaks the cache
      “Deferred tools are not included in the system-prompt prefix.”
      ↩︎ What breaks the cache
      “The cache follows a prefix hierarchy (tools → system → messages), so a change at one level invalidates that level and everything after it”
      ↩︎ Exam trap 3

    Continue to page 2 of 2

    Claude Agent Skills: Modular, Reusable Prompts with Progressive Disclosure