CertSafari
    CLAUDE-CERTIFIED-DEVELOPER-FOUNDATIONS-CCDV-F · Lessons

    Domain 2 · Lesson 6/25

    Claude Prompt Caching, Message Batches and Realtime vs Batch

    Claude API Mechanics

    11 min read
    5.52% of exam
    3 sources
    Published 29 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Place cache_control breakpoints and explain the rules that produce a 400 error or a skipped cache
    • Reason about cache lifetime and TTL cost when timing follow-up requests
    • Build, poll and read a Message Batch, and name the parameters a batch rejects
    • Choose between realtime Messages API calls and the Message Batches API for a workload

    1.Prompt caching: reusing a prefix

    Prompt caching lets a request pick up from a prompt prefix that was already processed, which cuts both time and cost when many requests share the same start: a long system prompt, many examples, large background documents, or a growing conversation. On each request the system checks whether the prefix up to a cache breakpoint is already cached. If it is, the cached version is used. If not, the full prompt is processed and the prefix is cached once the response begins.

    There are two ways to mark a breakpoint. Automatic caching puts one cache_control at the top level of the request. The breakpoint goes on the last cacheable block and moves forward as the conversation grows, so earlier turns are read from cache and new turns are written to it, and you never move markers yourself. Explicit breakpoints put cache_control on specific content blocks when you want exact control. The two can be combined:

    An explicit breakpoint on the system prompt combined with top-level automatic caching for the conversationjson
    {
      "model": "claude-opus-5-5",
      "max_tokens": 1024,
      "cache_control": { "type": "ephemeral" },
      "system": [
        {
          "type": "text",
          "text": "You are a helpful assistant.",
          "cache_control": { "type": "ephemeral" }
        }
      ],
      "messages": [{ "role": "user", "content": "What are the key terms?" }]
    }
    How automatic caching interacts with explicit breakpoints
    SituationResult
    Automatic plus explicit breakpointsThe automatic breakpoint takes one of the 4 breakpoint slots
    4 explicit block-level breakpoints already set400 error: no slot left for automatic caching
    Last block already has explicit cache_control with the same TTLAutomatic caching does nothing
    Last block has explicit cache_control with a different TTL400 error
    Last block not eligible as a breakpointWalks backward to the nearest eligible block; if none, caching is skipped

    Sources1

    2.Cache lifetime and what it costs

    By default a cache entry lives for 5 minutes, and each use refreshes it at no extra cost. The clock starts when the request that writes or reads the entry begins, not when its response ends, so generation time uses up the lifetime.

    Per-million-token prompt caching prices for the current headline models
    ModelBase input5m cache writes1h cache writesHits and refreshes
    Claude Opus 5.5$4 / MTok$5 / MTok$8 / MTok$0.20 / MTok
    Claude Sonnet 5$2 / MTok$2.50 / MTok$4 / MTok$0.20 / MTok
    Claude Haiku 4.5$1 / MTok$1.25 / MTok$2 / MTok$0.10 / MTok
    Claude Fable 5.1$10 / MTok$12.50 / MTok$20 / MTok$0.25 / MTok

    Writing to the cache costs more than a normal input token, and reading from it costs much less. Most models price hits at 0.1x base input. Claude Opus 5.5 uses 0.05x, and Claude Fable 5.1 and Claude Mythos 5.1 use 0.025x. The 1-hour TTL costs 2x base input to write. It pays off when your requests are too far apart for the 5-minute window.

    A developer configures `thinking={"type": "enabled", "budget_tokens": 32000}` alongside `max_tokens=16000` and the request fails validation. What is the correct fix?

    Sources1

    3.Message Batches API: asynchronous bulk requests

    The Message Batches API takes many Messages requests at once and processes them asynchronously, each on its own. Every entry has a custom_id (1–64 characters, only letters, digits, hyphens and underscores) and a params object with standard Messages API parameters. A batch is capped at 100,000 requests or 256 MB, whichever it hits first. Because requests are independent, one batch can mix vision, tool use (including server tools), system messages, multi-turn conversations, extended thinking and most beta features.

    A newly created batch: processing_status starts at in_progress and request_counts tracks each outcomejson
    {
      "id": "msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d",
      "type": "message_batch",
      "processing_status": "in_progress",
      "request_counts": {
        "processing": 2,
        "succeeded": 0,
        "errored": 0,
        "canceled": 0,
        "expired": 0
      },
      "ended_at": null,
      "created_at": "2024-09-24T18:37:24.100435Z",
      "expires_at": "2024-09-25T18:37:24.100435Z",
      "cancel_initiated_at": null,
      "results_url": null
    }

    Poll the batch by id, or watch it in the Console, until processing_status changes to ended. Most batches finish within an hour. Results are available once every request has finished or after 24 hours, whichever comes first, and a batch that hasn't finished within 24 hours expires. Each request ends in one of four states:

    Batch result types and billing
    Result typeMeaningBilled?
    succeededRequest succeeded; includes the message resultYes
    erroredInvalid request or internal server error; no message createdNo
    canceledBatch canceled before this request was sent to the modelNo
    expired24-hour expiry reached before this request was sent to the modelNo

    Results have their own access rules. Download them from the batch's results_url, streaming rather than loading everything at once, and match each one back to your input by custom_id. Results can be downloaded for 29 days after creation. After that the batch is still visible but its results are not. Batches are scoped to a Workspace, so anyone whose requests run in that Workspace can see all its batches and results. A few parameters are rejected with a validation error: stream: true, because results come back as one file; speed (Fast mode), because it tunes synchronous latency; and max_tokens: 0, because a cache entry pre-warmed during batch processing would probably expire before the follow-up request runs.

    A document-processing application needs to send up to 550 images in a single Messages API request. Which model configuration choice is required to support that volume?

    Sources2

    4.Realtime or batch: choosing the path

    Realtime Messages API calls compared with the Message Batches API
    DimensionMessages API (realtime)Message Batches API
    Response timingImmediate, optionally streamed over SSEMost batches within 1 hour; expire at 24 hours
    PriceStandard API prices50% of standard prices (Claude Opus 5.5: $2 / $10 per MTok in/out)
    stream: trueSupportedValidation error
    speed (Fast mode)Applies to synchronous latencyValidation error
    max_tokens: 0 cache pre-warmingAvailableValidation error
    Typical workloadsUser-facing chat, interactive agent loopsLarge-scale evaluations, content moderation, data analysis, bulk content generation

    The deciding question is whether anyone is waiting for the answer. Batch suits large volumes where immediate responses aren't needed and cost matters. You accept a turnaround of up to a day in exchange for half-price tokens and higher throughput. There are two operational caveats. Under heavy demand, processing can slow down and more requests can expire at 24 hours. And with high concurrency, a batch can slightly overshoot your Workspace's spend limit. Caching still stacks with batch pricing, but don't expect every request to hit the cache. Requests run independently and concurrently, and a 5-minute entry can expire in the middle of a batch.

    Sources2

    5.Integration surface, and what these sources leave out

    Everything above is the Messages API: direct model access where you write the loop, keep the state and choose realtime or batch. Anthropic also offers Claude Managed Agents, a pre-built agent harness that runs on managed infrastructure. The Messages API suits custom agent loops and fine-grained control. Managed Agents suit long-running tasks and asynchronous work.

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.The 5-minute cache lifetime starts counting when the previous response finishes.Why is that wrong?

      It counts from the start of the request that wrote or read the entry. A long-running response uses up most of the window before your next call.

      Covered in Cache lifetime and what it costs

    2. 2.You can pre-warm the cache inside a batch with max_tokens: 0 so later batch requests hit it.Why is that wrong?

      Batch requests need max_tokens of at least 1. max_tokens: 0 is rejected because the cache entry would probably expire before the follow-up runs.

      Covered in Message Batches API: asynchronous bulk requests

    3. 3.Batch requests that errored or expired are billed like the ones that succeeded.Why is that wrong?

      Only succeeded requests are charged. Errored, canceled and expired requests are not billed.

      Covered in Message Batches API: asynchronous bulk requests

    Practise it for real

    Submit a two-request Message Batch, poll it until it ends, and see which parameter a batch rejects.

    1. 1.Call client.messages.batches.create with two Request entries, each with its own custom_id (for example my-first-request and my-second-request) and params containing model, max_tokens and messages.

      Why: custom_id is how you match each result to its input, because requests are processed independently.

      You should see: A message_batch object with processing_status "in_progress" and request_counts.processing equal to 2.

    2. 2.Loop on client.messages.batches.retrieve(batch_id), sleeping 60 seconds between calls, until processing_status == "ended".

      Why: Batches are asynchronous. Results only exist once every request has finished or 24 hours have passed.

      You should see: The status changes to "ended" and results_url is populated.

    3. 3.Read request_counts and stream the results from results_url, printing each custom_id with its result type.

      Why: Streaming the results file is the recommended way to read large results, and only succeeded requests are billed.

      You should see: Two results, most likely both succeeded, whose counts add up to 2.

    4. 4.Create another batch in which one request's params include "stream": true.

      Why: Batch results come back as a single file, so streaming isn't a valid batch parameter.

      You should see: A validation error, and no batch is created.

    Stuck? Get a nudge

    Match results by custom_id and not by position. Requests are processed independently, so don't rely on the order they come back in.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Otherwise, it processes the full prompt and caches the prefix once the response begins.”
      ↩︎ Prompt caching: reusing a prefix
      “When used together, the automatic cache breakpoint uses one of the 4 available breakpoint slots.”
      ↩︎ Prompt caching: reusing a prefix
      “Pricing, minimum token thresholds, context ordering requirements, and the 20-block lookback window all apply the same as with explicit breakpoints.”
      ↩︎ Prompt caching: reusing a prefix
      “The cache is refreshed for no additional cost each time the cached content is used.”
      ↩︎ Cache lifetime and what it costs
      “You can specify a 1-hour TTL at 2x the base input token price”
      ↩︎ Cache lifetime and what it costs
      “Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.”
      ↩︎ Cache lifetime and what it costs
      “The lifetime is measured from the start of the request that writes or reads the cache entry, not from the end of its response.”
      ↩︎ Exam trap 1
    2. 2.
      “Because each request in the batch is processed independently, you can mix different types of requests within a single batch.”
      ↩︎ Message Batches API: asynchronous bulk requests
      “Batch results are available for 29 days after creation.”
      ↩︎ Message Batches API: asynchronous bulk requests
      “Batches are scoped to a Workspace.”
      ↩︎ Message Batches API: asynchronous bulk requests
      “it's recommended to stream results back rather than download them all at once.”
      ↩︎ Message Batches API: asynchronous bulk requests
      “All usage is charged at 50% of the standard API prices.”
      ↩︎ Realtime or batch: choosing the path
      “Batch results come back as a single file, not a stream.”
      ↩︎ Realtime or batch: choosing the path
      “processing may be slowed down based on current demand and your request volume.”
      ↩︎ Realtime or batch: choosing the path
      “max_tokens: 0 (cache pre-warming) is not supported inside a batch”
      ↩︎ Exam trap 2
      “You will not be billed for these requests.”
      ↩︎ Exam trap 3
    3. 3.
      “Anthropic offers two ways to build with Claude, each suited to different use cases:”
      ↩︎ Integration surface, and what these sources leave out

    Ready to test yourself?

    Practise the 20 questions on this subdomain.