What you will be able to do
- Place cache_control breakpoints and explain the rules that produce a 400 error or a skipped cache
- Reason about cache lifetime and TTL cost when timing follow-up requests
- Build, poll and read a Message Batch, and name the parameters a batch rejects
- Choose between realtime Messages API calls and the Message Batches API for a workload
1.Prompt caching: reusing a prefix
Prompt caching lets a request pick up from a prompt prefix that was already processed, which cuts both time and cost when many requests share the same start: a long system prompt, many examples, large background documents, or a growing conversation. On each request the system checks whether the prefix up to a cache breakpoint is already cached. If it is, the cached version is used. If not, the full prompt is processed and the prefix is cached once the response begins.
There are two ways to mark a breakpoint. Automatic caching puts one cache_control at the top level of the request. The breakpoint goes on the last cacheable block and moves forward as the conversation grows, so earlier turns are read from cache and new turns are written to it, and you never move markers yourself. Explicit breakpoints put cache_control on specific content blocks when you want exact control. The two can be combined:
{
"model": "claude-opus-5-5",
"max_tokens": 1024,
"cache_control": { "type": "ephemeral" },
"system": [
{
"type": "text",
"text": "You are a helpful assistant.",
"cache_control": { "type": "ephemeral" }
}
],
"messages": [{ "role": "user", "content": "What are the key terms?" }]
}| Situation | Result |
|---|---|
| Automatic plus explicit breakpoints | The automatic breakpoint takes one of the 4 breakpoint slots |
| 4 explicit block-level breakpoints already set | 400 error: no slot left for automatic caching |
| Last block already has explicit cache_control with the same TTL | Automatic caching does nothing |
| Last block has explicit cache_control with a different TTL | 400 error |
| Last block not eligible as a breakpoint | Walks backward to the nearest eligible block; if none, caching is skipped |
Sources1
2.Cache lifetime and what it costs
By default a cache entry lives for 5 minutes, and each use refreshes it at no extra cost. The clock starts when the request that writes or reads the entry begins, not when its response ends, so generation time uses up the lifetime.
About 1 minute. The 5-minute lifetime counts from the start of the 4-minute request, so only about 1 minute is left once the response completes. For slower loops, use the 1-hour TTL: { "cache_control": { "type": "ephemeral", "ttl": "1h" } }.
| Model | Base input | 5m cache writes | 1h cache writes | Hits and refreshes |
|---|---|---|---|---|
| Claude Opus 5.5 | $4 / MTok | $5 / MTok | $8 / MTok | $0.20 / MTok |
| Claude Sonnet 5 | $2 / MTok | $2.50 / MTok | $4 / MTok | $0.20 / MTok |
| Claude Haiku 4.5 | $1 / MTok | $1.25 / MTok | $2 / MTok | $0.10 / MTok |
| Claude Fable 5.1 | $10 / MTok | $12.50 / MTok | $20 / MTok | $0.25 / MTok |
Writing to the cache costs more than a normal input token, and reading from it costs much less. Most models price hits at 0.1x base input. Claude Opus 5.5 uses 0.05x, and Claude Fable 5.1 and Claude Mythos 5.1 use 0.025x. The 1-hour TTL costs 2x base input to write. It pays off when your requests are too far apart for the 5-minute window.
A developer configures `thinking={"type": "enabled", "budget_tokens": 32000}` alongside `max_tokens=16000` and the request fails validation. What is the correct fix?
Correct answer: A — Lower `budget_tokens` to a value strictly less than `max_tokens`, since the thinking budget must fit within the overall output token ceiling
- A. Correct. `budget_tokens` must be less than `max_tokens`; here 32000 exceeds 16000, so lowering the thinking budget below the max_tokens ceiling resolves the validation error.
- B. `max_tokens` is still a required parameter for standard (non-adaptive) requests; extended thinking doesn't remove the need to set an overall output ceiling.
- C. There's no negative-offset mechanism in the API; `budget_tokens` and `max_tokens` are both positive integers with a simple less-than constraint.
- D. The API does not enforce a fixed 2x ratio between `max_tokens` and `budget_tokens`; the only documented constraint is that the budget must be smaller than the max token ceiling.
Sources1
3.Message Batches API: asynchronous bulk requests
The Message Batches API takes many Messages requests at once and processes them asynchronously, each on its own. Every entry has a custom_id (1–64 characters, only letters, digits, hyphens and underscores) and a params object with standard Messages API parameters. A batch is capped at 100,000 requests or 256 MB, whichever it hits first. Because requests are independent, one batch can mix vision, tool use (including server tools), system messages, multi-turn conversations, extended thinking and most beta features.
{
"id": "msgbatch_01HkcTjaV5uDC8jWR4ZsDV8d",
"type": "message_batch",
"processing_status": "in_progress",
"request_counts": {
"processing": 2,
"succeeded": 0,
"errored": 0,
"canceled": 0,
"expired": 0
},
"ended_at": null,
"created_at": "2024-09-24T18:37:24.100435Z",
"expires_at": "2024-09-25T18:37:24.100435Z",
"cancel_initiated_at": null,
"results_url": null
}Poll the batch by id, or watch it in the Console, until processing_status changes to ended. Most batches finish within an hour. Results are available once every request has finished or after 24 hours, whichever comes first, and a batch that hasn't finished within 24 hours expires. Each request ends in one of four states:
| Result type | Meaning | Billed? |
|---|---|---|
| succeeded | Request succeeded; includes the message result | Yes |
| errored | Invalid request or internal server error; no message created | No |
| canceled | Batch canceled before this request was sent to the model | No |
| expired | 24-hour expiry reached before this request was sent to the model | No |
Results have their own access rules. Download them from the batch's results_url, streaming rather than loading everything at once, and match each one back to your input by custom_id. Results can be downloaded for 29 days after creation. After that the batch is still visible but its results are not. Batches are scoped to a Workspace, so anyone whose requests run in that Workspace can see all its batches and results. A few parameters are rejected with a validation error: stream: true, because results come back as one file; speed (Fast mode), because it tunes synchronous latency; and max_tokens: 0, because a cache entry pre-warmed during batch processing would probably expire before the follow-up request runs.
A document-processing application needs to send up to 550 images in a single Messages API request. Which model configuration choice is required to support that volume?
Correct answer: A — Use a model without a 200k-token context window, since only those models allow up to 600 images per API request, compared to a 100-image cap on 200k-context models
- A. Correct. The documented request limits are 100 images per request for models with a 200k-token context window, versus 600 images per request for other models, so reaching 550 images requires choosing a model outside the 200k-context tier.
- B. The cap is not uniform; it explicitly differs between 200k-context models (100) and other models (600), so this overstates the consistency across the lineup.
- C. 300 is not a documented cap; the actual limits are 100 or 600 depending on the model's context window, so this invented threshold doesn't reflect the real constraint.
- D. The Files API changes how image bytes are referenced (by `file_id` instead of resent base64 data) to reduce payload size, but it doesn't change the per-request image count limits, which are tied to context window size, not source type.
Sources2
4.Realtime or batch: choosing the path
| Dimension | Messages API (realtime) | Message Batches API |
|---|---|---|
| Response timing | Immediate, optionally streamed over SSE | Most batches within 1 hour; expire at 24 hours |
| Price | Standard API prices | 50% of standard prices (Claude Opus 5.5: $2 / $10 per MTok in/out) |
| stream: true | Supported | Validation error |
| speed (Fast mode) | Applies to synchronous latency | Validation error |
| max_tokens: 0 cache pre-warming | Available | Validation error |
| Typical workloads | User-facing chat, interactive agent loops | Large-scale evaluations, content moderation, data analysis, bulk content generation |
The deciding question is whether anyone is waiting for the answer. Batch suits large volumes where immediate responses aren't needed and cost matters. You accept a turnaround of up to a day in exchange for half-price tokens and higher throughput. There are two operational caveats. Under heavy demand, processing can slow down and more requests can expire at 24 hours. And with high concurrency, a batch can slightly overshoot your Workspace's spend limit. Caching still stacks with batch pricing, but don't expect every request to hit the cache. Requests run independently and concurrently, and a 5-minute entry can expire in the middle of a batch.
Sources2
5.Integration surface, and what these sources leave out
Everything above is the Messages API: direct model access where you write the loop, keep the state and choose realtime or batch. Anthropic also offers Claude Managed Agents, a pre-built agent harness that runs on managed infrastructure. The Messages API suits custom agent loops and fine-grained control. Managed Agents suit long-running tasks and asynchronous work.
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The 5-minute cache lifetime starts counting when the previous response finishes.Why is that wrong?
It counts from the start of the request that wrote or read the entry. A long-running response uses up most of the window before your next call.
Covered in Cache lifetime and what it costs
2.You can pre-warm the cache inside a batch with max_tokens: 0 so later batch requests hit it.Why is that wrong?
Batch requests need max_tokens of at least 1. max_tokens: 0 is rejected because the cache entry would probably expire before the follow-up runs.
3.Batch requests that errored or expired are billed like the ones that succeeded.Why is that wrong?
Only succeeded requests are charged. Errored, canceled and expired requests are not billed.
Practise it for real
Submit a two-request Message Batch, poll it until it ends, and see which parameter a batch rejects.
1.Call client.messages.batches.create with two Request entries, each with its own custom_id (for example my-first-request and my-second-request) and params containing model, max_tokens and messages.
Why: custom_id is how you match each result to its input, because requests are processed independently.
You should see: A message_batch object with processing_status "in_progress" and request_counts.processing equal to 2.
2.Loop on client.messages.batches.retrieve(batch_id), sleeping 60 seconds between calls, until processing_status == "ended".
Why: Batches are asynchronous. Results only exist once every request has finished or 24 hours have passed.
You should see: The status changes to "ended" and results_url is populated.
3.Read request_counts and stream the results from results_url, printing each custom_id with its result type.
Why: Streaming the results file is the recommended way to read large results, and only succeeded requests are billed.
You should see: Two results, most likely both succeeded, whose counts add up to 2.
4.Create another batch in which one request's params include "stream": true.
Why: Batch results come back as a single file, so streaming isn't a valid batch parameter.
You should see: A validation error, and no batch is created.
Stuck? Get a nudge
Match results by custom_id and not by position. Requests are processed independently, so don't rely on the order they come back in.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Otherwise, it processes the full prompt and caches the prefix once the response begins.”
↩︎ Prompt caching: reusing a prefix“When used together, the automatic cache breakpoint uses one of the 4 available breakpoint slots.”
↩︎ Prompt caching: reusing a prefix“Pricing, minimum token thresholds, context ordering requirements, and the 20-block lookback window all apply the same as with explicit breakpoints.”
↩︎ Prompt caching: reusing a prefix“The cache is refreshed for no additional cost each time the cached content is used.”
↩︎ Cache lifetime and what it costs“You can specify a 1-hour TTL at 2x the base input token price”
↩︎ Cache lifetime and what it costs“Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.”
↩︎ Cache lifetime and what it costs“The lifetime is measured from the start of the request that writes or reads the cache entry, not from the end of its response.”
↩︎ Exam trap 1 - 2.
“Because each request in the batch is processed independently, you can mix different types of requests within a single batch.”
↩︎ Message Batches API: asynchronous bulk requests“Batch results are available for 29 days after creation.”
↩︎ Message Batches API: asynchronous bulk requests“Batches are scoped to a Workspace.”
↩︎ Message Batches API: asynchronous bulk requests“it's recommended to stream results back rather than download them all at once.”
↩︎ Message Batches API: asynchronous bulk requests“All usage is charged at 50% of the standard API prices.”
↩︎ Realtime or batch: choosing the path“Batch results come back as a single file, not a stream.”
↩︎ Realtime or batch: choosing the path“processing may be slowed down based on current demand and your request volume.”
↩︎ Realtime or batch: choosing the path“max_tokens: 0 (cache pre-warming) is not supported inside a batch”
↩︎ Exam trap 2“You will not be billed for these requests.”
↩︎ Exam trap 3 - 3.
“Anthropic offers two ways to build with Claude, each suited to different use cases:”
↩︎ Integration surface, and what these sources leave out