What you will be able to do
- Explain the trade the Message Batches API makes: a 50% discount in exchange for results that can take up to 24 hours, with no latency guarantee
- Pick the synchronous Messages API or the Message Batches API based on whether anyone is waiting on the result
- Size a batch workload against the per-batch limits and spot requests a batch will reject, including tool loops that need your code to run mid-request
- Use custom_id to match each batch result back to the request that produced it, and work out how often to submit batches so a delivery SLA still holds when a batch takes the full 24 hours
- Read a batch's request_counts, resubmit only the requests that errored or expired with the fix each one needs, and test a prompt on a small sample before committing a large volume to a batch
Key concept
Latency-for-cost trade — The Message Batches API charges half the normal price because it does not promise when results will arrive. You get results within 24 hours at most. That makes it right for work nobody is waiting on and wrong for anything that blocks a person or a pipeline.
1.The trade: half price, no promise on timing
The Message Batches API takes a list of ordinary Messages requests and processes them asynchronously. Each entry is still a normal request with a model, max_tokens and messages. The difference is that you don't get an answer straight away. The system works through the batch, and you poll its status until processing has ended. In return, all batch usage is billed at 50% of standard API prices, and the discount covers every token, cached tokens included.
Each entry also carries a custom_id field that you choose, and custom_id fields are how you correlate batch request/response pairs. The custom_id must be unique within the batch, 1 to 64 characters long, and made only of letters, digits, hyphens and underscores. Correlating by custom_id matters because results come back as a .jsonl file in which the lines are not guaranteed to be in the same order as the requests you sent. Every response line carries the custom_id of the request that produced it, and that field is the only reliable way to pair a response with its request. A practical habit is to make the custom_id carry the identity of the underlying work item, such as a document id or a test-case name, so that each request/response pair can be routed straight back to the record it belongs to without a separate lookup table.
Exam questions often turn on the timing. Anthropic says most batches finish in under an hour, but the only firm limit is 24 hours. Results become available when every request has completed or after 24 hours, whichever comes first. Requests that haven't been processed by then expire. The docs also warn that when demand is high, processing can slow down and more requests may expire at the 24-hour mark. So "usually done within the hour" describes typical behaviour. It is not a service level. If you design around the one-hour case, you are relying on something nobody promised.
Calculating batch submission frequency based on SLA constraints starts from that 24-hour bound. A batch begins processing as soon as it is created and can take up to 24 hours to complete, so the worst case for any single document is the time it waits for the next submission plus 24 hours of batch processing. Suppose documents arrive continuously and each must be analysed within a 30-hour SLA. Submitting one batch a day fails: a document that arrives just after a submission waits almost 24 hours, then processing can take another 24, for up to 48 hours. Submitting in 4-hour windows works: the longest wait is 4 hours, plus 24 hours of batch processing, for 28 hours in the worst case, which guarantees the 30-hour SLA with 2 hours of margin. The general rule for the submission frequency is that the submission interval plus 24 hours must fit inside the SLA, and you should leave some slack rather than sit exactly on the limit.
No. The documented limit is 24 hours, not one hour. Processing can slow down when demand is high, and requests that miss the 24-hour window expire instead of completing. A three-hour deadline needs the synchronous API.
The interval plus 24 hours must fit inside 36 hours, so the longest interval is 12 hours. Sitting exactly on the limit leaves no room for a slow batch or a resubmission, so a shorter interval such as 8 hours is the safer design.
2.Choose the API by asking who is waiting
Anthropic's cost guidance boils the choice down to one rule: send every request that nobody is waiting on through a batch, and keep everything else on the interactive path. The useful question for any workload is whether a person or a process is blocked until the answer comes back. A pre-merge check is blocking because the developer can't merge until it returns, so it belongs on the synchronous Messages API. Overnight reports, weekly audits and nightly test generation are the opposite: nobody reads the output until the next day, so they can go through the Message Batches API. The docs give the same kinds of examples for batching: evaluation runs, backfills, scheduled jobs, content moderation and bulk content generation.
| Workload | Who is waiting | API to use |
|---|---|---|
| Pre-merge check on a pull request | A developer, blocked until it returns | Synchronous Messages API |
| Nightly test generation | No one; reviewed the next day | Message Batches API |
| Weekly audit or overnight report | No one; read later | Message Batches API |
| Evaluation runs, backfills, scheduled jobs | No one | Message Batches API |
| Claude Managed Agents session | A user, interactively | Interactive path; batching is not available for these sessions |
A platform team runs an automated code-review gate that blocks a pull request from merging until Claude returns a verdict on the diff, typically within a few seconds. Which approach should they use for this workflow?
Correct answer: A — The synchronous Messages API, since the merge gate blocks on an immediate response and the Message Batches API offers no guaranteed latency SLA
- A. Correct. Pre-merge checks block on a fast response, and the Batches API has no guaranteed turnaround time (it can take up to 24 hours), so the synchronous API is required.
- B. Incorrect. The Batches API has no SLA guaranteeing a few-second turnaround; individual requests can sit for hours during periods of high demand.
- C. Incorrect. A blocking merge gate cannot tolerate an unbounded wait; the cost savings do not offset stalling every pull request pipeline.
- D. Incorrect. Batch requests carry no latency guarantee at all, so treating typical fast completion as reliable for a blocking gate is unsafe.
Sources4
3.What a single batch can hold
A batch is capped at 100,000 Message requests or 256 MB, whichever limit you hit first. A nightly job over 150,000 documents therefore needs at least two batches. It may need more if the payloads are large, such as scanned documents, because the 256 MB limit can bite before the request count does. Rate limits also apply, both to Batches API HTTP requests and to the number of batched requests waiting to be processed. Batches belong to a Workspace, and their results can be downloaded for 29 days after creation.
Almost any Messages request can go in a batch: vision, system messages, multi-turn conversations, extended thinking, most beta features, and tool use including server tools. Each request is processed independently, so a single batch can mix these types. A few parameters are rejected with a validation error:
| Parameter | Why it is rejected in a batch |
|---|---|
| stream: true | Batch results come back as a single file, not a stream |
| speed (Fast mode) | Fast mode tunes synchronous latency, which doesn't apply to asynchronous processing |
| max_tokens: 0 | A cache pre-warming entry written during batch processing would likely expire before the follow-up request runs |
A compliance group reviews 15,000 vendor contracts once a week and publishes a findings report two business days later. Cost per document matters because the review runs across the entire vendor catalog every cycle. How should this recurring job be built?
Correct answer: A — Submit the contracts as a single Message Batch each week, since the two-day turnaround comfortably absorbs the batch processing window and the discount lowers per-cycle spend
- A. Correct. A weekly, non-blocking audit with a multi-day turnaround tolerates the batch processing window comfortably while capturing the 50% cost reduction across the full catalog.
- B. Incorrect. Parallel synchronous calls bypass the batch discount entirely and add operational complexity without a latency requirement that demands it.
- C. Incorrect. A single batch can hold up to 100,000 requests, so fragmenting 15,000 contracts into many small submissions adds overhead for no benefit.
- D. Incorrect. Result ordering is not guaranteed by either API in a way that requires sequential calls; documents are correlated by custom_id, not by submission order.
Sources1
4.Tool use in a batch: one request, one answer
The supported-features list includes both "tool use" and "multi-turn conversations", and it is easy to read more into that than it says. In the Messages API, multi-turn means you pass the earlier turns in the messages array yourself. The model then writes one next turn. A batch can't pause a request halfway, run a tool in your code, and feed the tool result back into that same request.
These sources never state that limit in so many words. It follows from what they do say: each request is handled independently, results only become available once processing has ended, and they arrive as a single file with no streaming. If Claude's answer ends by asking for one of your tools, that request is already finished. To continue the loop, you build a new request with the tool result appended. In a batch, that means a whole new batch cycle of up to 24 hours per tool round. The server tools the docs list as supported (web search, web fetch, code execution and others) don't need your code between turns. Any agent loop that relies on your own tools does.
5.Failures: resubmit only what failed, and test the prompt first
Handling batch failures starts from the fact that a batch does not succeed or fail as a whole. Once processing has ended, every request has its own result, and request_counts on the batch tells you how many landed in each state. Only errored, canceled and expired requests need attention, and you are not billed for any of them. Because every result line is identified by custom_id, you can read the results file, collect the custom_ids whose result type is not succeeded, and build a new batch containing only those failed documents. Resubmitting the whole batch would pay a second time for every request that already succeeded.
| Result type | What it means | What to do |
|---|---|---|
| succeeded | The request produced a message | Keep the result; do not resubmit |
| errored | Invalid request or internal server error; not billed | Fix what caused the error, then resubmit that request by custom_id |
| canceled | You canceled the batch before this request reached the model; not billed | Resubmit unchanged if the work is still wanted |
| expired | The 24-hour window passed before this request reached the model; not billed | Resubmit in a new batch; consider smaller batches if demand is high |
Resubmitting only failed documents is half the job; the other half is making appropriate modifications, and the kind of failure decides what those are. An expired request never reached the model, so it can go back in unchanged. An errored request was rejected as invalid or hit a server error, and an invalid request resubmitted unchanged is still invalid. Look at the error on the result line and fix the cause first. For documents that exceeded context limits, the appropriate modification is chunking: split each oversized document into chunks that each become their own request, with custom_ids such as doc-42-part-1 and doc-42-part-2 so the pieces can be reassembled. The custom_id is what makes this selective resubmission possible, because it identifies which document each failed line belongs to.
The same 24-hour bound argues for prompt refinement on a sample set before batch-processing large volumes. Results are only available once processing has ended for all requests, so if the prompt is wrong you find out after a full cycle, and each round of iterative resubmission costs another cycle plus the tokens of every request that has to run again. The cheaper path is to run a small, representative sample set through the synchronous API first, check the outputs against the criteria you care about, and refine the prompt until the sample passes. Anthropic's guidance is to audit prompts against the model you actually run, and again whenever you change models; the sample run is where that audit happens. Only then submit the full volume. Refining on the sample maximizes the first-pass success rate of the large batch, which is what reduces iterative resubmission costs: most documents succeed the first time, and only a small remainder needs a second batch.
Only the 1,500 requests that did not succeed, found by reading the custom_id on each non-succeeded result line. The 300 expired requests can be resubmitted as they were. The 1,200 errored requests need their errors inspected first, and oversized documents split into chunks, because an invalid request fails again if resubmitted unchanged. None of the 1,500 was billed.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Most batches finish within an hour, so batching is fine for a check that has to finish before a merge.Why is that wrong?
The one-hour figure describes typical behaviour. The only firm bound is 24 hours, and requests still unprocessed at that point expire. A blocking check needs the synchronous API.
Covered in The trade: half price, no promise on timing
2.Results come back in the order the requests were submitted, so you can match them by position.Why is that wrong?
The results file is explicitly not guaranteed to be in request order. Matching by position silently attaches outputs to the wrong inputs. The custom_id on each result line is the only reliable key.
Covered in The trade: half price, no promise on timing
3.A 30-hour SLA is met by submitting one batch a day, since most batches finish within an hour.Why is that wrong?
Size the schedule on the worst case: the wait for the next submission plus up to 24 hours of processing. Daily submission can reach 48 hours; 4-hour windows cap it at 28.
Covered in The trade: half price, no promise on timing
4.The batch docs list tool use as supported, so a batched request can run your tools mid-request and carry on with their results.Why is that wrong?
A batched request produces one result, delivered in a file after processing ends. To continue a tool loop you submit a new request with the tool result appended, which means another batch cycle.
Covered in Tool use in a batch: one request, one answer
5.One nightly batch can hold any volume, as long as processing finishes within 24 hours.Why is that wrong?
Each batch has a hard cap on request count and on total size. Workloads above either cap have to be split across several batches.
Covered in What a single batch can hold
6.If some requests in a batch errored, the simplest recovery is to resubmit the whole batch.Why is that wrong?
Errored and expired requests are not billed, but the succeeded ones were, and running them again pays for them twice. Use request_counts to see what failed and custom_id to resubmit only those requests, fixed where the error demands it.
Covered in Failures: resubmit only what failed, and test the prompt first
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“with most batches finishing in less than 1 hour while reducing costs by 50% and increasing throughput”
↩︎ The trade: half price, no promise on timing“Batches expire if processing does not complete within 24 hours.”
↩︎ The trade: half price, no promise on timing“you may see more requests expiring after 24 hours”
↩︎ The trade: half price, no promise on timing“A unique custom_id for identifying the Messages request. Must be 1 to 64 characters and contain only alphanumeric characters, hyphens, and underscores”
↩︎ The trade: half price, no promise on timing“A Message Batch is limited to either 100,000 Message requests or 256 MB in size, whichever is reached first.”
↩︎ What a single batch can hold“Rate limits apply to both Batches API HTTP requests and the number of requests within a batch waiting to be processed.”
↩︎ What a single batch can hold“Because each request in the batch is processed independently, you can mix different types of requests within a single batch.”
↩︎ What a single batch can hold“The batch is then processed asynchronously, with each request handled independently.”
↩︎ Tool use in a batch: one request, one answer“Batch results come back as a single file, not a stream.”
↩︎ Tool use in a batch: one request, one answer“The batch's request_counts shows an overview of your results, indicating how many requests reached each of these four states.”
↩︎ Failures: resubmit only what failed, and test the prompt first“Possible errors include invalid requests and internal server errors. You will not be billed for these requests.”
↩︎ Failures: resubmit only what failed, and test the prompt first“Batch reached its 24-hour expiration before this request could be sent to the model. You will not be billed for these requests.”
↩︎ Failures: resubmit only what failed, and test the prompt first“You can access batch results when all messages have completed or after 24 hours, whichever comes first.”
↩︎ Failures: resubmit only what failed, and test the prompt first“You can access batch results when all messages have completed or after 24 hours, whichever comes first.”
↩︎ Exam trap 1“Batch results come back as a single file, not a stream.”
↩︎ Exam trap 4“A Message Batch is limited to either 100,000 Message requests or 256 MB in size, whichever is reached first.”
↩︎ Exam trap 5“Possible errors include invalid requests and internal server errors. You will not be billed for these requests.”
↩︎ Exam trap 6 - 2.
“Results are not guaranteed to be in the same order as requests. Use the custom_id field to match results to requests.”
↩︎ The trade: half price, no promise on timing“Developer-provided ID created for each request in a Message Batch. Useful for matching results to requests, as results may be given out of request order.”
↩︎ Failures: resubmit only what failed, and test the prompt first“Results are not guaranteed to be in the same order as requests. Use the custom_id field to match results to requests.”
↩︎ Exam trap 2 - 3.
“Developer-provided ID created for each request in a Message Batch. Useful for matching results to requests, as results may be given out of request order.”
↩︎ The trade: half price, no promise on timing“Once a Message Batch is created, it begins processing immediately. Batches can take up to 24 hours to complete.”
↩︎ The trade: half price, no promise on timing“you specify the prior conversational turns with the messages parameter”
↩︎ Tool use in a batch: one request, one answer“Once a Message Batch is created, it begins processing immediately. Batches can take up to 24 hours to complete.”
↩︎ Exam trap 3 - 4.https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligenceOfficial docs
“Route every request no one is waiting on through a batch, and keep the interactive path for the rest.”
↩︎ Choose the API by asking who is waiting“Batching is the second-largest free lever after caching for unattended agent work: evaluation runs, backfills, and scheduled jobs”
↩︎ Choose the API by asking who is waiting“Auditing prompts against the model you run now, and again whenever you change models, is a free win.”
↩︎ Failures: resubmit only what failed, and test the prompt first“The Batch API takes 50% off every token of a request, including cached ones, in exchange for results arriving any time within 24 hours.”
↩︎ Key concept