CertSafari
    CLAUDE-CERTIFIED-ARCHITECT-FOUNDATIONS-CCAR-F · Lessons

    Domain 4 · Lesson 23/30

    Message Batches API: When to Batch and What It Costs

    Design efficient batch processing strategies

    14 min read
    3.33% of exam
    4 sources
    Published 29 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Explain the trade the Message Batches API makes: a 50% discount in exchange for results that can take up to 24 hours, with no latency guarantee
    • Pick the synchronous Messages API or the Message Batches API based on whether anyone is waiting on the result
    • Size a batch workload against the per-batch limits and spot requests a batch will reject, including tool loops that need your code to run mid-request
    • Use custom_id to match each batch result back to the request that produced it, and work out how often to submit batches so a delivery SLA still holds when a batch takes the full 24 hours
    • Read a batch's request_counts, resubmit only the requests that errored or expired with the fix each one needs, and test a prompt on a small sample before committing a large volume to a batch

    Key concept

    Latency-for-cost trade — The Message Batches API charges half the normal price because it does not promise when results will arrive. You get results within 24 hours at most. That makes it right for work nobody is waiting on and wrong for anything that blocks a person or a pipeline.

    1.The trade: half price, no promise on timing

    The Message Batches API takes a list of ordinary Messages requests and processes them asynchronously. Each entry is still a normal request with a model, max_tokens and messages. The difference is that you don't get an answer straight away. The system works through the batch, and you poll its status until processing has ended. In return, all batch usage is billed at 50% of standard API prices, and the discount covers every token, cached tokens included.

    Each entry also carries a custom_id field that you choose, and custom_id fields are how you correlate batch request/response pairs. The custom_id must be unique within the batch, 1 to 64 characters long, and made only of letters, digits, hyphens and underscores. Correlating by custom_id matters because results come back as a .jsonl file in which the lines are not guaranteed to be in the same order as the requests you sent. Every response line carries the custom_id of the request that produced it, and that field is the only reliable way to pair a response with its request. A practical habit is to make the custom_id carry the identity of the underlying work item, such as a document id or a test-case name, so that each request/response pair can be routed straight back to the record it belongs to without a separate lookup table.

    Exam questions often turn on the timing. Anthropic says most batches finish in under an hour, but the only firm limit is 24 hours. Results become available when every request has completed or after 24 hours, whichever comes first. Requests that haven't been processed by then expire. The docs also warn that when demand is high, processing can slow down and more requests may expire at the 24-hour mark. So "usually done within the hour" describes typical behaviour. It is not a service level. If you design around the one-hour case, you are relying on something nobody promised.

    Calculating batch submission frequency based on SLA constraints starts from that 24-hour bound. A batch begins processing as soon as it is created and can take up to 24 hours to complete, so the worst case for any single document is the time it waits for the next submission plus 24 hours of batch processing. Suppose documents arrive continuously and each must be analysed within a 30-hour SLA. Submitting one batch a day fails: a document that arrives just after a submission waits almost 24 hours, then processing can take another 24, for up to 48 hours. Submitting in 4-hour windows works: the longest wait is 4 hours, plus 24 hours of batch processing, for 28 hours in the worst case, which guarantees the 30-hour SLA with 2 hours of margin. The general rule for the submission frequency is that the submission interval plus 24 hours must fit inside the SLA, and you should leave some slack rather than sit exactly on the limit.

    Sources123

    2.Choose the API by asking who is waiting

    Anthropic's cost guidance boils the choice down to one rule: send every request that nobody is waiting on through a batch, and keep everything else on the interactive path. The useful question for any workload is whether a person or a process is blocked until the answer comes back. A pre-merge check is blocking because the developer can't merge until it returns, so it belongs on the synchronous Messages API. Overnight reports, weekly audits and nightly test generation are the opposite: nobody reads the output until the next day, so they can go through the Message Batches API. The docs give the same kinds of examples for batching: evaluation runs, backfills, scheduled jobs, content moderation and bulk content generation.

    Matching workloads to the API by who is waiting on the result
    WorkloadWho is waitingAPI to use
    Pre-merge check on a pull requestA developer, blocked until it returnsSynchronous Messages API
    Nightly test generationNo one; reviewed the next dayMessage Batches API
    Weekly audit or overnight reportNo one; read laterMessage Batches API
    Evaluation runs, backfills, scheduled jobsNo oneMessage Batches API
    Claude Managed Agents sessionA user, interactivelyInteractive path; batching is not available for these sessions

    A platform team runs an automated code-review gate that blocks a pull request from merging until Claude returns a verdict on the diff, typically within a few seconds. Which approach should they use for this workflow?

    Sources4

    3.What a single batch can hold

    A batch is capped at 100,000 Message requests or 256 MB, whichever limit you hit first. A nightly job over 150,000 documents therefore needs at least two batches. It may need more if the payloads are large, such as scanned documents, because the 256 MB limit can bite before the request count does. Rate limits also apply, both to Batches API HTTP requests and to the number of batched requests waiting to be processed. Batches belong to a Workspace, and their results can be downloaded for 29 days after creation.

    Almost any Messages request can go in a batch: vision, system messages, multi-turn conversations, extended thinking, most beta features, and tool use including server tools. Each request is processed independently, so a single batch can mix these types. A few parameters are rejected with a validation error:

    Messages API parameters the Message Batches API rejects, and why
    ParameterWhy it is rejected in a batch
    stream: trueBatch results come back as a single file, not a stream
    speed (Fast mode)Fast mode tunes synchronous latency, which doesn't apply to asynchronous processing
    max_tokens: 0A cache pre-warming entry written during batch processing would likely expire before the follow-up request runs

    A compliance group reviews 15,000 vendor contracts once a week and publishes a findings report two business days later. Cost per document matters because the review runs across the entire vendor catalog every cycle. How should this recurring job be built?

    Sources1

    4.Tool use in a batch: one request, one answer

    The supported-features list includes both "tool use" and "multi-turn conversations", and it is easy to read more into that than it says. In the Messages API, multi-turn means you pass the earlier turns in the messages array yourself. The model then writes one next turn. A batch can't pause a request halfway, run a tool in your code, and feed the tool result back into that same request.

    These sources never state that limit in so many words. It follows from what they do say: each request is handled independently, results only become available once processing has ended, and they arrive as a single file with no streaming. If Claude's answer ends by asking for one of your tools, that request is already finished. To continue the loop, you build a new request with the tool result appended. In a batch, that means a whole new batch cycle of up to 24 hours per tool round. The server tools the docs list as supported (web search, web fetch, code execution and others) don't need your code between turns. Any agent loop that relies on your own tools does.

    Sources31

    5.Failures: resubmit only what failed, and test the prompt first

    Handling batch failures starts from the fact that a batch does not succeed or fail as a whole. Once processing has ended, every request has its own result, and request_counts on the batch tells you how many landed in each state. Only errored, canceled and expired requests need attention, and you are not billed for any of them. Because every result line is identified by custom_id, you can read the results file, collect the custom_ids whose result type is not succeeded, and build a new batch containing only those failed documents. Resubmitting the whole batch would pay a second time for every request that already succeeded.

    The four batch result types and what each one calls for
    Result typeWhat it meansWhat to do
    succeededThe request produced a messageKeep the result; do not resubmit
    erroredInvalid request or internal server error; not billedFix what caused the error, then resubmit that request by custom_id
    canceledYou canceled the batch before this request reached the model; not billedResubmit unchanged if the work is still wanted
    expiredThe 24-hour window passed before this request reached the model; not billedResubmit in a new batch; consider smaller batches if demand is high

    Resubmitting only failed documents is half the job; the other half is making appropriate modifications, and the kind of failure decides what those are. An expired request never reached the model, so it can go back in unchanged. An errored request was rejected as invalid or hit a server error, and an invalid request resubmitted unchanged is still invalid. Look at the error on the result line and fix the cause first. For documents that exceeded context limits, the appropriate modification is chunking: split each oversized document into chunks that each become their own request, with custom_ids such as doc-42-part-1 and doc-42-part-2 so the pieces can be reassembled. The custom_id is what makes this selective resubmission possible, because it identifies which document each failed line belongs to.

    The same 24-hour bound argues for prompt refinement on a sample set before batch-processing large volumes. Results are only available once processing has ended for all requests, so if the prompt is wrong you find out after a full cycle, and each round of iterative resubmission costs another cycle plus the tokens of every request that has to run again. The cheaper path is to run a small, representative sample set through the synchronous API first, check the outputs against the criteria you care about, and refine the prompt until the sample passes. Anthropic's guidance is to audit prompts against the model you actually run, and again whenever you change models; the sample run is where that audit happens. Only then submit the full volume. Refining on the sample maximizes the first-pass success rate of the large batch, which is what reduces iterative resubmission costs: most documents succeed the first time, and only a small remainder needs a second batch.

    Sources124

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Most batches finish within an hour, so batching is fine for a check that has to finish before a merge.Why is that wrong?

      The one-hour figure describes typical behaviour. The only firm bound is 24 hours, and requests still unprocessed at that point expire. A blocking check needs the synchronous API.

      Covered in The trade: half price, no promise on timing

    2. 2.Results come back in the order the requests were submitted, so you can match them by position.Why is that wrong?

      The results file is explicitly not guaranteed to be in request order. Matching by position silently attaches outputs to the wrong inputs. The custom_id on each result line is the only reliable key.

      Covered in The trade: half price, no promise on timing

    3. 3.A 30-hour SLA is met by submitting one batch a day, since most batches finish within an hour.Why is that wrong?

      Size the schedule on the worst case: the wait for the next submission plus up to 24 hours of processing. Daily submission can reach 48 hours; 4-hour windows cap it at 28.

      Covered in The trade: half price, no promise on timing

    4. 4.The batch docs list tool use as supported, so a batched request can run your tools mid-request and carry on with their results.Why is that wrong?

      A batched request produces one result, delivered in a file after processing ends. To continue a tool loop you submit a new request with the tool result appended, which means another batch cycle.

      Covered in Tool use in a batch: one request, one answer

    5. 5.One nightly batch can hold any volume, as long as processing finishes within 24 hours.Why is that wrong?

      Each batch has a hard cap on request count and on total size. Workloads above either cap have to be split across several batches.

      Covered in What a single batch can hold

    6. 6.If some requests in a batch errored, the simplest recovery is to resubmit the whole batch.Why is that wrong?

      Errored and expired requests are not billed, but the succeeded ones were, and running them again pays for them twice. Use request_counts to see what failed and custom_id to resubmit only those requests, fixed where the error demands it.

      Covered in Failures: resubmit only what failed, and test the prompt first

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “with most batches finishing in less than 1 hour while reducing costs by 50% and increasing throughput”
      ↩︎ The trade: half price, no promise on timing
      “Batches expire if processing does not complete within 24 hours.”
      ↩︎ The trade: half price, no promise on timing
      “you may see more requests expiring after 24 hours”
      ↩︎ The trade: half price, no promise on timing
      “A unique custom_id for identifying the Messages request. Must be 1 to 64 characters and contain only alphanumeric characters, hyphens, and underscores”
      ↩︎ The trade: half price, no promise on timing
      “A Message Batch is limited to either 100,000 Message requests or 256 MB in size, whichever is reached first.”
      ↩︎ What a single batch can hold
      “Rate limits apply to both Batches API HTTP requests and the number of requests within a batch waiting to be processed.”
      ↩︎ What a single batch can hold
      “Because each request in the batch is processed independently, you can mix different types of requests within a single batch.”
      ↩︎ What a single batch can hold
      “The batch is then processed asynchronously, with each request handled independently.”
      ↩︎ Tool use in a batch: one request, one answer
      “Batch results come back as a single file, not a stream.”
      ↩︎ Tool use in a batch: one request, one answer
      “The batch's request_counts shows an overview of your results, indicating how many requests reached each of these four states.”
      ↩︎ Failures: resubmit only what failed, and test the prompt first
      “Possible errors include invalid requests and internal server errors. You will not be billed for these requests.”
      ↩︎ Failures: resubmit only what failed, and test the prompt first
      “Batch reached its 24-hour expiration before this request could be sent to the model. You will not be billed for these requests.”
      ↩︎ Failures: resubmit only what failed, and test the prompt first
      “You can access batch results when all messages have completed or after 24 hours, whichever comes first.”
      ↩︎ Failures: resubmit only what failed, and test the prompt first
      “You can access batch results when all messages have completed or after 24 hours, whichever comes first.”
      ↩︎ Exam trap 1
      “Batch results come back as a single file, not a stream.”
      ↩︎ Exam trap 4
      “A Message Batch is limited to either 100,000 Message requests or 256 MB in size, whichever is reached first.”
      ↩︎ Exam trap 5
      “Possible errors include invalid requests and internal server errors. You will not be billed for these requests.”
      ↩︎ Exam trap 6
    2. 2.
      “Results are not guaranteed to be in the same order as requests. Use the custom_id field to match results to requests.”
      ↩︎ The trade: half price, no promise on timing
      “Developer-provided ID created for each request in a Message Batch. Useful for matching results to requests, as results may be given out of request order.”
      ↩︎ Failures: resubmit only what failed, and test the prompt first
      “Results are not guaranteed to be in the same order as requests. Use the custom_id field to match results to requests.”
      ↩︎ Exam trap 2
    3. 3.
      “Developer-provided ID created for each request in a Message Batch. Useful for matching results to requests, as results may be given out of request order.”
      ↩︎ The trade: half price, no promise on timing
      “Once a Message Batch is created, it begins processing immediately. Batches can take up to 24 hours to complete.”
      ↩︎ The trade: half price, no promise on timing
      “you specify the prior conversational turns with the messages parameter”
      ↩︎ Tool use in a batch: one request, one answer
      “Once a Message Batch is created, it begins processing immediately. Batches can take up to 24 hours to complete.”
      ↩︎ Exam trap 3
    4. 4.
      “Route every request no one is waiting on through a batch, and keep the interactive path for the rest.”
      ↩︎ Choose the API by asking who is waiting
      “Batching is the second-largest free lever after caching for unattended agent work: evaluation runs, backfills, and scheduled jobs”
      ↩︎ Choose the API by asking who is waiting
      “Auditing prompts against the model you run now, and again whenever you change models, is a free win.”
      ↩︎ Failures: resubmit only what failed, and test the prompt first
      “The Batch API takes 50% off every token of a request, including cached ones, in exchange for results arriving any time within 24 hours.”
      ↩︎ Key concept

    Ready to test yourself?

    Practise the 16 questions on this subdomain.