What you will be able to do
- Match a job such as sending a message, counting tokens, processing in bulk or uploading files to the right Claude API endpoint and its HTTP verb
- Name the required request headers and say which ones an official SDK sends for you
- Follow the server-sent event flow of a streamed response, including partial-JSON tool input and in-stream errors
- Choose between streaming and the Message Batches API for asynchronous work, and explain how rate limits and pagination affect client code
Key concept
The Claude API is a REST service that speaks JSON — Claude is reached through ordinary HTTP endpoints that take and return JSON, so the usual REST habits apply: resource paths, verbs, headers, status codes and paginated lists. Streaming and batching are two asynchronous ways of using those same endpoints, not a separate protocol.
1.One base URL, one endpoint per job
Much of this subdomain is ordinary software engineering applied to one particular service. The Claude API is a RESTful API at https://api.anthropic.com, so what you already do with any HTTP service still applies: a verb plus a resource path, a JSON body, and a status code when something goes wrong. What you have to learn is which endpoint does which job. Exam questions often describe a need, such as "check the size before sending" or "process thousands of prompts cheaply", and expect you to pick the endpoint that was built for it.
| API | Method and path | Purpose | Max request size |
|---|---|---|---|
| Messages API | POST /v1/messages | Send messages to Claude for conversational interactions | 32 MB |
| Message Batches API | POST /v1/messages/batches | Process large volumes of Messages requests asynchronously with 50% cost reduction | 256 MB |
| Token Counting API | POST /v1/messages/count_tokens | Count tokens in a message before sending to manage costs and rate limits | 32 MB |
| Models API | GET /v1/models | List available Claude models and their details | Not listed |
| Files API | POST /v1/files, GET /v1/files | Upload and manage files for use across multiple API calls | 500 MB |
Two details in that table matter in practice. First, counting tokens has its own endpoint. You don't set a flag on the Messages call; you ask a separate endpoint before sending anything, which lets you decide ahead of time whether a large prompt needs splitting. Second, the size limits behave like normal REST limits: a request over the limit gets a 413 request_too_large error. The overview also lists beta APIs for Claude Managed Agents (Agents, Sessions and Environments), which you reach through beta headers.
An engineer is processing a nightly batch of 50,000 support tickets through Claude to generate categorization labels. Each ticket is independent, latency is not a concern until the next morning, and the team wants to minimize cost. Which approach best fits this asynchronous workload?
Correct answer: A — Submit the tickets through the Batch API, which processes large volumes of independent requests asynchronously at roughly half the cost of standard synchronous calls
- A. Correct. Batch processing is designed exactly for large volumes of independent, non-latency-sensitive requests and is billed at a discount versus standard calls.
- B. Incorrect. Streaming is meant for incremental delivery of a single response, not for fanning out thousands of unrelated ticket classifications in one conversation.
- C. Incorrect. Looping synchronous calls with no batching or concurrency control wastes cost and ignores the discount and throughput benefits the Batch API offers.
- D. Incorrect. Resuming a session replays conversation context for a single continued exchange; it does not deduplicate billing across unrelated tickets.
Sources1
2.Headers, JSON bodies and what the SDK does for you
Every request carries a small, fixed set of headers. Each one maps to a familiar engineering concern: authentication, versioning, wire format and tenancy.
| Header | Value | Required |
|---|---|---|
| Authorization | Bearer <token>: an API key or a short-lived token from Workload Identity Federation | Yes, unless x-api-key is set |
| x-api-key | Your API key from Console (legacy fallback) | No |
| anthropic-version | API version, for example 2023-06-01 | Yes |
| content-type | application/json | Yes |
| anthropic-workspace-id | ID of the workspace the request runs in | Required with a multi-workspace API key |
Because content-type is application/json, JSON is the wire format for the Messages API, just as it would be between any two internal microservices. The anthropic-version header pins the API version your code was written against. Treat it like any other version pin in your configuration: change it on purpose, never by accident. Responses carry headers too. The most useful is request-id, a globally unique identifier you include when contacting support about a specific call, so it belongs in your logs.
The official SDKs remove most of this boilerplate. They handle headers automatically (authentication, anthropic-version, content-type) and add type-safe request and response handling, built-in retry logic and error handling, streaming support, and request timeouts. There is one exception: when your key needs anthropic-workspace-id, you pass it yourself.
Sources1
3.Streaming: asynchronous delivery over server-sent events
Streaming is the first asynchronous pattern. The connection stays open while the response is still being produced, and the response arrives as a series of events. In the SDKs, .stream() keeps the HTTP connection alive with server-sent events. .get_final_message() in Python, or .finalMessage() in TypeScript, then collects every event and returns the complete Message object. The event order is fixed: message_start, which carries a Message with empty content; then, for each content block, a content_block_start, one or more content_block_delta events and a content_block_stop; then one or more message_delta events; and finally message_stop. Each block has an index that matches its position in the final Message content array. Ping events can appear anywhere in the stream.
event: content_block_delta
data: {"type": "content_block_delta","index": 1,"delta": {"type": "input_json_delta","partial_json": "{\"location\": \"San Fra"}}}It is something else. Tool input deltas are fragments of a JSON string, and only the final tool_use.input is an object. You collect the fragments and parse them once content_block_stop arrives, either with a library that handles partial JSON, such as Pydantic, or with the SDK helpers. A 200 response doesn't guarantee a clean stream, either. Errors can arrive inside the stream: an overloaded_error during heavy usage, for example, which would normally be an HTTP 529 in a non-streaming call.
event: error
data: {"type": "error", "error": {"type": "overloaded_error", "message": "Overloaded"}}Sources2
4.Batches, rate limits and pagination
The second asynchronous pattern is about volume, not latency. The Message Batches API processes large volumes of Messages requests asynchronously at a 50% cost reduction, and it accepts requests of up to 256 MB, against 32 MB for Messages. Streaming delivers one response bit by bit over an open connection. Batching hands over many requests and collects the results later.
Whichever you use, your client runs inside limits. Organisations are placed on usage tiers automatically. Each tier has a spend limit (maximum monthly cost) and rate limits measured in requests per minute (RPM) and tokens per minute (TPM). That is another reason the Token Counting endpoint is useful before a large submission.
List endpoints follow standard REST pagination. You set limit for the page size, pass an opaque page cursor, and read next_page from each response until it comes back null. In Python and TypeScript you can simply iterate over the list result and the SDK follows next_page for you. That auto-pagination only goes forward; to go back, you read prev_page yourself, and only GET /v1/sessions returns it.
Message Batches. It exists to process large volumes of Messages requests asynchronously, at half the cost. Streaming keeps one HTTP connection open to deliver one response incrementally, which doesn't help when nobody is waiting on an individual answer.
A team wants every file edit that an autonomous coding agent makes to be recorded in a version-control-friendly audit log before the change is written to disk, so that the change history stays traceable alongside normal git commits. Which Agent SDK mechanism should they use?
Correct answer: A — Register a PreToolUse hook matched to the Edit and Write tools that appends the intended change to the audit log before the tool executes
- A. Correct. A PreToolUse hook matched on Edit and Write runs custom logic, such as appending to an audit file, before the tool executes, which is exactly the interception point needed.
- B. Incorrect. The SDK does not automatically create git commits for tool calls; commits only happen if the agent or a hook explicitly runs git commands.
- C. Incorrect. acceptEdits only controls whether edit-type tool calls are auto-approved without a prompt; it does not produce any audit logging on its own.
- D. Incorrect. A subagent with no tools cannot observe tool calls made by a separate agent context; hooks, not passive subagents, are the interception mechanism.
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Each streamed tool_use delta contains a complete input object that you can parse on arrival.Why is that wrong?
The deltas are fragments of a JSON string. You collect them and parse once content_block_stop arrives; only the final tool_use.input is an object.
Covered in Streaming: asynchronous delivery over server-sent events
2.Once a streaming request returns HTTP 200, the stream can't fail, so error handling only needs to check status codes.Why is that wrong?
Errors such as overloaded_error can arrive as an error event inside a stream that has already started, so the stream consumer has to handle them.
Covered in Streaming: asynchronous delivery over server-sent events
3.When you use an official SDK, you still have to set the anthropic-version and content-type headers on every request yourself.Why is that wrong?
The SDKs send the authentication, version and content-type headers automatically. The one header you may still have to pass is anthropic-workspace-id, when your key needs it.
Covered in Headers, JSON bodies and what the SDK does for you
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/api/overviewOfficial docs
“Messages API: Send messages to Claude for conversational interactions (POST /v1/messages)”
↩︎ One base URL, one endpoint per job“Count tokens in a message before sending to manage costs and rate limits”
↩︎ One base URL, one endpoint per job“the SDK sends the authentication, version, and content-type headers automatically”
↩︎ Headers, JSON bodies and what the SDK does for you“Include it when you contact support about a specific request.”
↩︎ Headers, JSON bodies and what the SDK does for you“Process large volumes of Messages requests asynchronously with 50% cost reduction”
↩︎ Batches, rate limits and pagination“Rate limits: Maximum number of requests per minute (RPM) and tokens per minute (TPM)”
↩︎ Batches, rate limits and pagination“SDK auto-pagination is forward-only”
↩︎ Batches, rate limits and pagination“The Claude API is a RESTful API at https://api.anthropic.com that provides programmatic access to Claude models and Claude Managed Agents.”
↩︎ Key concept“the SDK sends the authentication, version, and content-type headers automatically”
↩︎ Exam trap 3 - 2.
“The .stream() call keeps the HTTP connection alive with server-sent events”
↩︎ Streaming: asynchronous delivery over server-sent events“your code should handle unknown event types gracefully”
↩︎ Streaming: asynchronous delivery over server-sent events“the deltas are partial JSON strings, whereas the final tool_use.input is always an object.”
↩︎ Exam trap 1“The API may occasionally send errors in the event stream.”
↩︎ Exam trap 2