What you will be able to do
- Name the main Claude API endpoints and the HTTP method and path each one uses
- List the headers a raw HTTP request needs, and say which ones an official SDK sends for you
- Explain what an official SDK adds on top of the REST API, including typed errors and automatic retries
- Read an error status code and pick the right response: fix the request, back off and retry, or switch to streaming
- Follow the basic engineering practices for integrating an official SDK that wraps the REST API: install it, load the key from the environment, choose a sync or async client, and map each SDK method to its endpoint
- Explain that Claude streaming runs over server-sent events on an ordinary HTTP response, and say how that differs from a websocket connection
Key concept
SDK as a wrapper over a stateless REST API — Every Claude API call is a single HTTP request to a REST endpoint, and the server keeps no conversation state between calls. The official SDKs add headers, typed responses, retries and streaming on top of those requests, but the requests underneath are unchanged.
1.The API is a set of REST endpoints
Start with what the Claude API is. It is a RESTful API served from https://api.anthropic.com. Each capability is an HTTP method plus a path, so an integration comes down to choosing the right endpoint and sending a JSON body to it. Everything else in this lesson builds on that. SDK methods map onto these endpoints, error codes are ordinary HTTP status codes, and even streaming is an HTTP response that stays open and delivers server-sent events. There is no separate websocket protocol to learn for the Messages API; the last section of this lesson explains how server-sent events differ from websockets and why the distinction matters when you integrate.
| API | Endpoint | Purpose |
|---|---|---|
| Messages API | POST /v1/messages | Send messages to Claude for conversational interactions |
| Message Batches API | POST /v1/messages/batches | Process large volumes of Messages requests asynchronously with 50% cost reduction |
| Token Counting API | POST /v1/messages/count_tokens | Count tokens in a message before sending it |
| Models API | GET /v1/models | List available Claude models and their details |
| Files API | POST /v1/files, GET /v1/files | Upload and manage files for use across multiple API calls |
Model selection, batching and token counting all show up in other parts of this domain. Here the point is only that each of them is a REST endpoint you can reach directly or through an SDK.
2.Headers every raw request must carry
If you skip the SDK and call the API with plain HTTP, you have to set the headers yourself. Three are always required: a credential, an API version, and a content type. For the credential, the documented header is Authorization: Bearer <token>, where the token is your API key or a short-lived token from Workload Identity Federation. The older x-api-key header still works as a fallback. Separately, a key that can access several workspaces must also say which workspace it is acting in.
| Header | Value | Required |
|---|---|---|
| Authorization | Bearer <token> (API key or short-lived access token) | Yes, unless x-api-key is set |
| x-api-key | Your API key from Console (legacy fallback) | No |
| anthropic-version | API version, for example 2023-06-01 | Yes |
| content-type | application/json | Yes |
| anthropic-workspace-id | ID of the workspace the request runs in | Required with a multi-workspace API key |
It should not, because anthropic-version is marked required. This is the main practical difference an SDK makes: it sends the authentication, version and content-type headers for you. It does not send anthropic-workspace-id. You still pass that yourself when your key needs it.
A frontend engineer building a custom SSE parser for streaming Messages API responses needs to know when it is safe to treat one content block as fully received and move on to parsing the next one. Which event signals that a specific content block, identified by its index, will receive no further delta events?
Correct answer: A — content_block_stop
- A. Correct — content_block_stop is emitted once a content block at a given index has received all of its delta events and will not change further.
- B. message_delta communicates top-level changes to the Message object, such as stop_reason and cumulative usage, not the completion of an individual content block.
- C. content_block_start marks the beginning of a new content block and precedes its delta events, not the end.
- D. message_stop signals the end of the entire stream, after all content blocks have already closed with their own content_block_stop events.
Sources1
3.Integrating with SDKs that wrap the REST API
The exam guide lists integrating with SDKs that wrap REST APIs as a basic engineering practice, and the Claude SDKs are a clean example of the pattern. Official client SDKs exist for Python, TypeScript, C#, Go, Java, PHP and Ruby. According to the API overview, they give you automatic header management, type-safe requests and responses, built-in retries and error handling, streaming support, and request timeouts and connection management. None of that changes the REST API underneath; each SDK method still turns into one HTTP request to one endpoint. The Python SDK needs Python 3.10 or later and has both a synchronous client (Anthropic) and an asynchronous one (AsyncAnthropic). The client reads ANTHROPIC_API_KEY from the environment by default.
The integration steps are the same in every language. Install the package (pip install anthropic or npm install @anthropic-ai/sdk). Keep the API key out of source code and let the client read it from the environment. Construct one client and reuse it. Then call the method that wraps the endpoint you want, and read the typed response object instead of parsing JSON yourself.
client = Anthropic(
# This is the default and can be omitted
api_key=os.environ.get("ANTHROPIC_API_KEY"),
)
message = client.messages.create(
max_tokens=1024,
messages=[
{
"role": "user",
"content": "Hello, Claude",
}
],
model="claude-opus-5-5",
)const client = new Anthropic({
apiKey: process.env["ANTHROPIC_API_KEY"] // This is the default and can be omitted
});
const message = await client.messages.create({
max_tokens: 1024,
messages: [{ role: "user", content: "Hello, Claude" }],
model: "claude-opus-5-5"
});
for (const block of message.content) {
if (block.type === "text") {
console.log(block.text);
}
}Read the two snippets against the endpoint table in the first section. client.messages.create is a wrapper around POST /v1/messages. In the same way, client.messages.count_tokens wraps the Token Counting endpoint, and client.messages.batches wraps the Batches endpoint. The parameter names (model, max_tokens, messages) are the JSON body fields of the REST request, so knowledge of the raw API transfers directly to any SDK. The TypeScript SDK follows the same shape and adds one safety default worth knowing: it does not run in web browsers unless you explicitly set dangerouslyAllowBrowser to true. The reason is that a browser would expose your secret API key. That default is itself an engineering practice: the SDK belongs on a server you control, not in code shipped to end users.
Header management (authentication, anthropic-version, content-type), typed request and response handling, retry logic with error handling, and streaming support. The official SDK supplies all of these, plus request timeouts and connection management.
4.Stateless requests and measuring tokens
REST calls to the Messages API are stateless. The server does not remember earlier turns, so each request has to include the whole conversation up to that point. That is also why earlier assistant turns in the messages array can be synthetic: the server has no record to compare them against.
Claude sees only that one message. The Messages API is stateless, so the client has to append every prior user and assistant turn to the messages array on every call.
Because the whole history is sent every time, input size grows as the conversation grows. The SDK lets you measure it in two ways. Each response has a usage property that reports input_tokens and output_tokens. You can also count tokens before sending a request.
count = client.messages.count_tokens(
model="claude-opus-5-5",
messages=[{"role": "user", "content": "Hello, world"}],
)
print(count.input_tokens) # 10Sources6
5.HTTP errors, typed exceptions and retries
Since this is a REST API, failures come back as HTTP status codes with a JSON body. That body has a top-level error object with a type and a message, plus a request_id you can give to support. The SDKs turn these into typed exceptions; in Python, for example, a 404 becomes anthropic.NotFoundError. The documentation says to catch those typed classes, most specific first, and not to string-match error messages.
| Status | Error type | Response |
|---|---|---|
| 400 | invalid_request_error | Fix the request format or content; also returned when a spend limit you set is reached |
| 401 | authentication_error | Check the API key (malformed, revoked or expired) |
| 413 | request_too_large | Shrink the request; the Messages API limit is 32 MB, Batch API 256 MB |
| 429 | rate_limit_error | Rate limit or spend cap hit; a tier spend-cap 429 has no retry-after header |
| 500 | api_error | Retry with exponential backoff |
| 504 | timeout_error | Consider the streaming Messages API for long-running requests |
| 529 | overloaded_error | The API is temporarily overloaded |
You usually don't need to write your own retry loop for transient failures. The official SDKs already retry connection errors, rate limits and 5xx errors with exponential backoff, twice by default, and they honor retry-after when it is present. Use the max_retries client option to change the count or turn retries off.
Sources7
6.Streaming transport: server-sent events, not websockets
The exam guide names websockets alongside SDK integration as a basic engineering practice, so it helps to be precise about which real-time mechanism the Claude API actually uses. A websocket is a persistent, bidirectional connection: after an HTTP handshake, both client and server can send frames to each other at any time over the same socket. Server-sent events (SSE) are simpler. The client makes an ordinary HTTP request, and the server keeps the response open and writes a sequence of named events to it, one way, until it closes. The documented streaming path for the Messages API is SSE: you set "stream": true on the same POST /v1/messages request, and the response arrives incrementally instead of as one JSON document.
client = anthropic.Anthropic()
with client.messages.stream(
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}],
model="claude-opus-5-5",
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)Each server-sent event has a named event type and a JSON payload. A stream opens with message_start, then delivers content blocks as content_block_start, content_block_delta and content_block_stop events, followed by message_delta events and a final message_stop. The SDK helpers accumulate these for you and hand back the complete Message through get_final_message() in Python or finalMessage() in TypeScript. Two consequences matter for integration. First, because SSE is one-directional, you cannot send a follow-up turn down the same connection the way you might over a websocket; a new turn is a new stateless POST with the full history. Second, an error can arrive as an error event after the API has already returned HTTP 200, so streaming code needs to watch the event stream, not only the status code.
They should send the normal Messages API POST with stream set to true, or use the SDK's messages.stream() helper, and read server-sent events from that HTTP response. The SDKs also require streaming for requests with large max_tokens values, to avoid HTTP timeouts, even when the caller only wants the final Message.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The SDK surfaces every 429 or 5xx straight to your code, so you have to build your own backoff loop before production.Why is that wrong?
The official SDKs already retry transient failures with exponential backoff (twice by default) and honor retry-after. You change this with max_retries.
Covered in HTTP errors, typed exceptions and retries
2.A 400 always means the JSON body is malformed.Why is that wrong?
400 is also returned when usage reaches an organization or workspace spend limit you set, so a correctly formed request can still get one.
Covered in HTTP errors, typed exceptions and retries
3.The API keeps conversation state on the server, so each request only needs the newest user message.Why is that wrong?
The Messages API is stateless. The client sends the full conversation history on every call.
Covered in Stateless requests and measuring tokens
4.Streaming a Claude response requires opening a websocket connection to the API.Why is that wrong?
Streaming is done with server-sent events on the ordinary Messages POST: set stream to true and read events from the HTTP response that stays open.
Covered in Streaming transport: server-sent events, not websockets
5.The TypeScript SDK can be dropped into a browser bundle like any other npm package.Why is that wrong?
Browser support is disabled by default because it would expose your secret API key; you must set dangerouslyAllowBrowser to true, and the sensible practice is to keep the SDK on a server.
Covered in Integrating with SDKs that wrap the REST API
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/api/overviewOfficial docs
“The Claude API is a RESTful API at https://api.anthropic.com that provides programmatic access to Claude models and Claude Managed Agents.”
↩︎ The API is a set of REST endpoints“Message Batches API: Process large volumes of Messages requests asynchronously with 50% cost reduction (POST /v1/messages/batches)”
↩︎ The API is a set of REST endpoints“If you are using the Client SDKs, the SDK sends the authentication, version, and content-type headers automatically”
↩︎ Headers every raw request must carry“Anthropic provides official SDKs that simplify API integration by handling authentication, request formatting, error handling, and more.”
↩︎ Integrating with SDKs that wrap the REST API“Anthropic provides official SDKs that simplify API integration by handling authentication, request formatting, error handling, and more.”
↩︎ Key concept - 2.
“The .stream() call keeps the HTTP connection alive with server-sent events”
↩︎ The API is a set of REST endpoints“When creating a Message, you can set "stream": true to incrementally stream the response using server-sent events (SSE).”
↩︎ Streaming transport: server-sent events, not websockets“Each server-sent event includes a named event type and associated JSON data.”
↩︎ Streaming transport: server-sent events, not websockets“This is especially useful for requests with large max_tokens values, where the SDKs require streaming to avoid HTTP timeouts.”
↩︎ Streaming transport: server-sent events, not websockets“When creating a Message, you can set "stream": true to incrementally stream the response using server-sent events (SSE).”
↩︎ Exam trap 4 - 3.
“Each SDK provides idiomatic interfaces, type safety, and built-in support for streaming, retries, and error handling.”
↩︎ Integrating with SDKs that wrap the REST API - 4.
“The Anthropic Python SDK provides convenient access to the Claude API from Python applications.”
↩︎ Integrating with SDKs that wrap the REST API - 5.
“Web browsers: disabled by default to avoid exposing your secret API credentials”
↩︎ Integrating with SDKs that wrap the REST API“Web browsers: disabled by default to avoid exposing your secret API credentials”
↩︎ Exam trap 5 - 6.
“The Messages API is stateless, which means that you always send the full conversational history to the API.”
↩︎ Stateless requests and measuring tokens“The Messages API is stateless, which means that you always send the full conversational history to the API.”
↩︎ Exam trap 3 - 7.https://platform.claude.com/docs/en/api/errorsOfficial docs
“Catch the SDK's typed classes rather than string-matching error messages, handling the most specific classes first.”
↩︎ HTTP errors, typed exceptions and retries“If you exceed these limits, you'll receive a 413 request_too_large error.”
↩︎ HTTP errors, typed exceptions and retries“When receiving a streaming response over server-sent events (SSE), an error can occur after the API returns a 200 response.”
↩︎ Streaming transport: server-sent events, not websockets“The official SDKs automatically retry transient failures (such as connection errors, rate limits, and 5xx server errors) with exponential backoff, twice by default”
↩︎ Exam trap 1“The API also returns a 400 when usage reaches an organization or workspace spend limit you set”
↩︎ Exam trap 2