CertSafari
    CLAUDE-CERTIFIED-DEVELOPER-FOUNDATIONS-CCDV-F · Lessons

    Domain 2 · Lesson 7/25

    Claude API as a REST and JSON Service: Endpoints, Headers, Streaming and Batches

    Software Engineering Foundations

    9 min read
    5.52% of exam
    2 sources
    Published 29 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Match a job such as sending a message, counting tokens, processing in bulk or uploading files to the right Claude API endpoint and its HTTP verb
    • Name the required request headers and say which ones an official SDK sends for you
    • Follow the server-sent event flow of a streamed response, including partial-JSON tool input and in-stream errors
    • Choose between streaming and the Message Batches API for asynchronous work, and explain how rate limits and pagination affect client code

    Key concept

    The Claude API is a REST service that speaks JSON — Claude is reached through ordinary HTTP endpoints that take and return JSON, so the usual REST habits apply: resource paths, verbs, headers, status codes and paginated lists. Streaming and batching are two asynchronous ways of using those same endpoints, not a separate protocol.

    1.One base URL, one endpoint per job

    Much of this subdomain is ordinary software engineering applied to one particular service. The Claude API is a RESTful API at https://api.anthropic.com, so what you already do with any HTTP service still applies: a verb plus a resource path, a JSON body, and a status code when something goes wrong. What you have to learn is which endpoint does which job. Exam questions often describe a need, such as "check the size before sending" or "process thousands of prompts cheaply", and expect you to pick the endpoint that was built for it.

    Generally available Claude API endpoints, what each is for, and its documented maximum request size
    APIMethod and pathPurposeMax request size
    Messages APIPOST /v1/messagesSend messages to Claude for conversational interactions32 MB
    Message Batches APIPOST /v1/messages/batchesProcess large volumes of Messages requests asynchronously with 50% cost reduction256 MB
    Token Counting APIPOST /v1/messages/count_tokensCount tokens in a message before sending to manage costs and rate limits32 MB
    Models APIGET /v1/modelsList available Claude models and their detailsNot listed
    Files APIPOST /v1/files, GET /v1/filesUpload and manage files for use across multiple API calls500 MB

    Two details in that table matter in practice. First, counting tokens has its own endpoint. You don't set a flag on the Messages call; you ask a separate endpoint before sending anything, which lets you decide ahead of time whether a large prompt needs splitting. Second, the size limits behave like normal REST limits: a request over the limit gets a 413 request_too_large error. The overview also lists beta APIs for Claude Managed Agents (Agents, Sessions and Environments), which you reach through beta headers.

    An engineer is processing a nightly batch of 50,000 support tickets through Claude to generate categorization labels. Each ticket is independent, latency is not a concern until the next morning, and the team wants to minimize cost. Which approach best fits this asynchronous workload?

    Sources1

    2.Headers, JSON bodies and what the SDK does for you

    Every request carries a small, fixed set of headers. Each one maps to a familiar engineering concern: authentication, versioning, wire format and tenancy.

    Request headers for the Claude API and when each is required
    HeaderValueRequired
    AuthorizationBearer <token>: an API key or a short-lived token from Workload Identity FederationYes, unless x-api-key is set
    x-api-keyYour API key from Console (legacy fallback)No
    anthropic-versionAPI version, for example 2023-06-01Yes
    content-typeapplication/jsonYes
    anthropic-workspace-idID of the workspace the request runs inRequired with a multi-workspace API key

    Because content-type is application/json, JSON is the wire format for the Messages API, just as it would be between any two internal microservices. The anthropic-version header pins the API version your code was written against. Treat it like any other version pin in your configuration: change it on purpose, never by accident. Responses carry headers too. The most useful is request-id, a globally unique identifier you include when contacting support about a specific call, so it belongs in your logs.

    The official SDKs remove most of this boilerplate. They handle headers automatically (authentication, anthropic-version, content-type) and add type-safe request and response handling, built-in retry logic and error handling, streaming support, and request timeouts. There is one exception: when your key needs anthropic-workspace-id, you pass it yourself.

    Sources1

    3.Streaming: asynchronous delivery over server-sent events

    Streaming is the first asynchronous pattern. The connection stays open while the response is still being produced, and the response arrives as a series of events. In the SDKs, .stream() keeps the HTTP connection alive with server-sent events. .get_final_message() in Python, or .finalMessage() in TypeScript, then collects every event and returns the complete Message object. The event order is fixed: message_start, which carries a Message with empty content; then, for each content block, a content_block_start, one or more content_block_delta events and a content_block_stop; then one or more message_delta events; and finally message_stop. Each block has an index that matches its position in the final Message content array. Ping events can appear anywhere in the stream.

    A tool_use delta: partial_json is a fragment of a JSON string, not an objecttext
    event: content_block_delta
    data: {"type": "content_block_delta","index": 1,"delta": {"type": "input_json_delta","partial_json": "{\"location\": \"San Fra"}}}

    It is something else. Tool input deltas are fragments of a JSON string, and only the final tool_use.input is an object. You collect the fragments and parse them once content_block_stop arrives, either with a library that handles partial JSON, such as Pydantic, or with the SDK helpers. A 200 response doesn't guarantee a clean stream, either. Errors can arrive inside the stream: an overloaded_error during heavy usage, for example, which would normally be an HTTP 529 in a non-streaming call.

    An error delivered inside an event stream rather than as an HTTP statustext
    event: error
    data: {"type": "error", "error": {"type": "overloaded_error", "message": "Overloaded"}}

    Sources2

    4.Batches, rate limits and pagination

    The second asynchronous pattern is about volume, not latency. The Message Batches API processes large volumes of Messages requests asynchronously at a 50% cost reduction, and it accepts requests of up to 256 MB, against 32 MB for Messages. Streaming delivers one response bit by bit over an open connection. Batching hands over many requests and collects the results later.

    Whichever you use, your client runs inside limits. Organisations are placed on usage tiers automatically. Each tier has a spend limit (maximum monthly cost) and rate limits measured in requests per minute (RPM) and tokens per minute (TPM). That is another reason the Token Counting endpoint is useful before a large submission.

    List endpoints follow standard REST pagination. You set limit for the page size, pass an opaque page cursor, and read next_page from each response until it comes back null. In Python and TypeScript you can simply iterate over the list result and the SDK follows next_page for you. That auto-pagination only goes forward; to go back, you read prev_page yourself, and only GET /v1/sessions returns it.

    A team wants every file edit that an autonomous coding agent makes to be recorded in a version-control-friendly audit log before the change is written to disk, so that the change history stays traceable alongside normal git commits. Which Agent SDK mechanism should they use?

    Sources1

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Each streamed tool_use delta contains a complete input object that you can parse on arrival.Why is that wrong?

      The deltas are fragments of a JSON string. You collect them and parse once content_block_stop arrives; only the final tool_use.input is an object.

      Covered in Streaming: asynchronous delivery over server-sent events

    2. 2.Once a streaming request returns HTTP 200, the stream can't fail, so error handling only needs to check status codes.Why is that wrong?

      Errors such as overloaded_error can arrive as an error event inside a stream that has already started, so the stream consumer has to handle them.

      Covered in Streaming: asynchronous delivery over server-sent events

    3. 3.When you use an official SDK, you still have to set the anthropic-version and content-type headers on every request yourself.Why is that wrong?

      The SDKs send the authentication, version and content-type headers automatically. The one header you may still have to pass is anthropic-workspace-id, when your key needs it.

      Covered in Headers, JSON bodies and what the SDK does for you

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Messages API: Send messages to Claude for conversational interactions (POST /v1/messages)”
      ↩︎ One base URL, one endpoint per job
      “Count tokens in a message before sending to manage costs and rate limits”
      ↩︎ One base URL, one endpoint per job
      “the SDK sends the authentication, version, and content-type headers automatically”
      ↩︎ Headers, JSON bodies and what the SDK does for you
      “Include it when you contact support about a specific request.”
      ↩︎ Headers, JSON bodies and what the SDK does for you
      “Process large volumes of Messages requests asynchronously with 50% cost reduction”
      ↩︎ Batches, rate limits and pagination
      “Rate limits: Maximum number of requests per minute (RPM) and tokens per minute (TPM)”
      ↩︎ Batches, rate limits and pagination
      “SDK auto-pagination is forward-only”
      ↩︎ Batches, rate limits and pagination
      “The Claude API is a RESTful API at https://api.anthropic.com that provides programmatic access to Claude models and Claude Managed Agents.”
      ↩︎ Key concept
      “the SDK sends the authentication, version, and content-type headers automatically”
      ↩︎ Exam trap 3
    2. 2.
      “The .stream() call keeps the HTTP connection alive with server-sent events”
      ↩︎ Streaming: asynchronous delivery over server-sent events
      “your code should handle unknown event types gracefully”
      ↩︎ Streaming: asynchronous delivery over server-sent events
      “the deltas are partial JSON strings, whereas the final tool_use.input is always an object.”
      ↩︎ Exam trap 1
      “The API may occasionally send errors in the event stream.”
      ↩︎ Exam trap 2

    Continue to page 2 of 2

    Claude in the SDLC: Version Control, CI, Code Review and Refactoring