What you will be able to do
- Read a Messages API response and explain its content array, stop_reason and usage fields
- Carry multi-turn state correctly through a stateless API, including system instructions and prefill
- Send images with the right source type and refer to several images across turns
- Consume a streamed response, including tool_use input deltas and thinking deltas
Key concept
The Messages API is stateless — The API keeps no conversation for you. Every request must carry the whole transcript you want Claude to see: earlier turns, images and system instructions. Storing, trimming and replaying that history is your application's job.
1.One request in, one Message out
A basic Messages API call has three parts: a model, a max_tokens cap, and a messages list of turns, each with a role and some content. The reply is a Message object. It has an id, role: "assistant", a content array of typed blocks, the model that served the request, a stop_reason, and a usage object that counts input and output tokens.
{
"id": "msg_01XFDUDYJgAACzvnptvVoYEL",
"type": "message",
"role": "assistant",
"content": [
{
"type": "text",
"text": "Hello!"
}
],
"model": "claude-opus-5-5",
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {
"input_tokens": 12,
"output_tokens": 6
}
}content is an array, not a string, because one reply can hold several blocks: text here, and elsewhere tool_use or thinking blocks, which the streaming section covers. stop_reason tells you why generation ended. "end_turn" means Claude finished naturally. "max_tokens" means the cap cut the reply off, which is sometimes intended. A refusal comes back as stop_reason: "refusal" with a stop_details object that names the policy category behind it.
Sources1
2.Conversation state is yours to carry
Since the API is stateless, a multi-turn conversation is a list your application keeps and resends in full each time. This is the main data access pattern of the Messages API. Nothing on the server lets a follow-up request refer back to an earlier one, so persisting, trimming and replaying the transcript are all up to you. Earlier assistant turns also don't have to come from Claude. You can write synthetic assistant messages to set up a conversation.
Instructions that apply from the first turn go in the top-level system field. On Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Opus 5.5, Claude Opus 4.8 and Claude Opus 5, you can also add a "role": "system" message after a user turn to bring in a new instruction partway through. It has the same authority as the top-level field. It cannot be the first entry in messages. Because it is appended at the end, it does not invalidate the cached prefix before it.
The content is the single text block "C" and stop_reason is "max_tokens". This technique is called prefill. The final assistant turn is Claude's partial reply, and Claude continues from it. Here the one-token cap is deliberate, so max_tokens is the expected stop reason, not an error.
Sources1
3.Sending images: three source types
Images go into a user turn as image content blocks, next to text blocks. Supported media types are image/jpeg, image/png, image/gif and image/webp. Vision also works in claude.ai and in the Claude Console Playground, but on the API the choice that matters is where the image bytes come from:
| Source type | What the request carries | When it fits |
|---|---|---|
| base64 | The encoded image data plus its media_type, inside the request body | One-off images you already hold in memory |
| url | A URL pointing to an image hosted online | Images that are already publicly reachable |
| file | A file_id returned by the Files API | Images you will send repeatedly: upload once, reference many times |
One request can hold several images, and Claude analyses them together, which suits comparisons or a series of document pages. The docs recommend putting a short text label before each image (Image 1:, Image 2:) so you and later turns can refer to them by name. Combine this with statelessness: a follow-up that mentions "Image 1" only works if the earlier turn containing that image is still in the history you resend. The Files API makes resending cheap, because the history holds a file_id reference instead of the base64 data again.
A compliance-document assistant sends the same 40,000-token reference manual as cached context, but users only ask questions roughly every 20-30 minutes, well beyond the default cache lifetime. Which caching configuration change best fits this access pattern?
Correct answer: A — Set the ephemeral cache breakpoint's `ttl` to `"1h"` on the reference manual block, accepting the higher write cost for reduced re-write frequency over the longer idle gaps
- A. Correct. The 1-hour cache duration is designed for content used less frequently than every 5 minutes but still within an hour; paying the 2x write cost once and reading at 0.1x for a subsequent question in the same hour is more efficient than repeatedly re-writing the full 5-minute cache.
- B. A second breakpoint helps manage the 20-block lookback for growing conversations, but it doesn't extend the underlying cache lifetime past 5 minutes on its own; the cache still expires if no request arrives within that window.
- C. Automatic caching simplifies breakpoint placement to a single parameter, but it does not change the underlying TTL options; the default lifetime is still 5 minutes unless the 1-hour TTL is explicitly requested.
- D. Shrinking content below the minimum threshold makes it ineligible for caching altogether rather than making it persist longer, which is the opposite of the desired outcome.
4.Streaming: events, tool input and thinking
Set "stream": true and the reply arrives as server-sent events (SSE). The Python and TypeScript SDKs wrap this for you:
with client.messages.stream(
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}],
model="claude-opus-5-5",
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)Streaming also matters when you don't show partial output. With large max_tokens values the SDKs require streaming to avoid HTTP timeouts. .get_final_message() (Python) or .finalMessage() (TypeScript) streams internally and returns the same complete Message that .create() would. The raw event order is: message_start (a Message with empty content), then for each content block a content_block_start, one or more content_block_delta events and a content_block_stop, then one or more message_delta events and a final message_stop. Each block's index matches its position in the final content array. There can be ping events at any point. Errors can also arrive inside the stream, for example an overloaded_error that would be an HTTP 529 in a non-streaming call. New event types may be added, so your parser should ignore types it doesn't recognise.
| Delta type | Block it updates | Payload |
|---|---|---|
| text_delta | text | A fragment of the text |
| input_json_delta | tool_use | partial_json: a fragment of the tool input as a JSON string |
| thinking_delta | thinking | A fragment of the thinking field |
| signature_delta | thinking | The signature, sent just before content_block_stop |
Tool calls need care when streamed. The deltas for a tool_use block are partial JSON strings, while the final tool_use.input is always an object. Collect the fragments and parse them when content_block_stop arrives, or use the SDK helpers that give you parsed values as they build up. Current models emit one complete key and value at a time, so there can be pauses between events while a tool call is being written. Thinking blocks stream as thinking_delta events and close with a signature_delta, which verifies the block's integrity. With display: "omitted", no thinking text is streamed: the block receives one empty thinking_delta and a signature, then closes.
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.After the first request, the API remembers the conversation, so a follow-up only needs to send the new user message.Why is that wrong?
The API is stateless. Each request must contain the full history you want Claude to consider, and your application stores it.
Covered in Conversation state is yours to carry
2.Once Claude has seen an image in an earlier turn, a later request can refer to 'the first image' without including it again.Why is that wrong?
The earlier turn, with its image, has to be in the resent history. To avoid re-encoding it, upload it to the Files API and reference the file_id in each turn.
Covered in Sending images: three source types
3.Each input_json_delta for a tool call is a JSON object you can parse on its own.Why is that wrong?
The deltas are fragments of a JSON string. Only the collected result, the final tool_use.input, is an object, so parse after content_block_stop or use the SDK helpers.
Covered in Streaming: events, tool input and thinking
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Refusal responses (stop_reason: "refusal") also include a stop_details object identifying the policy category that triggered the refusal, on every model.”
↩︎ One request in, one Message out“Earlier conversational turns don't necessarily need to actually originate from Claude.”
↩︎ Conversation state is yours to carry“A system message cannot be the first entry in messages.”
↩︎ Conversation state is yours to carry“You can pre-fill part of Claude's response in the last position of the input messages list.”
↩︎ Conversation state is yours to carry“You can supply images using the base64, url, or file source types.”
↩︎ Sending images: three source types“The Messages API is stateless, which means that you always send the full conversational history to the API.”
↩︎ Key concept“The Messages API is stateless, which means that you always send the full conversational history to the API.”
↩︎ Exam trap 1 - 2.
“You can include multiple images in a single request, and Claude analyzes them jointly.”
↩︎ Sending images: three source types“Upload the image once, then reference the returned file_id in subsequent messages instead of resending base64 data.”
↩︎ Exam trap 2 - 3.
“This is especially useful for requests with large max_tokens values, where the SDKs require streaming to avoid HTTP timeouts.”
↩︎ Streaming: events, tool input and thinking“your code should handle unknown event types gracefully.”
↩︎ Streaming: events, tool input and thinking“This signature is used to verify the integrity of the thinking block.”
↩︎ Streaming: events, tool input and thinking“the deltas are partial JSON strings, whereas the final tool_use.input is always an object.”
↩︎ Exam trap 3