What you will be able to do
- Explain the tool-use contract: Claude requests a tool call, and your code or Anthropic's servers run it
- Trace the client-side agentic loop keyed on stop_reason, and handle end_turn, pause_turn and the other exit reasons
- Describe how the Claude Agent SDK's query() loop runs turns and which message types it yields
- Decide when a task needs a tool loop and when a plain request is enough
- Explain how sub-agents keep intermediate work out of the parent's context, and define one with the agents option
- Choose a context-window management strategy for long-running agents: compaction, tool result clearing, or the memory tool
Key concept
The agent loop — An agent is a model called in a cycle. Claude decides on an action, a tool runs it, and the result goes back to Claude for its next decision. The cycle ends when Claude replies without requesting any tool.
1.The contract underneath every agent
Every agent pattern in this subdomain rests on one mechanism, tool use. Your application declares which operations exist and what their inputs and outputs look like. Claude decides when to call them and with what arguments. Claude never runs anything itself. It emits a structured request, something else executes it, and the result flows back into the conversation.
What changes from tool to tool is where the code runs, and that decides what your application has to do. There are three cases:
- User-defined tools. You write the schema, run the code and return the result in a tool_result block. Claude sees only your schema and your result, never your implementation.
- Anthropic-schema tools (memory, bash, text_editor, computer, browser). Anthropic publishes the schema and your application runs the operation. Claude has been trained on these exact signatures, so it calls them more reliably than a custom tool that does the same job.
- Server tools (web_search, web_fetch, code_execution, tool_search). Anthropic runs the code, and you never build a tool_result for them.
Sources1
2.The client-side loop, keyed on stop_reason
Claude can't run your code, so every client-executed tool call is a round trip: the model asks, you execute, you report back, the model continues. This round trip is the tool-use loop, the most basic agent pattern:
1. Send a request with your tools array and the user message.
2. Claude responds with stop_reason: "tool_use" and one or more tool_use blocks.
3. Execute each tool and format the outputs as tool_result blocks.
4. Send a new request containing the original messages, Claude's response, and a user message holding the tool_result blocks.
5. Repeat from step 2 while stop_reason is "tool_use".
The loop exits on any other stop reason: "end_turn", "max_tokens", "stop_sequence" or "refusal". Claude has either produced a final answer or stopped for a reason your application has to handle.
Server tools run their own loop inside Anthropic's infrastructure, so a single request can trigger several searches before any response reaches you. That internal loop has an iteration limit. If Claude is still working when it hits the cap, the response comes back with pause_turn instead of end_turn. The work is not done: send the conversation again, including the paused response, and Claude continues where it left off. Control also comes back to you early when Claude calls a server tool and a client tool in the same group of parallel tool calls. In that case the server tool runs only after you return the client tool's results.
| Aspect | Client-executed tools | Server-executed tools |
|---|---|---|
| Examples | User-defined tools; memory, bash, text_editor, computer, browser | web_search, web_fetch, code_execution, tool_search |
| Who runs the operation | Your application | Anthropic's servers |
| Who builds tool_result | You, on the next request | Nobody. The response already contains server_tool_use blocks |
| Stop reason while work remains | tool_use | pause_turn (the internal iteration limit was reached) |
Sources1
3.The same loop, run for you: the Agent SDK's query()
Writing that while loop yourself is one option. The Claude Agent SDK is the other: it runs the same cycle for you. Claude receives your prompt along with the system prompt, tool definitions and conversation history. It evaluates the state and replies with text, tool calls or both. The SDK runs each requested tool and feeds the results back. That repeats until Claude produces a response with no tool calls.
The SDK documentation walks through a failing-test example. In turn 1 Claude runs npm test with Bash and sees three failures. In turn 2 it reads auth.ts and auth.test.ts. In turn 3 it edits auth.ts and runs the tests again, and all three pass. The final turn is text only, with no tool calls. Nobody scripted that sequence in advance: each tool result shaped the next decision.
The SDK hands the loop's progress to you as a stream of typed messages:
- SystemMessage. Session lifecycle events: init at the start, and compact_boundary after compaction.
- AssistantMessage. One per content block, such as a text block or a tool call.
- UserMessage. Tool results being sent back to Claude.
- ResultMessage. Marks the end of the loop and carries the final text, token usage, cost and session ID.
You can also use hooks to intercept, modify or block tool calls before they run.
async for message in query(prompt="Summarize this project"):
if isinstance(message, AssistantMessage):
# Each AssistantMessage carries one content block
for block in message.content:
if isinstance(block, TextBlock):
print(f"Claude: {block.text}")
elif isinstance(block, ToolUseBlock):
print(f"Tool call: {block.name}")
if isinstance(message, ResultMessage):
if message.subtype == "success":
print(message.result)
else:
print(f"Stopped: {message.subtype}")No. A few trailing system events, such as prompt_suggestion, can arrive after the ResultMessage, so iterate the stream to completion. Also check the ResultMessage's subtype: it tells you whether the task succeeded or hit a limit.
An engineering team is building an agent to refactor a large legacy codebase. The exact files needing changes, and how many edits are required, cannot be known until the model inspects the code and encounters each dependency. Which workflow pattern from Anthropic's agent design guidance best fits this task?
Correct answer: A — Deploy an orchestrator LLM that decomposes the refactor into subtasks at runtime and delegates each one to worker LLMs, then synthesizes their combined edits
- A. Correct — orchestrator-workers fits exactly this unpredictable, dependency-driven decomposition where subtasks emerge as the model inspects the code.
- B. Wrong — prompt chaining requires the subtasks to be known and fixed in advance, which conflicts with the scenario's unpredictable dependencies.
- C. Wrong — routing assumes categories are known upfront and doesn't handle a task where subtasks emerge dynamically during execution.
- D. Wrong — parallel voting works for independent attempts at the same task, not for decomposing an unpredictable set of dependent edits.
Sources2
4.When a tool loop is the right shape
A loop adds cost, so use one only when the task needs something the model can't do from text alone:
- actions with side effects, such as sending an email or writing a file - fresh or external data - output with a guaranteed structure - calls into existing systems such as databases and internal APIs
A practical test from the documentation: if you are writing a regex to pull a decision out of model output, that decision should have been a tool call.
The opposite also holds. Summarization, translation and general-knowledge questions can be answered from training alone. One-shot Q&A with no side effects gives a tool nothing to do. For a trivial response, the extra round trip can cost more than the work itself.
Sources1
5.Sub-agents and context-window management
A long-running agent has two linked problems: the loop keeps adding to the context window, and one agent may have to do several kinds of work. Sub-agents and context management address both.
Sub-agents. A subagent is a separate agent that the main agent delegates a task to. In the Agent SDK you define one programmatically with the agents parameter in your query() options, as a markdown file in .claude/agents/, or you rely on the built-in general-purpose subagent, which Claude can invoke through the Agent tool. A programmatic definition needs a description, which tells Claude when to use the subagent, and a prompt, which is its system prompt. It can also set tools to restrict what the subagent can do and model to override the model. The Agent tool must be available to the parent, which is why the SDK example lists it in allowedTools. Sub-agents give you four things:
- Context isolation. Each subagent runs in its own conversation. Its intermediate tool calls and results stay inside it, and only its final message returns to the parent, so a research subagent can read dozens of files without filling the main conversation. - Parallelization. Independent subagents, such as a style checker, a security scanner and a test-coverage checker, can run concurrently. - Specialized instructions. Each subagent can have a tailored system prompt. - Tool restrictions. A doc-reviewer subagent limited to Read and Grep can analyze files but never modify them.
Context-window management. Everything in the request counts toward the window: the system prompt, every message including tool results, and your tool definitions, plus the output Claude generates. More context is not automatically better, because accuracy and recall degrade as the token count grows (context rot), so curating what is in context matters as much as how much space is left. The main tools are:
- Server-side compaction. The primary strategy for long-running conversations and agentic workflows. It summarizes earlier parts of the conversation on the server so the conversation can continue past the window limit. In the Agent SDK, a compact_boundary SystemMessage fires after compaction.
- Context editing. Tool result clearing removes old tool results in agentic workflows, and thinking block clearing manages thinking blocks.
- Context awareness. Some models track their remaining token budget during a conversation, so they can pace long-running tasks against the space that remains.
- The memory tool. Claude stores what it learns in files and reads them back on demand instead of loading everything up front. This keeps the active context focused on the current task and lets work carry across sessions. It is client-side, so your application executes each memory operation.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.When Claude calls a tool, Claude runs that tool.Why is that wrong?
Claude only emits a structured tool_use request. Your application runs client tools, and Anthropic's servers run server tools. The result then goes back into the conversation.
Covered in The contract underneath every agent
2.Any stop_reason other than tool_use means Claude has finished the task.Why is that wrong?
With server tools, pause_turn means the internal loop hit its iteration limit and the work is unfinished. Send the conversation again so Claude can continue.
Covered in The client-side loop, keyed on stop_reason
3.The ResultMessage is the last item the SDK stream yields, so you can break as soon as it arrives.Why is that wrong?
A few trailing system events can arrive after the ResultMessage, so read the stream to the end.
Covered in The same loop, run for you: the Agent SDK's query()
4.Every tool call and file a subagent reads is added to the parent agent's context window.Why is that wrong?
A subagent runs in its own conversation. Its intermediate tool calls and results stay inside it, and only its final message returns to the parent.
Covered in Sub-agents and context-window management
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The model never executes anything on its own.”
↩︎ The contract underneath every agent“these schemas are trained-in.”
↩︎ The contract underneath every agent“Client-executed tools (both user-defined and Anthropic-schema) require your application to drive a loop.”
↩︎ The client-side loop, keyed on stop_reason“In practice this reads as: while stop_reason == "tool_use", execute the tools and continue the conversation.”
↩︎ The client-side loop, keyed on stop_reason“Every tool call is at least one extra round trip; for lightweight tasks the overhead can exceed the work.”
↩︎ When a tool loop is the right shape“The model never executes anything on its own.”
↩︎ Exam trap 1“If the model is still iterating when it hits the cap, the response comes back with stop_reason: "pause_turn" instead of "end_turn".”
↩︎ Exam trap 2 - 2.https://code.claude.com/docs/en/agent-sdk/agent-loopOfficial docs
“Each full cycle is one turn. Claude continues calling tools and processing results until it produces a response with no tool calls.”
↩︎ The same loop, run for you: the Agent SDK's query()“You can use hooks to intercept, modify, or block tool calls before they run.”
↩︎ The same loop, run for you: the Agent SDK's query()“ResultMessage: marks the end of the agent loop. Contains the final text result, token usage, cost, and session ID.”
↩︎ The same loop, run for you: the Agent SDK's query()“Each full cycle is one turn. Claude continues calling tools and processing results until it produces a response with no tool calls.”
↩︎ Key concept“so iterate the stream to completion rather than breaking on the result.”
↩︎ Exam trap 3 - 3.https://code.claude.com/docs/en/agent-sdk/subagentsOfficial docs
“intermediate tool calls and results stay inside the subagent; only its final message returns to the parent.”
↩︎ Sub-agents and context-window management“multiple subagents can run concurrently, so independent subtasks finish in the time of the slowest one”
↩︎ Sub-agents and context-window management“intermediate tool calls and results stay inside the subagent; only its final message returns to the parent.”
↩︎ Exam trap 4 - 4.
“As token count grows, accuracy and recall degrade, a phenomenon known as context rot.”
↩︎ Sub-agents and context-window management“Compaction automatically summarizes earlier parts of the conversation on the server, so the conversation can continue past the context window limit.”
↩︎ Sub-agents and context-window management - 5.
“Rather than loading all relevant information up front, an agent records what it learns in memory files and reads them back on demand.”
↩︎ Sub-agents and context-window management