What you will be able to do
- Turn a business interaction into the discrete tasks a Claude solution must perform and be evaluated on
- Trace one request through input, context assembly, model call, tool execution and the feedback of results
- Explain who executes each kind of tool and what your application must do in each case
- Identify when the loop has really finished, and where human approval belongs before an output takes effect
Key concept
The agentic loop — The feedback cycle at the centre of every end-to-end Claude design. Claude gets the input and its context and may request tool calls. Something outside the model runs those calls, and the results go back to Claude. The cycle repeats until Claude replies without asking for a tool.
1.Start from the interaction, not the model
An end-to-end architecture starts before any API call, with a picture of what the user actually does. Anthropic's customer support guide asks you to write out an ideal interaction first, because that outline decides the technical requirements. In its car insurance example, a customer is greeted, asks about cover for an electric car, drifts off topic, asks for a quote, and finally asks follow-up questions.
That single conversation is really several jobs: greeting and general guidance, product information, conversation management (staying on topic and redirecting), and quote generation. The quote step shows the full architecture in small form. Claude collects the facts through questions, sends a request to a quote generation API tool, gets the result back, and turns it into a natural answer. So the input is the customer's message, the processing is Claude plus a tool, and the output is a quote that Claude has written up from the tool's response.
Breaking the interaction into tasks has a second purpose: each task becomes something you can prompt for and test. The guide pairs this breakdown with success criteria and measurable benchmarks set with the support team. This is the first feedback loop in the design, and it runs outside the model. Evaluations show whether each task does its job.
Sources1
2.Processing: context in, tool calls out, results back
Once the tasks are known, the processing stage is the same for all of them. The Agent SDK describes it step by step. First, Claude receives the prompt along with the system prompt, tool definitions and conversation history; together these are the assembled context. Next, Claude evaluates the current state and replies with text, one or more tool call requests, or both. Then the tools run and their results go back to Claude. These steps repeat, and each full cycle counts as one turn.
The key architectural fact is that Claude never executes a tool itself. Tool use is a contract: you declare which operations exist and what shape they take, and Claude decides when to call them. Claude only ever sees the schema and the result, never your implementation. Where the code runs decides which parts of the loop your application owns.
| Tool category | Examples | Who executes | Your application's job |
|---|---|---|---|
| User-defined | A database query, an HTTP call, a file write | Your code | Write the schema, run the operation, return a tool_result |
| Anthropic-schema | memory, bash, text_editor, computer, browser | Your code | Run the operation against a trained-in schema and return a tool_result |
| Server-executed | web_search, web_fetch, code_execution, tool_search | Anthropic's servers | Enable the tool and read the final answer; never build a tool_result |
For client-executed tools, each call is a round trip: the model asks, you execute, you report back, and the model continues. Your application runs a while loop keyed on stop_reason. Server-executed tools run their own loop inside Anthropic's infrastructure, so a single request may trigger several searches before any response reaches you.
The guidance says that if you are writing a regex to pull a decision out of model output, that decision should have been a tool call. Structured intent belongs in a tool schema, which enforces its shape, rather than in prose you parse afterwards.
A team is designing the processing stage of a customer-support architecture. Support tickets arrive continuously and must be triaged with strong reasoning about edge cases, but the team has a strict per-ticket cost ceiling and wants to avoid over-provisioning intelligence. Which approach best fits the input-to-processing handoff for this architecture?
Correct answer: A — Start triage on Claude Haiku 4.5, benchmark it against real tickets, and only route to a stronger model when a capability gap appears
- A. Correct. This mirrors Anthropic's documented model-selection approach: start with a fast, cost-effective model, benchmark against real use-case data, and upgrade only when a concrete capability gap is found, which fits a cost-constrained triage stage.
- B. Incorrect. Routing everything through the most capable, highest-effort model ignores the stated cost ceiling and is not the recommended starting point for high-volume, straightforward triage.
- C. Incorrect. Random alternation is not a model-selection strategy Anthropic documents; it introduces unpredictable accuracy without any evaluation basis.
- D. Incorrect. Replacing Claude entirely with an external classifier abandons the reasoning capability the scenario requires for edge cases and isn't a supported architecture pattern for this trade-off.
3.Output: knowing when the loop is done, and what it may do
The loop ends when Claude replies with no tool calls. The SDK then produces a final message and a result holding the final text, token usage, cost and session ID. Those figures are what you use to check the architecture against its cost budget. On the raw API, the stop_reason field tells you why the response ended, and not every exit means the work is done.
| stop_reason | Meaning for the architecture |
|---|---|
| tool_use | Execute the requested tools, send back tool_result blocks, loop again |
| end_turn | Claude has produced its final answer |
| max_tokens, stop_sequence, refusal | Loop exits; Claude stopped for a reason your application should handle |
| pause_turn | The server-side loop hit its iteration cap; the work is not finished, so re-send the conversation to continue |
An output that changes real systems needs one more design decision: whether it applies straight away. Anthropic's commerce agent guide shows the careful version. The storefront agent states prices and availability only from tool results in the conversation. Every write the merchant agent proposes becomes a staged change with a server-generated ID, shown to the operator as a preview card. Guardrails are checked when the change is staged and again when it is applied. The change takes effect only after a person approves it outside the conversation, through a portal button, a console prompt, or an always-ask permission policy.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Once tools are defined, Claude runs them itself, so the application only has to send the prompt and read the reply.Why is that wrong?
Claude only produces a structured request. For client tools, your code runs the operation and returns a tool_result; for server tools, Anthropic's servers run it.
Covered in Processing: context in, tool calls out, results back
2.Any response that is not stop_reason tool_use means Claude has finished the task.Why is that wrong?
pause_turn means the server-side loop hit its iteration limit mid-task. You must re-send the conversation, including the paused response, so the model can carry on.
Covered in Output: knowing when the loop is done, and what it may do
3.If the operator types approval in the chat, the agent may apply the staged price change.Why is that wrong?
In the commerce reference design, changes apply only after approval outside the conversation, and guardrails are checked again at apply time.
Covered in Output: knowing when the loop is done, and what it may do
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Outline an ideal customer interaction to define how and when you expect the customer to interact with Claude.”
↩︎ Start from the interaction, not the model“break down your ideal customer interaction into every task you want Claude to be able to perform”
↩︎ Start from the interaction, not the model“Work with your support team to define success criteria and write detailed evaluations with measurable benchmarks and goals.”
↩︎ Start from the interaction, not the model - 2.https://code.claude.com/docs/en/agent-sdk/agent-loopOfficial docs
“Claude receives your prompt, along with the system prompt, tool definitions, and conversation history.”
↩︎ Processing: context in, tool calls out, results back“Claude continues calling tools and processing results until it produces a response with no tool calls.”
↩︎ Output: knowing when the loop is done, and what it may do“a ResultMessage with the final text, token usage, cost, and session ID.”
↩︎ Output: knowing when the loop is done, and what it may do“Each set of tool results feeds back to Claude for the next decision.”
↩︎ Key concept - 3.
“It emits a structured request, your code (or Anthropic's servers) runs the operation, and the result flows back into the conversation.”
↩︎ Processing: context in, tool calls out, results back“Claude never sees your implementation; it only sees the schema you provided and the result you returned.”
↩︎ Processing: context in, tool calls out, results back“Client-executed tools (both user-defined and Anthropic-schema) require your application to drive a loop.”
↩︎ Processing: context in, tool calls out, results back“Server-executed tools run their own loop inside Anthropic's infrastructure.”
↩︎ Processing: context in, tool calls out, results back“if you're writing a regex to extract a decision from model output, that decision should have been a tool call.”
↩︎ Processing: context in, tool calls out, results back“The model never executes anything on its own.”
↩︎ Exam trap 1“A paused turn means the work isn't finished; re-send the conversation (including the paused response) to let the model continue where it left off.”
↩︎ Exam trap 2 - 4.
“The agent states products, prices, availability, and store terms only from tool results in the conversation”
↩︎ Output: knowing when the loop is done, and what it may do“are checked when the change is staged and again when it is applied”
↩︎ Output: knowing when the loop is done, and what it may do“The change applies only after a person approves it outside the conversation”
↩︎ Output: knowing when the loop is done, and what it may do“An approval typed in chat approves nothing.”
↩︎ Exam trap 3