What you will be able to do
- Explain what a subagent provides: context isolation, parallel work, specialized instructions and tool restrictions
- Define subagents programmatically with AgentDefinition and choose tools, model and maxTurns
- List what counts toward the context window and pick a strategy for long-running agents
- Use the memory tool to keep knowledge across sessions without loading all of it into context
1.Subagents: splitting work across separate conversations
A single agent loop builds up everything it touches: every file read and every tool result. Subagents are the orchestrator-and-worker answer to that. The main agent hands a subtask to a subagent, and the subagent runs in a conversation of its own. The Agent SDK documentation lists four benefits:
- Context isolation. The subagent's intermediate tool calls and results stay inside it. A research subagent can explore dozens of files, and the parent receives a short summary rather than every file it read. - Parallelization. Several subagents can run at once, so independent subtasks finish in the time of the slowest one instead of the sum. A code review can run style-checker, security-scanner and test-coverage subagents at the same time. - Specialized instructions. Each subagent can have its own system prompt. A database-migration subagent can carry detailed SQL and rollback guidance that would only be noise in the main agent's instructions. - Tool restrictions. A doc-reviewer subagent limited to Read and Grep can analyze your files but never change them.
There are three ways to get a subagent: pass the agents parameter in your query() options, write markdown files in .claude/agents/, or rely on the built-in general-purpose subagent. Claude can call the built-in one through the Agent tool without you defining anything.
Sources1
2.Defining subagents with AgentDefinition
In the programmatic form, each entry in agents maps a name to a definition. Two things in the example below are easy to miss. First, Agent appears in allowed_tools: that is the tool Claude uses to delegate. Second, the description is what Claude reads when deciding whether to use a subagent, so write it as guidance on when to call it.
allowed_tools=["Read", "Grep", "Glob", "Agent"],
agents={
"code-reviewer": AgentDefinition(
# description tells Claude when to use this subagent
description="Expert code review specialist. Use for quality, security, and maintainability reviews.",
# prompt defines the subagent's behavior and expertise
prompt="""You are a code review specialist with expertise in security, performance, and best practices.
When reviewing code:
- Identify security vulnerabilities
- Check for performance issues
- Verify adherence to coding standards
- Suggest specific improvements
Be thorough but concise in your feedback.""",
# tools restricts what the subagent can do (read-only here)
tools=["Read", "Grep", "Glob"],
# model overrides the default model for this subagent
model="sonnet",| Field | Required | What it controls |
|---|---|---|
| description | Yes | When Claude should use this agent |
| prompt | Yes | The subagent's system prompt: its role and behavior |
| tools | No | The tools it may use. If omitted, it gets every tool available to subagents |
| model | No | A model override, such as 'sonnet', 'haiku' or 'inherit' |
| maxTurns | No | The turn limit. When reached, output is returned marked as partial, and the agent can be resumed |
| background | No | Runs the agent as a non-blocking background task |
A content team wants an agent that first drafts a product announcement in English, then, only once the draft passes a length and tone check, translates it into French, German, and Japanese. Each step should read only the previous step's validated output. Which workflow pattern matches this design?
Correct answer: B — Prompt chaining, where the draft and each translation are separate LLM calls in a fixed sequence with a programmatic gate verifying the draft before translation begins
- A. Wrong — the languages here are fixed in advance, not dynamically decided by an orchestrator based on unpredictable requirements.
- B. Correct — this is prompt chaining: fixed sequential steps with a gate between the draft and translation stages.
- C. Wrong — routing selects one path among alternatives; here all three translations run in the same predetermined sequence.
- D. Wrong — no evaluator LLM iteratively critiques the translations in this design; the gate is a one-time check, not a feedback loop.
Sources1
3.Context-window management
Subagents are one way to manage context. It helps to know exactly what the context window holds. It is the model's working memory for a request: everything the model can refer to while it generates a response, including the response itself. Everything in the request counts toward it: the system prompt, every message (tool results, images and documents included), your tool definitions, and the output Claude generates, extended thinking included.
A bigger window is not automatically better. As the token count grows, accuracy and recall degrade, which is called context rot. Choosing what goes into context matters as much as how much space you have.
Some models (Claude Sonnet 5, Sonnet 4.6, Sonnet 4.5 and Haiku 4.5) have context awareness. The API tells them their total budget and, after each tool call, how much capacity remains, so they can plan long tasks against the space left. You don't have to enable anything.
| Strategy | What it does |
|---|---|
| Server-side compaction | Summarizes earlier parts of the conversation on the server so the conversation can continue past the limit. It is the primary strategy for long-running agentic work |
| Tool result clearing | Clears old tool results in agentic workflows |
| Thinking block clearing | Manages thinking blocks when extended thinking is used |
| Tool search tool | Defers tool definitions to reduce the context they consume |
Sources2
4.Memory: knowledge that lasts across sessions
Compaction keeps one conversation going. The memory tool does something different: it lets Claude keep knowledge across separate conversations, in a directory of files that Claude can create, read, update and delete. This is just-in-time retrieval. The agent writes down what it learns and reads it back when it needs it, rather than loading everything at the start. That keeps the active context focused on the current task.
When the memory tool is enabled, Claude checks its memory directory before starting a task. As it works, it saves what it learns in files under /memories and reads them in later conversations to continue the work.
The memory tool runs on the client side: Claude requests file operations, and your application carries them out against storage you control. The /memories path is a prefix that your handler maps onto real storage, such as a per-user directory or database keys. For security, your handler must reject any path outside /memories. Configuration is a single entry in tools:
message = client.messages.create(
model="claude-opus-5-5",
max_tokens=2048,
messages=[
{
"role": "user",
"content": "Help me respond to this customer service ticket.",
}
],
tools=[{"type": "memory_20250818", "name": "memory"}],
)The request must send the same memory tools entry, and your handler must serve the same store. Memory lives entirely in your application, not on Anthropic's side.
A support platform wants incoming tickets first classified as billing, technical, or account-access, then handled by a prompt written specifically for that category, since the resolution steps differ substantially between them. Which pattern should the team implement?
Correct answer: C — Routing, where a classifier step directs each ticket to one of three specialized downstream prompts built for its category
- A. Wrong — sectioning parallelizes independent subtasks of the same input; it doesn't select a single specialized path based on category.
- B. Wrong — the categories here are fixed and known in advance, not dynamically invented by an orchestrator.
- C. Correct — routing separates concerns by classifying input and sending it to a specialized prompt for that category.
- D. Wrong — chaining reuses sequential steps on the same output; it doesn't select a single specialized prompt based on classification.
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The subagent's full transcript, including every file it read, is merged back into the parent's conversation.Why is that wrong?
Intermediate tool calls and results stay inside the subagent. Only its final message returns to the parent, which is what keeps the parent's context small.
Covered in Subagents: splitting work across separate conversations
2.Prompt caching frees up room in the context window.Why is that wrong?
Caching changes the price of those tokens, not whether they count. Cached prefixes still take up context space.
Covered in Context-window management
3.The memory tool stores Claude's memories on Anthropic's servers.Why is that wrong?
The memory tool runs on the client side. Claude requests operations, and your application runs them against storage you control.
Covered in Memory: knowledge that lasts across sessions
Practise it for real
Give Claude memory that lasts across two separate conversations by using the SDK's local-filesystem memory helper.
1.In Python, create
memory = BetaLocalFilesystemMemoryTool(base_path="./memory")and pass it intoolstoclient.beta.messages.tool_runnerwith modelclaude-opus-5-5and the message "Remember that customer Acme Corp prefers email follow-ups."Why: The helper handles the memory tool interface and the tool-use loop, and stores memories as files on disk.
You should see:
runner.until_done()returns a final message after Claude has made one or more memory tool calls.2.List the contents of
./memoryon disk.Why: It confirms the memory lives in storage you control, not in the conversation.
You should see: One or more files that Claude wrote, recording the Acme Corp preference.
3.Start a new tool_runner conversation with the same memory tool and ask how Acme Corp prefers to be contacted.
Why: A later conversation carries on from the same memory when it sends the same tools entry and the handler serves the same store.
You should see: Claude checks its memory directory first and answers from the saved file.
Stuck? Get a nudge
If the second conversation doesn't find the memory, check that both runs use the same base_path.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://code.claude.com/docs/en/agent-sdk/subagentsOfficial docs
“only its final message returns to the parent.”
↩︎ Subagents: splitting work across separate conversations“independent subtasks finish in the time of the slowest one rather than the sum of all of them.”
↩︎ Subagents: splitting work across separate conversations“subagents can be limited to specific tools, reducing the risk of unintended actions.”
↩︎ Subagents: splitting work across separate conversations“description tells Claude when to use this subagent”
↩︎ Defining subagents with AgentDefinition“When the agent reaches the limit, Claude Code returns its output marked as partial, and you can resume the agent to continue.”
↩︎ Defining subagents with AgentDefinition“only its final message returns to the parent.”
↩︎ Exam trap 1 - 2.
“As token count grows, accuracy and recall degrade, a phenomenon known as context rot.”
↩︎ Context-window management“For long-running conversations and agentic workflows, server-side compaction is the primary strategy for context management.”
↩︎ Context-window management“Compaction automatically summarizes earlier parts of the conversation on the server, so the conversation can continue past the context window limit.”
↩︎ Context-window management“Cached prompt prefixes still occupy the context window”
↩︎ Exam trap 2 - 3.
“Memory supports just-in-time context retrieval.”
↩︎ Memory: knowledge that lasts across sessions“When the memory tool is enabled, Claude automatically checks its memory directory before starting a task.”
↩︎ Memory: knowledge that lasts across sessions“Your handler must reject paths outside /memories”
↩︎ Memory: knowledge that lasts across sessions“The memory tool operates client-side: Claude requests file operations, and your application executes them.”
↩︎ Exam trap 3