What you will be able to do
- Explain the two ways a large, always-loaded toolset degrades an agent: context consumption and tool selection accuracy
- Recognise the tool-count thresholds the documentation gives, and say which page each one comes from
- Configure tool search with defer_loading so that tool definitions load on demand
- Choose the right context-management approach for where an agent's tokens are actually going
Key concept
Capability bloat — Giving an agent more tool definitions than a request needs, all loaded up front. Each extra definition takes context and gives the model another wrong choice, so past a certain size the agent gets worse at picking tools, not better.
1.The two costs of loading every tool
It seems safe to give an agent every tool it might possibly need. The documentation names two separate costs of doing that, and it helps to keep them apart, because each has its own fix.
The first cost is context. Every tool definition, meaning its name, description and input schema, sits in the context window before the agent does anything. The Agent SDK documentation puts a figure on it: 50 tools can use 10-20K tokens, leaving less room for actual work. The API documentation gives a realistic example: a setup connected to several MCP servers (GitHub, Slack, Sentry, Grafana and Splunk) can consume about 55k tokens in definitions before Claude does any work. On a long-running agent, those tokens are gone before the first tool_result arrives.
The second cost is selection accuracy, and adding context does not fix it. The more tools there are, the harder the model finds it to choose the right one. The documentation places the degradation at 30 to 50 loaded tools. A bigger context window gives you room for the definitions, but it does not make the choice between 60 similar-looking tools any easier.
A developer wants to build a code-review agent that must never edit files or run shell commands, even if Claude decides such actions would help, and must not fall back to interactive approval because it runs headless in CI. Which configuration achieves this?
Correct answer: A — Set allowedTools: ["Read", "Glob", "Grep"] with permissionMode: "dontAsk" so listed tools are approved and everything else is denied outright.
- A. Correct. In dontAsk mode, tools pre-approved by allowedTools run, and anything not pre-approved is denied outright without calling canUseTool, giving a fixed, headless-safe tool surface.
- B. Incorrect. This does block Edit, Write, and Bash, but bypassPermissions approves every other tool, including any MCP or future tools, which is broader than the intended read-only scope and not headless-safe by design.
- C. Incorrect. In default mode, unlisted tools fall through to the canUseTool callback, which requires interactive confirmation, defeating the headless requirement.
- D. Incorrect. acceptEdits auto-approves file edits and filesystem commands like rm and mv, which is the opposite of preventing edits, and Bash still prompting does not satisfy a fixed, non-interactive tool surface.
2.Tool search: load definitions on demand
The documented fix for both costs is tool search. You don't put every definition into the model's context up front. Instead, the model starts with a search tool plus a small set of tools it always needs, and it discovers the rest when a request calls for them. According to the documentation, this typically cuts definition tokens by over 85 percent, loading only the few tools a given request needs. Because the model only ever chooses among a focused set, selection accuracy stays high even when the catalog runs to thousands of tools.
There are two variants. The regex variant (tool_search_tool_regex_20251119) has Claude write regex patterns. The BM25 variant (tool_search_tool_bm25_20251119) uses natural-language queries. Both search tool names, descriptions, argument names and argument descriptions. In practice, that means your descriptions are the text discovery matches against. A vague description makes a tool harder to find as well as harder to choose.
To mark a tool for on-demand loading, add defer_loading: true to its definition:
{
"name": "get_weather",
"description": "Get current weather for a location",
"input_schema": {
"type": "object",
"properties": {
"location": { "type": "string" },
"unit": { "type": "string", "enum": ["celsius", "fahrenheit"] }
},
"required": ["location"]
},
"defer_loading": true
}Three configuration rules decide whether this works well:
- You still send every definition. defer_loading controls what goes into the model's context, not what you send in the request. The API needs the full definitions server-side so it can run the search and expand the matches.
- Never defer the search tool itself. At least one tool has to stay non-deferred, and that is normally the search tool.
- Keep your 3–5 most frequently used tools non-deferred, so Claude can call them without searching first.
When Claude searches, the API returns the matching tools as tool_reference blocks (up to 5 by default) and expands them into full definitions. Deferred tools are left out of the system-prompt prefix, so discovering one doesn't break prompt caching.
A lead agent has allowedTools: ["Read", "Write", "Edit", "Bash", "Agent"] for its own broad development work. The team also defines a code-reviewer subagent that should only ever read and search code, never write files or run shell commands, regardless of what the lead agent is permitted to do. How should the subagent be configured to enforce this?
Correct answer: A — Define the subagent with its own tools: ["Read", "Glob", "Grep"] field in its AgentDefinition, independent of the lead agent's allowedTools.
- A. Correct. Custom subagents define their own tools field, which scopes exactly what that subagent can call, independent of the lead agent's broader allowedTools.
- B. Incorrect. Subagents do not automatically inherit a subset of the lead agent's tool list; without an explicit tools field, a subagent's capability is defined by its own configuration.
- C. Incorrect. CLAUDE.md is contextual guidance, not a tool permission mechanism, and subagents do not inherit tool permissions purely from prose instructions.
- D. Incorrect. permissionMode is a session-level setting; switching it on the lead agent affects approval behavior generally, it does not selectively restrict only the subagent's tool access.
Sources2
3.Match the fix to where the tokens go
Tool search is one of four documented ways to manage tool context. Bloat from definitions and bloat from results are different problems, and an exam scenario will often describe one of them while offering fixes for the other. The table shows which approach addresses which pressure.
| Approach | What it reduces | When it fits |
|---|---|---|
| Tool search | Tool definitions loaded upfront | Large toolsets (20+ tools) where most tools aren't needed every turn |
| Programmatic tool calling | tool_result roundtrips | Chains of tool calls that can execute as a single script |
| Prompt caching | Token cost of repeated tool definitions | Stable toolsets across many requests |
| Context editing | Old tool_result blocks in history | Long conversations where early results are no longer relevant |
Tool search is not free. It costs a small amount of latency, one extra turn to look up a tool, in exchange for a large reduction in baseline context. Prompt caching lowers the *cost* of repeated definitions, but the definitions are all still in context and the model still chooses among all of them. So caching helps your bill and does nothing for selection accuracy.
In the Agent SDK, tool search is controlled by the ENABLE_TOOL_SEARCH environment variable. If you leave it unset, tool search is on and definitions are deferred. auto counts the deferrable definition tokens and turns tool search on only once they reach 10% of the context window. auto:N does the same with a custom percentage, and false loads every definition on every turn.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If a tool has defer_loading: true, you can leave its full definition out of the request to save tokens.Why is that wrong?
Deferral decides what enters the model's context, not what you send. Every definition still goes in the tools array, because the API uses them server-side to run the search and expand tool_reference blocks.
Covered in Tool search: load definitions on demand
2.For the biggest saving, defer every tool, including the tool search tool.Why is that wrong?
The search tool must stay non-deferred, or Claude has no way to discover anything. It's also best to keep your most-used tools loaded so they can be called without a search first.
Covered in Tool search: load definitions on demand
3.In the Agent SDK, leaving ENABLE_TOOL_SEARCH unset behaves like auto, so tool search only starts once definitions reach 10% of the window.Why is that wrong?
Unset means tool search is on and definitions are deferred. The 10% threshold applies only when you explicitly choose auto, and auto:N lets you set a different percentage.
Covered in Match the fix to where the tokens go
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Tool definitions can consume large portions of the context window (50 tools can use 10-20K tokens), leaving less room for actual work.”
↩︎ The two costs of loading every tool“When the total reaches 10% of the window, tool search activates.”
↩︎ Match the fix to where the tokens go“Tool selection accuracy degrades with more than 30-50 tools loaded at once.”
↩︎ Key concept“Tool search is on. Tool definitions are deferred and discovered on demand.”
↩︎ Exam trap 3 - 2.
“A typical multiserver setup (GitHub, Slack, Sentry, Grafana, and Splunk) can consume ~55k tokens in definitions before Claude does any work.”
↩︎ The two costs of loading every tool“Both tool search variants (regex and bm25) search tool names, descriptions, argument names, and argument descriptions.”
↩︎ Tool search: load definitions on demand“Keep your 3–5 most frequently used tools non-deferred so Claude can call them without searching first.”
↩︎ Tool search: load definitions on demand“The prefix is untouched, so prompt caching is preserved.”
↩︎ Tool search: load definitions on demand“defer_loading controls what enters the context window, not what you send in the request”
↩︎ Exam trap 1“Never set defer_loading: true on the tool search tool itself.”
↩︎ Exam trap 2 - 3.
“This trades a small amount of latency (one extra turn to look up a tool) for a large reduction in baseline context usage.”
↩︎ Match the fix to where the tokens go