CertSafari
    CCAR-P · Lessons

    Domain 3 · Lesson 12/38

    Capability Bloat: Why Too Many Tools Hurt Agents

    Evaluate tool/agent configuration for capability bloat

    8 min read
    2.38% of exam
    3 sources
    Published 27 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Explain the two ways a large, always-loaded toolset degrades an agent: context consumption and tool selection accuracy
    • Recognise the tool-count thresholds the documentation gives, and say which page each one comes from
    • Configure tool search with defer_loading so that tool definitions load on demand
    • Choose the right context-management approach for where an agent's tokens are actually going

    Key concept

    Capability bloat — Giving an agent more tool definitions than a request needs, all loaded up front. Each extra definition takes context and gives the model another wrong choice, so past a certain size the agent gets worse at picking tools, not better.

    1.The two costs of loading every tool

    It seems safe to give an agent every tool it might possibly need. The documentation names two separate costs of doing that, and it helps to keep them apart, because each has its own fix.

    The first cost is context. Every tool definition, meaning its name, description and input schema, sits in the context window before the agent does anything. The Agent SDK documentation puts a figure on it: 50 tools can use 10-20K tokens, leaving less room for actual work. The API documentation gives a realistic example: a setup connected to several MCP servers (GitHub, Slack, Sentry, Grafana and Splunk) can consume about 55k tokens in definitions before Claude does any work. On a long-running agent, those tokens are gone before the first tool_result arrives.

    The second cost is selection accuracy, and adding context does not fix it. The more tools there are, the harder the model finds it to choose the right one. The documentation places the degradation at 30 to 50 loaded tools. A bigger context window gives you room for the definitions, but it does not make the choice between 60 similar-looking tools any easier.

    A developer wants to build a code-review agent that must never edit files or run shell commands, even if Claude decides such actions would help, and must not fall back to interactive approval because it runs headless in CI. Which configuration achieves this?

    Sources12

    2.Tool search: load definitions on demand

    The documented fix for both costs is tool search. You don't put every definition into the model's context up front. Instead, the model starts with a search tool plus a small set of tools it always needs, and it discovers the rest when a request calls for them. According to the documentation, this typically cuts definition tokens by over 85 percent, loading only the few tools a given request needs. Because the model only ever chooses among a focused set, selection accuracy stays high even when the catalog runs to thousands of tools.

    There are two variants. The regex variant (tool_search_tool_regex_20251119) has Claude write regex patterns. The BM25 variant (tool_search_tool_bm25_20251119) uses natural-language queries. Both search tool names, descriptions, argument names and argument descriptions. In practice, that means your descriptions are the text discovery matches against. A vague description makes a tool harder to find as well as harder to choose.

    To mark a tool for on-demand loading, add defer_loading: true to its definition:

    A tool definition marked for on-demand loading with defer_loadingjson
    {
      "name": "get_weather",
      "description": "Get current weather for a location",
      "input_schema": {
        "type": "object",
        "properties": {
          "location": { "type": "string" },
          "unit": { "type": "string", "enum": ["celsius", "fahrenheit"] }
        },
        "required": ["location"]
      },
      "defer_loading": true
    }

    Three configuration rules decide whether this works well:

    - You still send every definition. defer_loading controls what goes into the model's context, not what you send in the request. The API needs the full definitions server-side so it can run the search and expand the matches. - Never defer the search tool itself. At least one tool has to stay non-deferred, and that is normally the search tool. - Keep your 3–5 most frequently used tools non-deferred, so Claude can call them without searching first.

    When Claude searches, the API returns the matching tools as tool_reference blocks (up to 5 by default) and expands them into full definitions. Deferred tools are left out of the system-prompt prefix, so discovering one doesn't break prompt caching.

    A lead agent has allowedTools: ["Read", "Write", "Edit", "Bash", "Agent"] for its own broad development work. The team also defines a code-reviewer subagent that should only ever read and search code, never write files or run shell commands, regardless of what the lead agent is permitted to do. How should the subagent be configured to enforce this?

    Sources2

    3.Match the fix to where the tokens go

    Tool search is one of four documented ways to manage tool context. Bloat from definitions and bloat from results are different problems, and an exam scenario will often describe one of them while offering fixes for the other. The table shows which approach addresses which pressure.

    Four approaches to tool-context pressure, and which source of tokens each one reduces
    ApproachWhat it reducesWhen it fits
    Tool searchTool definitions loaded upfrontLarge toolsets (20+ tools) where most tools aren't needed every turn
    Programmatic tool callingtool_result roundtripsChains of tool calls that can execute as a single script
    Prompt cachingToken cost of repeated tool definitionsStable toolsets across many requests
    Context editingOld tool_result blocks in historyLong conversations where early results are no longer relevant

    Tool search is not free. It costs a small amount of latency, one extra turn to look up a tool, in exchange for a large reduction in baseline context. Prompt caching lowers the *cost* of repeated definitions, but the definitions are all still in context and the model still chooses among all of them. So caching helps your bill and does nothing for selection accuracy.

    In the Agent SDK, tool search is controlled by the ENABLE_TOOL_SEARCH environment variable. If you leave it unset, tool search is on and definitions are deferred. auto counts the deferrable definition tokens and turns tool search on only once they reach 10% of the context window. auto:N does the same with a custom percentage, and false loads every definition on every turn.

    Sources31

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.If a tool has defer_loading: true, you can leave its full definition out of the request to save tokens.Why is that wrong?

      Deferral decides what enters the model's context, not what you send. Every definition still goes in the tools array, because the API uses them server-side to run the search and expand tool_reference blocks.

      Covered in Tool search: load definitions on demand

    2. 2.For the biggest saving, defer every tool, including the tool search tool.Why is that wrong?

      The search tool must stay non-deferred, or Claude has no way to discover anything. It's also best to keep your most-used tools loaded so they can be called without a search first.

      Covered in Tool search: load definitions on demand

    3. 3.In the Agent SDK, leaving ENABLE_TOOL_SEARCH unset behaves like auto, so tool search only starts once definitions reach 10% of the window.Why is that wrong?

      Unset means tool search is on and definitions are deferred. The 10% threshold applies only when you explicitly choose auto, and auto:N lets you set a different percentage.

      Covered in Match the fix to where the tokens go

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Tool definitions can consume large portions of the context window (50 tools can use 10-20K tokens), leaving less room for actual work.”
      ↩︎ The two costs of loading every tool
      “When the total reaches 10% of the window, tool search activates.”
      ↩︎ Match the fix to where the tokens go
      “Tool selection accuracy degrades with more than 30-50 tools loaded at once.”
      ↩︎ Key concept
      “Tool search is on. Tool definitions are deferred and discovered on demand.”
      ↩︎ Exam trap 3
    2. 2.
      “A typical multiserver setup (GitHub, Slack, Sentry, Grafana, and Splunk) can consume ~55k tokens in definitions before Claude does any work.”
      ↩︎ The two costs of loading every tool
      “Both tool search variants (regex and bm25) search tool names, descriptions, argument names, and argument descriptions.”
      ↩︎ Tool search: load definitions on demand
      “Keep your 3–5 most frequently used tools non-deferred so Claude can call them without searching first.”
      ↩︎ Tool search: load definitions on demand
      “The prefix is untouched, so prompt caching is preserved.”
      ↩︎ Tool search: load definitions on demand
      “defer_loading controls what enters the context window, not what you send in the request”
      ↩︎ Exam trap 1
      “Never set defer_loading: true on the tool search tool itself.”
      ↩︎ Exam trap 2
    3. 3.
      “This trades a small amount of latency (one extra turn to look up a tool) for a large reduction in baseline context usage.”
      ↩︎ Match the fix to where the tokens go

    Continue to page 2 of 2

    Trimming Agent Toolsets: Disable, Narrow, and Scope Tools