CertSafari
    CCAR-P · Lessons

    Domain 3 · Lesson 19/38

    Tool Search and defer_loading: When to Load Tools on Demand

    Evaluate progressive discovery vs. monolithic context strategy

    9 min read
    2.38% of exam
    3 sources
    Published 27 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Configure deferred tool loading with the tool search tool and defer_loading
    • Weigh the latency cost of tool search against its context and accuracy benefits
    • Choose between upfront loading, tool search and prompt caching for a given toolset

    1.How deferred loading works

    For tools, progressive discovery is built into the API as the tool search tool. Claude starts with a small searchable catalog instead of every schema, and loads only the tools a request needs. The documentation says this typically cuts definition tokens by over 85 percent, loading only the 3–5 tools Claude needs for a given request.

    There are two variants. With tool_search_tool_regex_20251119, Claude writes regex patterns. With tool_search_tool_bm25_20251119, it writes natural-language queries. Both search tool names, descriptions, argument names and argument descriptions. To use either, add it to your tools list and mark every tool that should not load up front with defer_loading: true.

    A tool marked for on-demand loadingjson
    {
      "name": "get_weather",
      "description": "Get current weather for a location",
      "input_schema": {
        "type": "object",
        "properties": {
          "location": { "type": "string" },
          "unit": { "type": "string", "enum": ["celsius", "fahrenheit"] }
        },
        "required": ["location"]
      },
      "defer_loading": true
    }

    Note what the flag does and does not change. You still send every tool's full definition on every request, because the API needs them server-side to run the search and expand results. The flag only decides whether a definition enters Claude's context up front. At least one tool, normally the search tool itself, must stay non-deferred.

    At runtime, Claude calls the search tool. The API runs the search and returns matches as tool_reference blocks (up to 5 by default), then expands them into full definitions. Deferred tools are left out of the system-prompt prefix, and discovered ones are added inline in the conversation, so the prefix stays the same and prompt caching still works.

    Sources1

    2.The trade-off, and when discovery switches on

    Progressive discovery has a cost. When Claude needs a deferred tool, it spends one extra turn looking it up. The documentation describes the trade as a small amount of latency in exchange for a large reduction in baseline context usage. That is why it recommends keeping your 3–5 most frequently used tools non-deferred: the tools Claude needs most often are then available without a search. In practice this is a hybrid. A small fixed set is loaded up front and the long tail is discovered on demand.

    The Claude Agent SDK can make this decision for you. Tool search is on by default for Claude Opus 4.5, Sonnet 4.5, Haiku 4.5 and later models, and the ENABLE_TOOL_SEARCH environment variable controls it.

    ENABLE_TOOL_SEARCH values in the Agent SDK
    ValueBehavior
    (unset)Tool search is on; definitions are deferred and discovered on demand, with fallbacks to upfront loading on some deployments
    trueTool search is always on, except where the serving stack forces upfront loading
    autoTool search activates when deferrable tool definitions reach 10% of the model's context window; below that, all definitions load upfront
    auto:NSame as auto with a custom percentage; lower values activate sooner
    falseTool search is off; all tool definitions load into context on every turn

    auto sets the choice between the two strategies by size. While definitions are a small share of the window, upfront loading costs little and avoids the search turn. Once they pass the threshold, the SDK switches to discovery.

    Agent SDK: switch to tool search once remote MCP tool definitions reach 5% of contextpython
        options = ClaudeAgentOptions(
            mcp_servers={
                "enterprise-tools": {
                    "type": "http",
                    "url": "https://tools.example.com/mcp",
                }
            },
            allowed_tools=[
                "mcp__enterprise-tools__*"
            ],  # Wildcard pre-approves all tools from this server
            env={
                "ENABLE_TOOL_SEARCH": "auto:5"  # Activate tool search when deferrable definitions reach 5% of context
            },
        )

    An enterprise agent aggregates tool catalogs from GitHub, Slack, Sentry, Grafana, and Splunk MCP servers, and the combined tool definitions consume about 55,000 tokens before the agent does any work. Which strategy best addresses this context bloat?

    Sources213

    3.Choosing a strategy: match the fix to the pressure

    Tool search is one of four context-management approaches in the documentation. Each one targets a different source of token pressure, so picking the right one starts with finding where your tokens are going.

    Which approach addresses which kind of context pressure
    ApproachWhat it reducesWhen it fits
    Tool searchTool definitions loaded upfrontLarge toolsets (20+ tools) where most tools aren't needed every turn
    Prompt cachingToken cost of repeated tool definitionsStable toolsets across many requests
    Programmatic tool callingtool_result roundtripsChains of tool calls that can execute as a single script
    Context editingOld tool_result blocks in historyLong conversations where early results are no longer relevant

    The key distinction is between tool search and prompt caching. Caching does not reduce the number of tokens in context. It reduces what you pay for them on later requests. It is the right choice when a toolset is large but fixed, but it leaves every definition in front of the model, so it cannot fix degraded tool selection. Only reducing what Claude sees does that.

    The approaches are not alternatives. The documentation's suggested starting point for a high-volume agent is to enable prompt caching on tool definitions from day one, add tool search once the toolset grows past roughly 20 tools or baseline context usage becomes noticeable, and add context editing when conversations run long enough that early results stop mattering.

    During testing, an engineering team notices that Claude increasingly selects the wrong tool as more MCP servers are connected to their agent. At what approximate number of simultaneously loaded tools does tool selection accuracy begin to noticeably degrade?

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Setting defer_loading: true means you can leave those tool definitions out of the request payload.Why is that wrong?

      You still send every full definition on every request. The flag only controls whether a definition enters Claude's context up front.

      Covered in How deferred loading works

    2. 2.Prompt caching the full tool list solves the same problem as tool search.Why is that wrong?

      Caching lowers the price of repeated definitions, but every token stays in context, so the selection-accuracy problem remains.

      Covered in Choosing a strategy: match the fix to the pressure

    3. 3.The best progressive configuration defers every tool, including the most-used ones.Why is that wrong?

      Each lookup costs an extra turn, so the most frequently used tools should stay loaded. The search tool itself must never be deferred.

      Covered in The trade-off, and when discovery switches on

    Practise it for real

    Convert a Messages API request from upfront tool loading to on-demand tool discovery.

    1. 1.Add {"type": "tool_search_tool_regex_20251119", "name": "tool_search_tool_regex"} to the tools array, and do not give it defer_loading.

      Why: At least one tool must stay non-deferred, normally the search tool itself.

      You should see: The request is accepted, with the search tool loaded into context.

    2. 2.Add "defer_loading": true to each tool that is rarely used, keeping your 3–5 most frequently used tools non-deferred.

      Why: Deferred tools stay out of context until they are found, and the frequent ones can be called without a search turn.

      You should see: You still send every tool's full definition, but only the non-deferred ones start in context.

    3. 3.Send a user message that needs one of the deferred tools, such as a weather question when get_weather is deferred.

      Why: This triggers discovery instead of a direct call.

      You should see: The response has a server_tool_use block, a tool_search_tool_result with tool_references, then a tool_use for the discovered tool, and ends with stop_reason "tool_use".

    4. 4.Run the discovered tool and return a tool_result for its toolu_ ID. Keep the tool_search_tool_result block in the message history unchanged.

      Why: The search runs on Anthropic's servers. You only handle the discovered tool, exactly as in standard tool use.

      You should see: Claude continues with the tool's output. You never return a tool_result for the srvtoolu_ ID.

    Stuck? Get a nudge

    If Claude never finds the tool, look at its description and argument names. Both search variants match against those fields.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Tool search typically reduces this by over 85 percent, loading only the 3–5 tools Claude needs for a given request.”
      ↩︎ How deferred loading works
      “defer_loading controls what enters the context window, not what you send in the request”
      ↩︎ How deferred loading works
      “At least one tool, normally the tool search tool itself, must stay non-deferred.”
      ↩︎ How deferred loading works
      “The prefix is untouched, so prompt caching is preserved.”
      ↩︎ How deferred loading works
      “Keep your 3–5 most frequently used tools non-deferred so Claude can call them without searching first.”
      ↩︎ The trade-off, and when discovery switches on
      “defer_loading controls what enters the context window, not what you send in the request”
      ↩︎ Exam trap 1
      “Keep your 3–5 most frequently used tools non-deferred so Claude can call them without searching first.”
      ↩︎ Exam trap 3
    2. 2.
      “This trades a small amount of latency (one extra turn to look up a tool) for a large reduction in baseline context usage.”
      ↩︎ The trade-off, and when discovery switches on
      “Prompt caching doesn't reduce the number of tokens in context, but it reduces what you pay for them on subsequent requests.”
      ↩︎ Choosing a strategy: match the fix to the pressure
      “This is the right choice when the toolset is large but fixed.”
      ↩︎ Choosing a strategy: match the fix to the pressure
      “Add tool search once your toolset grows past roughly 20 tools or your baseline context usage becomes noticeable.”
      ↩︎ Choosing a strategy: match the fix to the pressure
      “Prompt caching doesn't reduce the number of tokens in context, but it reduces what you pay for them on subsequent requests.”
      ↩︎ Exam trap 2
    3. 3.
      “When the total reaches 10% of the window, tool search activates.”
      ↩︎ The trade-off, and when discovery switches on

    Ready to test yourself?

    Practise the 12 questions on this subdomain.