What you will be able to do
- Explain why the tool description, more than the name or the implementation, decides whether Claude selects a tool correctly
- Write a description that covers what the tool does, when to use it and when not to, what each parameter means, and its limitations
- Use parameter descriptions and input_examples to pin down input formats, and recognise when a failure is really a schema gap
- Diagnose misrouting caused by overlapping tools, and fix it by renaming, rewriting the description, or splitting a generic tool into purpose-specific tools with defined input and output contracts
- Review a system prompt for keyword-sensitive instructions that could override a well-written tool description
Key concept
The description is the tool's interface to the model — Claude cannot see how a tool is implemented. It chooses a tool, and decides how to call it, from the name, description and schema you supply. Whatever the description leaves out, the model has to guess.
1.Selection happens from the description
When you give Claude a set of tools, it makes two decisions on every turn: whether to call a tool at all, and which one. It makes both from text. The API documentation says Claude chooses based on the user's request and the tool's description. Your code, your database and the service behind the tool are all invisible to the model. The name, the description and the input schema are the only evidence it has. In other words, the tool description is the primary mechanism an LLM uses for tool selection. The model matches the request against each tool's described capability, and calls the tool whose description fits.
This explains why Anthropic ranks the description above everything else in a tool definition. Its guidance puts it plainly: detailed descriptions are "by far the most important factor in tool performance." The reference table for a tool definition sets the same expectation for the description field. It should be a detailed plaintext account of what the tool does, when to use it, and how it behaves.
The effect of a minimal description is easiest to see when tools look alike. Suppose there is one weather tool and the user asks about the weather. A one-line description may be enough, because nothing else competes for the request. Add a second tool that also deals with locations or dates, and one-liners no longer separate the two. Claude then has to guess which tool the request belongs to, and routing starts to drift. Minimal descriptions lead to unreliable selection among similar tools, and the unreliability grows with every tool you add. The fix is not a smarter model. The fix is a description that answers the questions the model would otherwise have to guess at.
Anthropic's troubleshooting guide names the pattern directly. The symptom "Claude calls tool A when you wanted tool B" has one likely cause listed: description ambiguity. The fix is to sharpen the descriptions. The same page lists a second selection failure, where Claude never calls a tool at all. There the likely cause is a tool name collision or an overly generic schema, and the fix is to check for duplicate names and add input examples that make the intended use concrete. Both failures are description and naming problems, not model problems.
2.What a good description contains
Anthropic lists four things a description should explain: what the tool does, when it should be used and when it should not, what each parameter means and how it changes the tool's behaviour, and any important caveats or limitations. The guidance adds that a description should say what information the tool does *not* return. It suggests at least three to four sentences per tool, and more for complex tools.
{
"name": "get_stock_price",
"description": "Retrieves the current stock price for a given ticker symbol. The ticker symbol must be a valid symbol for a publicly traded company on a major US stock exchange like NYSE or NASDAQ. The tool will return the latest trade price in USD. It should be used when the user asks about the current or most recent price of a specific stock. It will not provide any other information about the stock or company.",
"input_schema": {
"type": "object",
"properties": {
"ticker": {
"type": "string",
"description": "The stock ticker symbol, e.g. AAPL for Apple Inc."
}
},
"required": ["ticker"]
}
}{
"name": "get_stock_price",
"description": "Gets the stock price for a ticker.",
"input_schema": {
"type": "object",
"properties": {
"ticker": {
"type": "string"
}
},
"required": ["ticker"]
}
}The poor version leaves real questions open. Does it accept a company name or only a ticker? Which exchanges does it cover? Is the price live or the previous close, and in which currency? Should Claude use it for questions about a company's earnings? The good version answers each of these. Its last sentence is a boundary: "It will not provide any other information about the stock or company." A sentence like that keeps a tool from being picked for requests just outside its purpose. That matters most when a similar tool sits right next to it.
| Element | What the good description says |
|---|---|
| What the tool does | Retrieves the current stock price for a given ticker symbol |
| When to use it | When the user asks about the current or most recent price of a specific stock |
| Input meaning and format | A valid ticker on a major US exchange such as NYSE or NASDAQ; the parameter gives AAPL as an example |
| Caveats and limitations | Returns the latest trade price in USD and no other information about the stock or company |
Those four elements cover one tool in isolation. When a similar tool exists, a good description also has to differentiate. Anthropic's troubleshooting advice is to differentiate tools by *when* to use them, not only by *what* they do. Two tools can do overlapping things, but their descriptions should make clear which one to reach for in a given situation, and that means naming the alternative. A description that says "Use this for the current price of one ticker. For historical prices or company fundamentals, use get_company_profile instead" tells Claude the purpose, the expected input, the output, and when to pick this tool versus the similar one. Purpose, inputs, outputs and the when-versus-alternatives boundary are the four things a description of a tool with neighbours must clearly differentiate.
3.Input formats, examples and edge cases
Choosing the right tool is half the job. The other half is calling it with well-formed input. Format expectations go in the parameter descriptions. The documented get_weather tool describes its location parameter as "The city and state, e.g. San Francisco, CA". That one example shows Claude the shape of a multi-word place name, so it does not have to invent one. An enum on unit limits that parameter to two values.
When a sentence of text cannot carry the shape of the input, add the optional input_examples field. Anthropic recommends it for complex tools with nested objects, optional parameters, or format-sensitive inputs. Each example must be valid against the tool's input_schema. The examples are placed in the prompt next to the schema, so Claude sees concrete, well-formed calls. For instance, the documented weather examples include one call that omits the optional unit, which shows Claude that leaving it out is allowed. Examples add to the description. They do not replace it.
Edge cases often show up as bad parameters rather than as misrouting. Anthropic's troubleshooting table covers two common ones. If Claude sends a parameter that does not exist in your schema, the cause is over-generation, and the fix is strict mode where the schema supports it. If a user leaves out a required value, some models will not ask for it. Claude Sonnet in particular may fill in a plausible value of its own. A description that states what is required, and what the tool should do when it is missing, closes both gaps.
Two tools, `analyze_content` (summarizes pasted text) and `analyze_document` (summarizes uploaded documents), share nearly identical one-line descriptions. Users report the model frequently calls `analyze_document` for pasted web article text instead of `analyze_content`. What is the most effective fix?
Correct answer: A — Rewrite each description to state the expected input type and add a note on when the other tool applies.
- A. Correct. Tool descriptions are the model's primary signal for tool selection; stating the expected input type and contrasting the two tools removes the ambiguity causing the misrouting.
- B. Incorrect. Sampling temperature affects token-level randomness in generated text, not the semantic reasoning the model uses to distinguish two functionally similar tool descriptions.
- C. Incorrect. Promotional wording adds length without adding differentiating information about input type or scope, so the ambiguity between the two tools remains.
- D. Incorrect. Removing a tool eliminates its intended use case rather than fixing the underlying description ambiguity, and pasted text would now be mishandled by a document-only tool.
4.Overlapping tools: rename, rewrite, or split
Misrouting is what happens when two tool descriptions overlap. Consider a tool library with analyze_content, described as "Analyzes content and returns key information", and analyze_document, described as "Analyzes a document and returns key information". The two names share a verb and the two descriptions are near-identical. Nothing in either says what kind of content, what the input looks like, what comes back, or when one applies and the other does not. Claude has no basis for choosing, so it picks one, and which one it picks can change from request to request. That is ambiguous or overlapping tool descriptions causing misrouting, and it is the exact symptom Anthropic's troubleshooting table attributes to description ambiguity.
There are three fixes, and a scenario question usually wants you to pick the right one. The first is renaming plus rewriting the description to eliminate the functional overlap. If analyze_content is really used to pull results out of web pages, rename it extract_web_results and give it a web-specific description: it takes a URL or fetched page, returns a list of results, and is not for uploaded files. Anthropic's naming guidance makes the same point about namespacing. Prefixing names by service or resource, such as github_list_prs or slack_send_message, makes tool selection unambiguous as the library grows. A name that carries the domain does part of the description's work, and it stops a duplicate or near-duplicate name from causing Claude to skip the tool altogether.
The second fix is splitting a generic tool into purpose-specific tools with defined input and output contracts. A generic analyze_document hides several different jobs behind one vague verb. Split it into extract_data_points, which takes a document and returns structured fields; summarize_content, which takes a document and returns a short summary; and verify_claim_against_source, which takes a claim and a source and returns whether the source supports it. Each of the three now has a clear purpose, a specific input shape and a specific output shape, and a description that can say when it applies. The generic tool could not say any of that, because it had no single answer.
Splitting has a counterweight you must know for the exam. Anthropic also recommends consolidating related operations into fewer tools: rather than separate create_pr, review_pr and merge_pr tools, use one tool with an action parameter, because fewer, more capable tools reduce selection ambiguity. The two pieces of advice are not in conflict. Consolidate when operations act on the same resource and differ only by action. Split when one tool covers genuinely different purposes with different inputs and outputs, so that no honest single description can be written for it. The test in both directions is the same: after the change, can each description clearly say what the tool does, what it takes, what it returns, and when to use it instead of its neighbours?
The third fix, which applies to all of the above, is to watch the evidence. Anthropic's engineering guidance on writing tools recommends reading evaluation transcripts and tool-calling metrics to find where agents get confused, and gives an example from its own web search tool. Claude was needlessly appending a year to the query parameter and biasing results, and the team steered it back by improving the tool description. Lots of errors on one tool, or a tool that is picked when a neighbour should have been, are the signal to rename, rewrite or split.
A single tool, `analyze_document`, extracts data points, produces summaries, and verifies claims against a source. Users report it inconsistently performs only one of these behaviors for ambiguous requests. Which redesign best addresses this?
Correct answer: C — Split the tool into `extract_data_points`, `summarize_content`, and `verify_claim_against_source`, each with a narrow contract.
- A. Incorrect. Forcing three calls per request wastes invocations on tasks that need only one behavior and does not resolve the ambiguity about which behavior was intended.
- B. Incorrect. Merging with unrelated tools increases the surface area of a single tool's responsibilities, worsening overlapping-purpose confusion rather than resolving it.
- C. Correct. Splitting a generic multi-purpose tool into purpose-specific tools with defined input/output contracts removes the ambiguity about which behavior a call should trigger.
- D. Incorrect. Adding a mode parameter without updating the description still leaves the model guessing at the tool's overall purpose and when each mode applies.
5.System prompt wording can override the description
Tool descriptions do not reach the model on their own. When you call the API with a tools parameter, the API constructs a special system prompt from the tool definitions, the tool configuration, and your own system prompt. All of it sits in one place, and the model reads it as one set of instructions. That is why the wording of your system prompt has an impact on tool selection. Anthropic documents that the decision to call a tool at all is steerable through the system prompt: "Use the tools to investigate before responding" increases tool use, "Always call a tool first before responding" pushes further, and "Use your judgment about whether to call a tool or respond directly" keeps it conservative. If a sentence in the system prompt can change whether a tool is called, a sentence can also change which one.
The failure mode to look for is a keyword-sensitive instruction. Suppose the system prompt says "Whenever the user mentions a document, analyze it thoroughly" and the tool library contains analyze_document. The instruction and the tool name share a keyword, and the model can form an unintended tool association: every request containing the word "document" gets routed to analyze_document, even when the user wants a summary and summarize_content has a precise description saying it handles exactly that. The well-written tool description loses to the system prompt because the instruction is more direct and shares the model's matching keyword. Anthropic's tool search guidance shows the same mechanism working in your favour: it recommends using keywords in descriptions that match how users describe tasks, because the model matches on the words it sees. Keywords in the system prompt pull in the same way, whether you intended them to or not.
So reviewing tool descriptions is not enough. When a well-described tool is still misrouted, review the system prompt for keyword-sensitive instructions that might override the descriptions. Look for sentences that name a tool, a verb from a tool name, or a noun that appears in only one tool's description, and that attach a blanket rule to it. Rewrite them to describe the goal rather than the keyword, or move the guidance into the relevant tool's description, where it is scoped to that tool. If you genuinely need one tool called regardless of wording, use tool_choice to require it instead of relying on prompt phrasing.
A support agent has both `search_web` and `fetch_webpage_results` as separate tools. Testing shows the model almost always calls `search_web`, even for tasks better suited to `fetch_webpage_results`. The system prompt contains "Always prefer searching for the most up to date information." What most likely explains the bias?
Correct answer: D — Keyword-sensitive system prompt wording creates an unintended association that overrides the more accurate tool description.
- A. Incorrect. Parallel tool use restrictions are not described here, and there is no indication the fetch tool is structurally excluded from the candidate set.
- B. Incorrect. Tool schemas do not carry a temperature value; temperature is a model-level sampling parameter unrelated to individual tool definitions.
- C. Incorrect. There is no fixed token-length threshold that categorically excludes longer tool descriptions from being selected by the model.
- D. Correct. The phrase 'always prefer searching' repeatedly emphasizes the word 'search,' creating a keyword-driven bias that overrides the more suitable tool's description.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If the tool name is self-explanatory, a one-line description is enough.Why is that wrong?
A one-line description leaves out scope, output, when-not-to-use and limitations. Anthropic treats detailed descriptions as the main factor in tool performance, and its own weak example is a one-liner with a clear name.
Covered in What a good description contains
2.Adding input_examples is a substitute for writing a fuller description.Why is that wrong?
Examples help with complex or format-sensitive inputs, but the description still comes first. The examples add to it.
Covered in Input formats, examples and edge cases
3.Splitting every tool into the smallest possible operation always improves tool selection.Why is that wrong?
Anthropic recommends consolidating related operations on one resource into a single tool with an action parameter. Split only when a generic tool hides different purposes with different input and output contracts.
Covered in Overlapping tools: rename, rewrite, or split
4.Once the tool descriptions are well written, the system prompt cannot affect which tool is chosen.Why is that wrong?
The user system prompt is assembled into the same constructed prompt as the tool definitions, and Anthropic documents that its wording steers tool-calling behaviour. A keyword-sensitive instruction can override a good description.
Covered in System prompt wording can override the description
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Claude determines when to call a tool based on the user's request and the tool's description.”
↩︎ Selection happens from the description“It calls a tool when the request maps to that tool's described capability and the answer isn't already in context.”
↩︎ Selection happens from the description“Claude (particularly Claude Sonnet) might guess values you didn't supply”
↩︎ Input formats, examples and edge cases“This boundary is steerable through your system prompt.”
↩︎ System prompt wording can override the description“To require a tool call rather than rely on prompting, set tool_choice.”
↩︎ System prompt wording can override the description“This boundary is steerable through your system prompt.”
↩︎ Exam trap 4 - 2.
“Provide extremely detailed descriptions. This is by far the most important factor in tool performance.”
↩︎ Selection happens from the description“Any important caveats or limitations, such as what information the tool does not return if the tool name is unclear.”
↩︎ What a good description contains“This is particularly useful for complex tools with nested objects, optional parameters, or format-sensitive inputs.”
↩︎ Input formats, examples and edge cases“Each example must be valid according to the tool's input_schema”
↩︎ Input formats, examples and edge cases“This makes tool selection unambiguous as your library grows, and is especially important when using tool search.”
↩︎ Overlapping tools: rename, rewrite, or split“Fewer, more capable tools reduce selection ambiguity and make your tool surface easier for Claude to navigate.”
↩︎ Overlapping tools: rename, rewrite, or split“the API constructs a special system prompt from the tool definitions, tool configuration, and any user-specified system prompt.”
↩︎ System prompt wording can override the description“The poor description is too brief and leaves Claude with many open questions about the tool's behavior and usage.”
↩︎ Exam trap 1“Prioritize descriptions, but consider using input_examples for complex tools.”
↩︎ Exam trap 2“Fewer, more capable tools reduce selection ambiguity and make your tool surface easier for Claude to navigate.”
↩︎ Exam trap 3 - 3.
“A detailed plaintext description of what the tool does, when it should be used, and how it behaves.”
↩︎ Selection happens from the description“When it should be used (and when it shouldn't)”
↩︎ What a good description contains - 4.
“Claude calls tool A when you wanted tool B | Description ambiguity | Sharpen descriptions.”
↩︎ Selection happens from the description“Check for duplicate names across your tool list. Add input_examples to make the intended use concrete.”
↩︎ Selection happens from the description“Sharpen descriptions. Differentiate tools by WHEN to use them, not only WHAT they do.”
↩︎ What a good description contains“Add strict: true if your schema is in the supported subset.”
↩︎ Input formats, examples and edge cases“Claude never calls your tool | Tool name collision or overly-generic schema”
↩︎ Overlapping tools: rename, rewrite, or split - 5.https://www.anthropic.com/engineering/writing-tools-for-agentsSecondary source
“we steered Claude in the right direction by improving the tool description”
↩︎ Overlapping tools: rename, rewrite, or split“lots of tool errors for invalid parameters might suggest tools could use clearer descriptions or better examples”
↩︎ Overlapping tools: rename, rewrite, or split - 6.
“Use keywords in descriptions that match how users describe tasks.”
↩︎ System prompt wording can override the description
Also cited
“Claude never sees your implementation; it only sees the schema you provided and the result you returned.”
↩︎ Key concept