What you will be able to do
- Explain tool use as a contract in which Claude requests an operation and your code or Anthropic's servers carry it out
- Write a user-defined tool definition with name, description, input_schema and optional input_examples
- Rewrite a weak tool description so Claude picks the tool only when it should
- Apply consolidation, namespacing and high-signal responses when building a tool set
Key concept
Tool use as a contract — You declare which operations exist and what their inputs look like. Claude decides when to call them, but it never runs anything itself. The schema and description you write are all Claude knows about your system.
1.Tool use as a contract with external systems
A tool is a typed interface between Claude and your application. You say which operations exist and what their inputs look like, and Claude chooses when to call them. The model never runs anything itself. It emits a structured request, your code (or Anthropic's servers) runs the operation, and the result comes back into the conversation. Claude never sees your implementation. It sees only the schema you wrote and the result you returned. So the definition is the entire interface Claude has to your system.
You configure tools to let Claude reach systems outside the model. The documentation lists four cases where tools fit:
- Actions with side effects, such as sending an email, writing a file or updating a record. - Fresh or external data, such as current prices, today's weather or the contents of a database. - Structured, guaranteed-shape outputs, where a schema has to enforce the fields. - Calls into existing systems, such as databases, internal APIs and filesystems.
One practical sign: if you are writing a regex to pull a decision out of model output, that decision should have been a tool call.
Tools are the wrong choice in three cases:
- The model can answer from training alone (summarization, translation, general knowledge). - A one-shot Q&A has no side effects and nothing to execute. - The extra round trip a tool call costs would take longer than the trivial task itself.
Sources1
2.Anatomy of a tool definition
You declare client tools in the tools top-level parameter of the API request. A user-defined tool has three required fields and one optional field:
| Field | What it holds |
|---|---|
| name | The tool's identifier; must match ^[a-zA-Z0-9_-]{1,128}$ |
| description | Plaintext: what the tool does, when it should be used, how it behaves |
| input_schema | A JSON Schema object defining the expected parameters |
| input_examples | Optional array of example input objects to help Claude use the tool |
{
"name": "get_weather",
"description": "Get the current weather in a given location",
"input_schema": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g. San Francisco, CA"
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "The unit of temperature, either 'celsius' or 'fahrenheit'"
}
},
"required": ["location"]
}
}location is required. unit is optional, but the enum limits it to "celsius" or "fahrenheit", so the schema itself restricts what Claude can send. Further optional properties, including cache_control, strict, defer_loading and allowed_callers, are documented in the Tool reference. Anthropic-schema client tools such as bash and the text editor are declared differently: you give a date-versioned type instead of writing your own schema.
Claude does not receive these definitions raw. The API builds a special system prompt from your tool definitions, the tool configuration and your own system prompt. The description you write is therefore prompt text that Claude reads when it decides what to do.
Sources2
3.Writing descriptions that steer tool selection
The documentation is direct about this: detailed descriptions matter more to tool performance than anything else. A good description covers four things:
- what the tool does - when to use it, and when not to - what each parameter means and how it changes the tool's behaviour - caveats and limits, such as information the tool does not return
The guideline is at least 3–4 sentences per tool, and more for complex tools. The documentation gives this example of a poor one:
{
"name": "get_stock_price",
"description": "Gets the stock price for a ticker.",
"input_schema": {
"type": "object",
"properties": {
"ticker": {
"type": "string"
}
},
"required": ["ticker"]
}
}The documented good version keeps the same schema and only rewrites the text:
- The ticker must be a valid symbol for a company on a major US exchange such as NYSE or NASDAQ. - The tool returns the latest trade price in USD. - It should be used when the user asks for the current or most recent price of a stock. - It provides no other information about the stock or company.
The ticker parameter also gets its own description ("e.g. AAPL for Apple Inc."). Scope, output and exclusions are all spelled out, so Claude can tell when a request falls outside the tool.
input_examples is a secondary option. Descriptions come first. Examples help with complex inputs, nested objects, optional parameters and format-sensitive values. Each example must validate against the tool's input_schema. The examples go into the prompt next to the schema and show Claude when to include optional parameters and how to format them.
An application must guarantee that Claude calls the `submit_refund` tool on every turn of a specific workflow step, and that no other tool is used in its place. Which `tool_choice` configuration achieves this?
Correct answer: A — `{"type": "tool", "name": "submit_refund"}`
- A. Correct. Setting `tool_choice` to the `tool` type with the specific tool name forces Claude to always use that particular tool, which is the only option that guarantees `submit_refund` specifically rather than any tool.
- B. `any` forces Claude to use one of the provided tools but does not force a particular tool, so Claude could call a different tool in the set instead of `submit_refund`.
- C. `auto` lets Claude decide whether to call any tool at all; a prompt instruction can steer behavior but provides no guarantee, unlike a forced `tool_choice`.
- D. `none` prevents Claude from using any tools at all, and array position has no effect on tool selection, so this configuration guarantees the opposite of the desired outcome.
Sources2
4.Constructing the tool set
Good individual descriptions are only part of the job. The design of the whole tool set also affects how well Claude uses it. The same guidance gives three rules for building the set.
Consolidate related operations. Don't make create_pr, review_pr and merge_pr separate tools. Expose one tool with an action parameter. Fewer, more capable tools reduce selection ambiguity.
Namespace by service. When tools span several services or resources, prefix each name with the service, for example github_list_prs or slack_send_message. This keeps selection unambiguous as the library grows, and it matters most when you use tool search.
Return only high-signal information. Tool results also use context. Return semantic, stable identifiers such as slugs or UUIDs, not opaque internal references. Include only the fields Claude needs for its next step. Bloated responses waste context and make the important parts harder to find.
| Symptom | Documented practice |
|---|---|
| A separate tool for every action (create_pr, review_pr, merge_pr) | One tool with an action parameter |
| Tools from several services with generic names | Service prefix, e.g. github_list_prs, slack_send_message |
| Results full of opaque internal references and extra fields | Semantic, stable identifiers (slugs or UUIDs) and only the fields Claude needs |
A custom `fetch_invoice` client tool calls a downstream billing API. During a load spike the API begins returning HTTP 503 responses. The engineering team wants Claude to recognize the failure, explain it to the user, and avoid treating the 503 body as valid invoice data. How should the tool result be constructed for this call?
Correct answer: A — Return a `tool_result` block with `tool_use_id` set to the original call's id, `content` describing the 503 and suggesting a retry window, and `is_error` set to `true`
- A. Correct. Setting `is_error: true` explicitly marks the call as failed rather than letting Claude parse ambiguous data, and an instructive message (what failed, what to try next) gives Claude the context needed to inform the user or retry appropriately.
- B. Omitting `is_error` forces Claude to guess whether the returned content is valid invoice data or an error body, risking the 503 payload being treated as real data instead of being surfaced as a failure.
- C. Every `tool_use` block must be immediately followed by a matching `tool_result`; omitting it produces a malformed message sequence and the API will reject the request.
- D. Marking a failed call as `is_error: false` with empty content misrepresents the outcome, and Claude has no signal that the billing API call did not succeed.
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.When Claude picks the wrong tool, the first fix is to add input_examples.Why is that wrong?
Wrong tool selection is a description problem. Descriptions are the most important factor. input_examples help mainly with complex, nested or format-sensitive inputs.
2.Giving each action its own narrowly named tool makes selection clearer for Claude.Why is that wrong?
The guidance says the opposite. Group related operations into one tool with an action parameter, because fewer tools reduce selection ambiguity.
Covered in Constructing the tool set
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Claude never sees your implementation; it only sees the schema you provided and the result you returned.”
↩︎ Tool use as a contract with external systems“Tool use is the bridge between natural-language requests and the systems that fulfill them.”
↩︎ Tool use as a contract with external systems“if you're writing a regex to extract a decision from model output, that decision should have been a tool call.”
↩︎ Tool use as a contract with external systems“You specify what operations are available and what shape their inputs and outputs take; Claude determines when and how to call them.”
↩︎ Key concept - 2.
“Must match the regex ^[a-zA-Z0-9_-]{1,128}$.”
↩︎ Anatomy of a tool definition“the API constructs a special system prompt from the tool definitions, tool configuration, and any user-specified system prompt.”
↩︎ Anatomy of a tool definition“Provide extremely detailed descriptions. This is by far the most important factor in tool performance.”
↩︎ Writing descriptions that steer tool selection“Aim for at least 3–4 sentences for each tool description, more if the tool is complex.”
↩︎ Writing descriptions that steer tool selection“Each example must be valid according to the tool's input_schema”
↩︎ Writing descriptions that steer tool selection“prefix names with the service (for example, github_list_prs, slack_send_message)”
↩︎ Constructing the tool set“Return semantic, stable identifiers (for example, slugs or UUIDs) rather than opaque internal references”
↩︎ Constructing the tool set“Bloated responses waste context and make it harder for Claude to extract what matters.”
↩︎ Constructing the tool set“Clear descriptions are most important, but for tools with complex inputs, nested objects, or format-sensitive parameters, you can use the input_examples field”
↩︎ Exam trap 1“Fewer, more capable tools reduce selection ambiguity and make your tool surface easier for Claude to navigate.”
↩︎ Exam trap 2