What you will be able to do
- Explain why a tool call with a JSON Schema is the most reliable way to get schema-compliant structured output from Claude
- Define an extraction tool and read the structured data from the tool_use block in the response
- Describe what strict: true guarantees, and what it leaves out
- Choose between tool_choice auto, any and a forced named tool for a given extraction pipeline
- Decide which schema fields to make required, optional or nullable so the model is never forced to fabricate a value the source does not contain
- Design enum fields with an "unclear" value for ambiguous cases and an "other" value plus a detail string for categories the enum does not list
- Pair a strict output schema with format normalization rules in the prompt so inconsistently formatted source data comes out in one consistent form
Key concept
Tool input as the output channel — In structured extraction, you define a tool only to receive data. The arguments Claude passes to it, which are shaped by your JSON Schema, are the structured result. The tool never has to run.
1.Why tool use is the structured-output channel
If you ask Claude in plain text to "reply in JSON", your code still has to parse whatever text comes back. The structured outputs documentation lists what goes wrong even with careful prompting: parsing errors from invalid JSON syntax, missing required fields, inconsistent data types, and schema violations that need error handling and retries.
Tool use avoids this. You define a tool that exists only to receive data. Anthropic's extraction cookbook defines custom tools for this purpose, getting well-structured JSON for summarization, entity extraction, sentiment analysis and text classification. Your code never has to perform an action with the tool. The arguments Claude supplies are the result, and because a JSON Schema describes those arguments, you decide the shape of the output.
The documentation also offers a second, complementary route: JSON outputs through output_config.format. This returns schema-valid JSON in a text block instead of a tool call. You can use the two separately or together in the same request. This lesson focuses on the tool-use route, because that is where the tool_choice control applies.
A schema settles the shape of the output, not the formatting of the values inside it. A field typed as string will accept "2024-01-15", "15/01/2024" and "Jan 15, 2024" equally, and source documents rarely agree on a format. That is why the schema and the prompt work as a pair. The extraction cookbook's open-ended example provides an input_schema and instructs Claude via prompting how to interact with the tool, and the tool-choice cookbook says that even when a tool is forced, you should still employ some basic prompt engineering. Put your format normalization rules in that prompt, next to the strict schema: state the date format, the currency and decimal convention, the casing for names and codes, and how to treat missing punctuation or abbreviations. The define-tools page names format-sensitive parameters as a case where input_examples help, so you can also show a normalized value in an example. The schema then guarantees the structure and the prompt makes the values consistent.
2.Defining an extraction tool and reading the tool_use block
A user-defined tool has three core fields: name, description and input_schema. The input_schema is a JSON Schema object defining the expected parameters. The documentation's standard example contains the three parts you reuse for extraction: typed properties, an enum that limits a field to fixed values, and a required array that separates mandatory fields from optional ones.
{
"name": "get_weather",
"description": "Get the current weather in a given location",
"input_schema": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g. San Francisco, CA"
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "The unit of temperature, either 'celsius' or 'fahrenheit'"
}
},
"required": ["location"]
}
}In this example, location is required. unit is optional, but if it appears it must be one of the two enum values. The description fields are part of your extraction instructions. The documentation calls detailed descriptions by far the most important factor in tool performance. It recommends at least 3–4 sentences per tool, covering what the tool does, when to use it, and what each parameter means.
When Claude calls the tool, the response has a stop_reason of tool_use and one or more tool_use content blocks. Each block has an id, the tool's name, and input, an object that conforms to the tool's input_schema. That input object is your structured data.
{
"type": "tool_use",
"name": "get_weather",
"input": {
"location": "San Francisco, CA"
}
}Read them from the input of the tool_use block, not from the text. The content array can mix text and tool_use blocks, so select the block by its type rather than by its position. The documentation's round-trip example does exactly this, as shown below.
tool_use = next(block for block in response.content if block.type == "tool_use")
print(f"Claude called {tool_use.name} with {json.dumps(tool_use.input)}")Schema design: required versus optional. The required array is a promise that every extraction will carry those fields. That is the right promise for data every document has, such as an invoice number on an invoice. It is the wrong promise for data a document may lack. With strict mode the schema is enforced by constrained sampling, so a required field cannot be omitted, and if the source has no value for it the model has to put something there. The result is a fabricated value that looks exactly like a real one. The defence is to leave such fields out of required, as the weather example does with unit, so the model can simply omit them, as the tool_use block above shows. A nullable type is the other option when downstream code wants the key present but empty. The structured outputs page's API-response example declares errors as list[dict] | None for exactly that reason.
Schema design: enums that admit uncertainty and growth. An enum keeps a field to fixed values, which is what makes it useful for categorisation. The same fixed list is a trap when the source is ambiguous or the category is one you did not anticipate. Two additions solve this. Add an "unclear" value so the model has an honest answer when the document does not settle the question, instead of guessing between the listed values. And when the categories are open-ended, add an "other" value together with a separate optional detail string, for example category_other_detail, where the model writes what it actually found. Both patterns keep the enum strict, which strict mode requires, while giving the model a schema-valid way to say "none of these". Describe in the enum's description when each value applies, since the description is what the model reads to choose.
Strict mode guarantees the required field is present, so for a contract with no governing-law clause the model must still fill it, and it will invent a plausible jurisdiction. Remove governing_law from required, or declare it nullable, and say in the description that it should be omitted or null when the contract does not state one. Absence is then a valid, truthful output.
3.Strict mode: eliminating syntax and type errors
With a plain tool definition, input that matches the schema is expected but not guaranteed. Adding strict: true makes it a guarantee. Token sampling is constrained so that only schema-valid output can be produced, a technique the documentation calls grammar-constrained sampling. The documentation's example is a booking system that needs passengers as an integer. Without strict mode, Claude might send "two" or "2". With strict: true, the response always contains passengers: 2.
Strict tool use compiles tool schemas into grammars using the same pipeline as structured outputs. The schema uses standard JSON Schema with some limitations. The guarantees it lists are about form: tool input strictly follows the input_schema, and the tool name is always valid. The structured outputs page describes the benefits in the same terms: no more JSON.parse() errors, guaranteed field types and required fields, and no retries for schema violations. All of these are syntax and type failures. Whether the values themselves are correct is a separate question, and schema design is where you start to answer it.
4.tool_choice: auto, any, or a named tool
Defining an extraction tool does not make Claude call it. The tool_choice parameter controls how much freedom Claude has. It is part of the tool configuration that the API includes in the system prompt it builds from your tools.
| tool_choice | Must Claude call a tool? | Who picks the tool? | Extraction use |
|---|---|---|---|
| {"type": "auto"} (the default) | No: it may reply with text instead | Claude | Unsafe when the output must be structured |
| {"type": "any"} | Yes | Claude, from the provided tools | Several extraction schemas and an unknown document type |
| {"type": "tool", "name": "print_sentiment_scores"} | Yes | You | One specific extraction must run on this turn |
auto is the default. Claude decides whether a tool is needed, so an extraction pipeline running on auto can get a plain-text reply instead of data. The tool-choice cookbook shows this. With a sentiment tool and a calculator available on auto, Claude answered one tweet in prose (stop_reason end_turn) and called the calculator for another. If your pipeline must always produce structured output, do not use auto.
any tells Claude it must call one of the provided tools but lets it choose which. This fits extraction when you have several schemas, for example one per document type, and you don't know in advance which kind of document you have. Every response is a tool call, and Claude picks the schema that fits. The cookbook's SMS assistant uses any so that it never gives a non-tool response. Even a gibberish message still produced a tool call.
A named tool, {"type": "tool", "name": ...}, takes the choice away completely. You must set type to tool and also give the tool's name. In the cookbook, once print_sentiment_scores was forced, Claude called it for every tweet, including one written to tempt it toward the calculator. Forcing is also how you control order in a multi-step pipeline. If a step such as extract_metadata must run before any enrichment tools, force that tool on the turn where it has to happen, then use a less restrictive tool_choice on later turns.
A team is building a pipeline that asks Claude to read free-form support tickets and return a JSON object with fields like priority, category, and summary. Early prototypes used a prompt asking Claude to "reply with only JSON", but downstream parsing occasionally failed on malformed brackets and stray commentary text. Which approach most reliably eliminates these JSON syntax failures?
Correct answer: A — Define an extraction tool with an input_schema describing the fields, and parse the structured arguments from the resulting tool_use block instead of parsing free text
- A. Correct. Tool use with a JSON schema constrains the model's output to match the schema's structure, guaranteeing well-formed structured data extracted from the tool_use block rather than relying on the model to format free text correctly.
- B. Stronger wording in the system prompt can reduce but not eliminate malformed output, and retry loops add latency and cost without guaranteeing correctness.
- C. Lowering temperature makes responses more consistent but does not enforce a schema; the model can still emit prose, missing brackets, or truncated JSON.
- D. Markdown fencing is a formatting convention, not a schema constraint; the enclosed content itself can still be malformed or contain extra commentary.
A document-processing service receives files that could be invoices, resumes, or contracts, but the type is not known ahead of time. The service defines three separate extraction tools (extract_invoice, extract_resume, extract_contract) and needs Claude to always call exactly one of them so the pipeline never falls back to plain text. Which tool_choice configuration should be used?
Correct answer: A — Set tool_choice to {"type": "any"} so Claude must call one of the three tools but can pick whichever matches the document
- A. Correct. tool_choice: "any" forces Claude to call some tool but leaves the choice of which tool up to the model, which is exactly what's needed when the document type is unknown but a tool call must always happen.
- B. "auto" is the default and permits Claude to respond with plain text instead of calling any tool, which does not guarantee structured output.
- C. Forcing a single named tool would incorrectly run invoice extraction even on resumes or contracts, since the document type is unknown in advance.
- D. Without the tools parameter there is no schema enforcement at all; Claude would be generating free-form text that happens to look like JSON, reintroducing syntax risk.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Once an extraction tool is provided, tool_choice auto is enough to guarantee Claude will call it.Why is that wrong?
auto lets Claude decide whether to call a tool at all, and it may answer in plain text. To guarantee a tool call, use any, or force the tool by name.
Covered in tool_choice: auto, any, or a named tool
2.With several extraction tools and an unknown document type, you have to force one specific tool to be sure you get structured output.Why is that wrong?
any already guarantees a tool call and still lets Claude choose the schema that matches the document. Forcing one named tool would push every document into that single schema.
Covered in tool_choice: auto, any, or a named tool
3.Marking every extraction field as required makes the output more complete and therefore more accurate.Why is that wrong?
Strict mode guarantees required fields are present, so when the source lacks the information the model must invent a value to satisfy the schema. Fields the document may not contain belong outside required, or should be nullable, so absence is a valid answer.
Covered in Defining an extraction tool and reading the tool_use block
4.A strict schema with a date field typed as string ensures every extracted date comes back in the same format.Why is that wrong?
The schema constrains the type and structure, not the formatting of a string's contents. Consistent formatting comes from normalization rules in the prompt, and optionally input_examples, alongside the schema.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Parsing errors from invalid JSON syntax”
↩︎ Why tool use is the structured-output channel“You can use these features independently or together in the same request.”
↩︎ Why tool use is the structured-output channel“errors: list[dict] | None”
↩︎ Defining an extraction tool and reading the tool_use block“Always valid: No more JSON.parse() errors”
↩︎ Strict mode: eliminating syntax and type errors“Type safe: Guaranteed field types and required fields”
↩︎ Exam trap 3 - 2.
“We'll define custom tools that prompt Claude to generate well-structured JSON output”
↩︎ Why tool use is the structured-output channel“we provide an open ended input_schema and instruct Claude via prompting how to interact with the tool”
↩︎ Why tool use is the structured-output channel“we provide an open ended input_schema and instruct Claude via prompting how to interact with the tool”
↩︎ Exam trap 4 - 3.https://platform.claude.com/cookbook/tool-use-tool-choiceSecondary source
“Even though we're forcing Claude to call our print_sentiment_scores tool, we should still employ some basic prompt engineering”
↩︎ Why tool use is the structured-output channel“This is the default behavior when working with tools.”
↩︎ tool_choice: auto, any, or a named tool“The idea is to create a chatbot that always calls one of these tools and never responds with a non-tool response.”
↩︎ tool_choice: auto, any, or a named tool“In addition to setting type to tool, we must provide a particular tool name.”
↩︎ tool_choice: auto, any, or a named tool“auto allows Claude to decide whether to call any provided tools or not”
↩︎ Exam trap 1“any tells Claude that it must use one of the provided tools, but doesn't force a particular tool”
↩︎ Exam trap 2 - 4.
“for tools with complex inputs, nested objects, or format-sensitive parameters, you can use the input_examples field”
↩︎ Why tool use is the structured-output channel“This helps Claude understand when to include optional parameters, what formats to use, and how to structure complex inputs.”
↩︎ Why tool use is the structured-output channel“A JSON Schema object defining the expected parameters for the tool.”
↩︎ Defining an extraction tool and reading the tool_use block“Provide extremely detailed descriptions. This is by far the most important factor in tool performance.”
↩︎ Defining an extraction tool and reading the tool_use block“expects an input object with a required location string and an optional unit string that must be either "celsius" or "fahrenheit"”
↩︎ Defining an extraction tool and reading the tool_use block“What each parameter means and how it affects the tool's behavior”
↩︎ Defining an extraction tool and reading the tool_use block“the API constructs a special system prompt from the tool definitions, tool configuration, and any user-specified system prompt”
↩︎ tool_choice: auto, any, or a named tool - 5.
“The response will have a stop_reason of tool_use and one or more tool_use content blocks”
↩︎ Defining an extraction tool and reading the tool_use block“input: An object containing the input being passed to the tool, conforming to the tool's input_schema.”
↩︎ Key concept - 6.
“Without strict mode, Claude might return incompatible types ("2" instead of 2) or omit required fields”
↩︎ Defining an extraction tool and reading the tool_use block“guarantees Claude's tool inputs match your JSON Schema by constraining the model's token sampling to schema-valid outputs”
↩︎ Strict mode: eliminating syntax and type errors“With strict: true, the response always contains passengers: 2.”
↩︎ Strict mode: eliminating syntax and type errors“Strict tool use compiles tool input_schema definitions into grammars using the same pipeline as structured outputs.”
↩︎ Strict mode: eliminating syntax and type errors“Tool name is always valid (from provided tools or server tools)”
↩︎ tool_choice: auto, any, or a named tool