What you will be able to do
- Write a prompt template that states the output format explicitly, including a system prompt and few-shot examples
- Choose between json_schema, json_object and an unconstrained response_format for a given requirement
- Design a JSON schema that stays within the subset Foundation Model APIs support, including the extra limits on Claude models
- Enforce an output schema from SQL with ai_query's responseFormat, and say when Databricks prefers structured outputs over function calling
- Check format compliance across prompt versions with a custom judge
Key concept
Structured outputs (response_format) — Prompt wording asks the model for a format. A response_format on the chat request makes the model return JSON, optionally checked against a schema you supply, and Databricks translates that one request shape for every supported model provider.
1.Start with the prompt: state the format you want
There are two ways to get a formatted response from an LLM. You can ask for the format in the prompt, or you can enforce it on the request. This lesson covers the prompt first, because every other technique builds on it. Even when you enforce a JSON schema later on, the instructions in the prompt still tell the model what goes into each field.
The MLflow Prompt Registry docs show the difference with two versions of the same summarization prompt. Version 1 is a single line: Summarize this text: {{content}}. It sets no length, structure or tone, so the model picks its own. Version 2 adds a role, a hard length requirement that it states twice, and a bulleted list of guidelines. It also ends with a Summary: cue, so the response begins right where the format begins.
prompt_v2 = mlflow.genai.register_prompt(
name=PROMPT_NAME,
template="""You are an expert summarizer. Create a summary of the following content in *exactly* 2 sentences (no more, no less - be very careful about the number of sentences).
Guidelines:
- Include ALL core facts and key findings
- Use clear, concise language
- Maintain factual accuracy
- Cover all main points mentioned
- Write for a general audience
- Use exactly 2 sentences
Content: {{content}}
Summary:""",
commit_message="v2: Added comprehensive fact coverage with 2-sentence requirement"
)The format can also be a template variable. Another registry example asks the model to condense content "into exactly {{ num_sentences }} clear and informative sentences", so one registered template can serve callers who need different lengths. Databricks' guidance for Genie Code follows the same rule: say how much detail you want and what structure the answer should have, for example "Provide instructions in numbered steps".
A chat request also has a system message, which is the natural place for the role and the conversion task. The structured-outputs examples put the extraction role in the system message and the raw text in the user message:
messages = [{
"role": "system",
"content": "You are an expert at structured data extraction. You will be given unstructured text from a research paper and should convert it into the given structure."
},
{
"role": "user",
"content": "..."
}]Describing a format in words has limits. Few-shot examples show the format instead. Databricks' RAG quality guidance recommends putting well-formed example queries and their ideal responses inside the prompt template, because the examples teach the model the format, style and content you expect. Choose examples from the query types your application struggles with, and make them representative of real traffic.
Checkpoint 1 of 8· Check yourself
Your support-bot answers are accurate, but their layout changes from one answer to the next. You want to add something to the prompt template that shows the model exactly what a finished answer looks like. What does Databricks' guidance recommend?
Few-shot examples are the technique the guidance names for teaching the desired format, style and content of responses. The other options either leave the format unspecified or have no support in the sources.
“This helps the model understand the desired format, style, and content of the responses.”Source: docs.databricks.com
Checkpoint 2 of 8· Exam question
A team is building a pipeline that sends customer support tickets to an LLM and expects a JSON object with the keys `customer_name`, `issue_category`, and `urgency` so a downstream Python script can parse the result reliably. The zero-shot instruction "Extract the customer name, issue category, and urgency as JSON" sometimes produces valid JSON and sometimes produces prose with an embedded JSON snippet. Which change to the prompt design would most reliably fix the downstream parsing failures?
Correct answer: A — Add a short system instruction stating the exact JSON schema and keys required, followed by one or two example tickets paired with correctly formatted JSON output
- A. Stating the exact schema and keys, then showing one or two worked examples of ticket-to-JSON mappings, gives the model both the rule and a concrete pattern to imitate, which is the most reliable way to lock in a specific output format. Few-shot examples paired with explicit schema instructions reduce the chance of stray prose being added before or after the JSON.
- B. Encouraging step-by-step reasoning is useful for improving answer quality on complex problems, but it does not constrain the output structure and can actually introduce more free-form text into the response, worsening the parsing problem here.
- C. Raising temperature increases randomness in word choice and phrasing, which works against consistent formatting rather than improving it, since the model becomes less likely to reproduce the same structural pattern each time.
- D. Duplicating the ticket text gives the model no new information about the desired output shape and does not address why the model sometimes wraps the JSON in explanatory prose.
2.Enforce JSON with response_format
Prompt instructions make the format likely. They do not make it certain. When a program parses the output, as in batch extraction, a data pipeline or any downstream code, you need machine-readable JSON every time. That is the job of structured outputs. You set a response_format field on the chat request, and the request looks the same whichever provider serves the model. Databricks translates it into each provider's own mechanism, so you never write a provider-specific format. Structured outputs work on Foundation Model APIs pay-per-token and provisioned throughput endpoints.
Databricks recommends structured outputs for three kinds of work: extracting data from large numbers of documents (for example, labelling product reviews as negative, positive or neutral), batch inference that needs outputs in a specified format, and turning unstructured data into structured data.
You can choose between three output modes:
| Setting | What the model returns | When to use it |
|---|---|---|
| "type": "json_schema" (with a schema and "strict": True) | A JSON object that follows the schema you supply | You know the fields in advance, e.g. title, authors, abstract, keywords |
| "type": "json_object" | A valid JSON object with no fixed schema | You need JSON but the schema is not known before hand |
| response_format omitted | Text, with no format enforced | Free-form output; the only way to get unconstrained output from Claude models |
In the json_object example from the docs, the prompt and the request setting do separate jobs. The user message names the fields and asks for "a JSON object", and wraps the input in <description> tags so the model can tell instructions from data. The response_format then guarantees the result is JSON.
response_format = {
"type": "json_object",
}
messages = [
{
"role": "user",
"content": "Extract the name, size, price, and color from this product description as a JSON object:\n<description>\nThe SmartHome Mini is a compact smart home assistant available in black or white for only $49.99. It's 5 inches wide.\n</description>"
}]When you know the fields in advance, use json_schema. You give the schema a name (the docs use research_paper_extraction), describe it as an object with typed properties, and set "strict": True. To run the same request on a Claude model, the docs change only the model value, from databricks-gpt-oss-20b to databricks-claude-sonnet-4-5.
Checkpoint 3 of 8· Match them up
Match each response_format choice to the output it produces
Tap a term, then the definition that fits it.
Structured outputs offer three modes: text, unstructured JSON objects, and JSON that follows a specific schema. You get plain text by leaving response_format out.
“You can choose to generate text, unstructured JSON objects, and JSON objects that adhere to a specific JSON schema.”Source: docs.databricks.com
Checkpoint 4 of 8· Exam question
An engineer is assembling a LangChain-style chain: a prompt template feeds an LLM, and the raw text response must become a validated Python object with fields `title`, `author`, and `year` before being written to a table. Which chain component should the engineer add immediately after the LLM call to reliably convert the model's text response into that structured object?
Correct answer: A — An output parser configured with the target schema, which validates and converts the model's response into the structured object
- A. An output parser bound to the expected schema is designed specifically to take the LLM's raw text and validate or coerce it into the structured object the downstream code needs, making it the correct component to place right after the LLM call.
- B. A retriever supplies grounding documents to the LLM before generation and plays no role in converting the LLM's already-generated text into a structured object afterward.
- C. Adding another prompt template would restart the generation step with different phrasing, but it does not perform validation or conversion of the existing response into a typed object.
- D. A memory component tracks prior turns of a conversation for context continuity, which is unrelated to parsing a single response into a structured schema.
Sources4
3.Design a schema the model can actually follow
Foundation Model APIs broadly accept the structured outputs format that OpenAI accepts. They deliberately support only part of JSON Schema, though, because simpler schemas produce better JSON. Plan around these limits when you design the schema:
- Not supported: regular expressions using pattern; composition and validation using anyOf, oneOf, allOf, prefixItems or $ref; and lists of types, except for the special case [type, "null"].
- Not enforced: length and size keywords such as maxProperties, minProperties and maxLength.
- Capped: a schema can specify at most 64 keys.
- Quality: deeply nested schemas produce worse output, so flatten the schema where you can.
Structured outputs also have a cost. Databricks uses prompt injection and other techniques to improve them, which changes how many input and output tokens you are billed for.
Claude models have further limits. They support only json_schema, not json_object; for unconstrained output you leave response_format out. Structured outputs do not work with streaming, so stream must be false. And on Claude, response_format cannot be combined with tools or tool_choice.
Checkpoint 5 of 8· Check yourself
A pipeline runs on databricks-gpt-oss-20b with response_format {"type": "json_object"} and streaming turned off. You switch the model to databricks-claude-sonnet-4-5 and change nothing else. What must you change?
Databricks does translate the request shape across providers, but Claude accepts only the json_schema type. Streaming must stay off, and tools or tool_choice cannot be combined with response_format on Claude.
“Only the json_schema structured output type is supported. json_object is not supported. For unconstrained output, omit response_format.”Source: docs.databricks.com
Sources4
4.The same enforcement from SQL, and when not to use function calling
In batch work the prompt often runs from SQL. Databricks recommends starting with a task-specific AI Function when one fits your goal. Use ai_query when you need more precise control of the prompt, the model parameters or the output format. ai_query takes a responseFormat argument that enforces a schema so downstream processing gets predictable output. You can write that argument as a JSON schema like the one above, or more compactly as a DDL-style type string:
Checkpoint 6 of 8· Fill the gap
Which named argument makes this ai_query call enforce the DDL-style schema?
SELECT ai_query(
"system.ai.gpt-oss-20b",
"Extract research paper details from the following abstract: " || abstract,
? => 'STRUCT<research_paper_extraction:STRUCT<title:STRING, authors:ARRAY<STRING>, abstract:STRING, keywords:ARRAY<STRING>>>'
)
FROM research_papers;In SQL the argument is responseFormat, in camelCase. response_format is the field name in the chat REST/Python request, and failOnError handles errors on individual rows.
Source: docs.databricks.comWith either form, the docs show output like { "title": "Understanding AI Functions in Databricks", "authors": ["Alice Smith", "Bob Jones"], ... }. That is one JSON object per row, ready to parse.
Function calling also produces structured JSON, and this confuses some candidates. With function calling, you describe functions in the tools parameter, and each function's arguments are a JSON schema. The model does not run those functions. It returns a JSON object of arguments, and your code calls the function. Function calling is for a model that has to decide whether to use an API, such as get_customers(min_revenue: int, created_before: string, limit: int). For batch inference or for turning unstructured data into structured data, Databricks recommends structured outputs instead.
Checkpoint 7 of 8· Put it in order
Put the basic function-calling sequence on Databricks in order
- 1.Call the model with the query and a set of functions defined in the tools parameter
- 2.Call the model again with the structured response appended as a new message, so it can summarize the results for the user
- 3.The model decides whether to call a function; if it does, it returns a JSON object that follows your schema
- 4.Parse the JSON in your code and call your function with the arguments it provides
The model only proposes arguments. Your code runs the function, and a second model call turns the result into the answer the user sees.
“Call the model using the submitted query and a set of functions defined in the tools parameter.”Source: docs.databricks.com
5.Measure format compliance across prompt versions
A schema can be enforced by the API. A rule like "exactly 2 sentences" or "answer in Markdown" cannot, so you measure it. The prompt-evaluation tutorial registers both summary prompts, runs each version against an evaluation dataset, and scores the outputs with two scorers: the built-in Correctness scorer and a custom judge that checks only the format rule.
sentence_count_judge = make_judge(
name="sentence_count_compliance",
instructions="""Evaluate if this summary follows the 2-sentence requirement.
Summary: {{ outputs }}
Count the sentences carefully. Return true if the summary has exactly 2 sentences, and false otherwise.""",
feedback_value_type=bool,
)MLflow averages boolean feedback, so each prompt version ends up with a pass rate (sentence_count_compliance/mean) that you can compare with its correctness score. A related notebook takes this a step further. It starts with the prompt Answer this question: {{question}}, which does not produce Markdown, and uses a judge that rates Markdown quality as high, medium or low. Those ratings become scores of 1.0, 0.5 and 0.0, and a prompt optimizer uses them to rewrite the prompt until the output follows the format.
Checkpoint 8 of 8· Check yourself
You are writing a make_judge scorer that checks whether each response follows a required format, and you want a single pass rate to compare prompt versions. What does Databricks recommend?
Boolean feedback is averaged into a pass rate that you can compare across versions. Correctness checks expected facts, not format.
“Databricks recommends feedback_value_type=bool for pass/fail judges.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Putting maxLength or minProperties in the JSON schema makes the model's output respect those limits.Why is that wrong?
Foundation Model APIs do not enforce length or size keywords. Put such limits in the prompt and check them afterwards.
2.Because Databricks translates response_format for every provider, a Claude model can stream structured output and use tools in the same request.Why is that wrong?
On Claude, structured outputs need stream set to false, and response_format cannot be combined with tools or tool_choice.
3.Function calling is the best way to turn a large batch of unstructured documents into structured JSON.Why is that wrong?
Function calling is for a model that decides whether to call your APIs. For batch inference and converting unstructured data into structured data, Databricks recommends structured outputs.
Covered in The same enforcement from SQL, and when not to use function calling
4.A richer schema, with anyOf branches, $ref reuse and deep nesting, gives the model more guidance and produces better JSON.Why is that wrong?
anyOf, oneOf, allOf, prefixItems and $ref are not supported, and heavily nested schemas lower output quality. Flatten the schema where you can.
Practise it for real
Run ai_query from SQL and enforce an output schema on the result, first as a DDL type string and then as a JSON schema
1.On a SQL warehouse or cluster (not a Classic SQL warehouse) running Databricks Runtime 15.4 LTS or above, run the documented ai_query against system.ai.gpt-oss-20b with responseFormat => 'STRUCT<research_paper_extraction:STRUCT<title:STRING, authors:ARRAY<STRING>, abstract:STRING, keywords:ARRAY<STRING>>>' over a table that has an abstract column.
Why: The DDL form is the most compact way to declare an output schema from SQL.
You should see: One JSON string per row containing title, authors, abstract and keywords, like the example output in the ai_query docs.
2.Replace the DDL string with the documented JSON schema responseFormat: type json_schema, name research_paper_extraction, the same four properties, and strict true.
Why: This is the same schema in the json_schema form that the chat API uses, so you can see the two forms are interchangeable.
You should see: The output has the same shape as in step 1.
3.Add failOnError => false to the call.
Why: On large workloads, one failed row should not stop the whole query.
You should see: The query finishes and returns error messages for failed rows next to the successful results.
Stuck? Get a nudge
If you are working in Python instead of SQL, remember the parameter there is response_format, not responseFormat.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Include examples of well-formed queries and their corresponding ideal responses within the prompt template itself (few-shot learning).”
↩︎ Start with the prompt: state the format you want“This helps the model understand the desired format, style, and content of the responses.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/genie-code/tipsOfficial docs
“Specify the structure of the response you want.”
↩︎ Start with the prompt: state the format you want - 3.https://docs.databricks.com/aws/en/mlflow3/genai/prompt-version-mgmt/prompt-registry/create-and-edit-promptsOfficial docs
“Condense the following content into exactly {{ num_sentences }} clear and informative sentences that capture the key points.”
↩︎ Start with the prompt: state the format you want - 4.
“Databricks handles any provider-specific translation for you, so you do not need to use a provider's native structured outputs format.”
↩︎ Enforce JSON with response_format“Batch inference tasks that require outputs to be in a specified format.”
↩︎ Enforce JSON with response_format“The following is an example of JSON extraction, but the JSON schema is not known before hand.”
↩︎ Enforce JSON with response_format“using a simpler JSON schema for JSON schema definitions results in higher quality JSON generation.”
↩︎ Design a schema the model can actually follow“The maximum number of keys specified in the JSON schema is 64.”
↩︎ Design a schema the model can actually follow“Structured outputs are not supported with streaming. Set stream to false when you specify a response_format.”
↩︎ Design a schema the model can actually follow“Structured outputs on Databricks let you generate responses in a defined JSON format as part of your AI application workflows.”
↩︎ Key concept“Foundation Model APIs does not enforce length or size constraints for objects and arrays.”
↩︎ Exam trap 1“The response_format parameter for Claude structured outputs cannot be combined with tools or tool_choice.”
↩︎ Exam trap 2“Heavily nested JSON schemas result in lower quality generation. If possible, try flattening the JSON schema for better results.”
↩︎ Exam trap 4“You can choose to generate text, unstructured JSON objects, and JSON objects that adhere to a specific JSON schema.”
↩︎ Checkpoint“Foundation Model APIs does not enforce length or size constraints for objects and arrays.”
↩︎ Prediction“Only the json_schema structured output type is supported. json_object is not supported. For unconstrained output, omit response_format.”
↩︎ Checkpoint - 5.
“Ensure that the output conforms to a specific schema for easier downstream processing using responseFormat.”
↩︎ The same enforcement from SQL, and when not to use function calling - 6.
“Control the prompt, model parameters, or output format more precisely”
↩︎ The same enforcement from SQL, and when not to use function calling - 7.
“it creates a JSON object that users can use to call the functions in their code”
↩︎ The same enforcement from SQL, and when not to use function calling“For batch inference or data processing tasks, like converting unstructured data into structured data. Databricks recommends using structured outputs.”
↩︎ Exam trap 3“Call the model using the submitted query and a set of functions defined in the tools parameter.”
↩︎ Checkpoint - 8.https://docs.databricks.com/aws/en/mlflow3/genai/prompt-version-mgmt/prompt-registry/evaluate-promptsOfficial docs
“Databricks recommends feedback_value_type=bool for pass/fail judges.”
↩︎ Measure format compliance across prompt versions - 9.
“The notebook walks you through a markdown judge that optimizes a prompt to output in a more markdown format.”
↩︎ Measure format compliance across prompt versions