CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 1 · Lesson 1/56

    Prompting for formatted LLM output: instructions, few-shot examples and structured outputs

    Design a prompt that elicits a specifically formatted response

    17 min read
    1.79% of exam
    9 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Write a prompt template that states the output format explicitly, including a system prompt and few-shot examples
    • Choose between json_schema, json_object and an unconstrained response_format for a given requirement
    • Design a JSON schema that stays within the subset Foundation Model APIs support, including the extra limits on Claude models
    • Enforce an output schema from SQL with ai_query's responseFormat, and say when Databricks prefers structured outputs over function calling
    • Check format compliance across prompt versions with a custom judge

    Key concept

    Structured outputs (response_format) — Prompt wording asks the model for a format. A response_format on the chat request makes the model return JSON, optionally checked against a schema you supply, and Databricks translates that one request shape for every supported model provider.

    1.Start with the prompt: state the format you want

    There are two ways to get a formatted response from an LLM. You can ask for the format in the prompt, or you can enforce it on the request. This lesson covers the prompt first, because every other technique builds on it. Even when you enforce a JSON schema later on, the instructions in the prompt still tell the model what goes into each field.

    The MLflow Prompt Registry docs show the difference with two versions of the same summarization prompt. Version 1 is a single line: Summarize this text: {{content}}. It sets no length, structure or tone, so the model picks its own. Version 2 adds a role, a hard length requirement that it states twice, and a bulleted list of guidelines. It also ends with a Summary: cue, so the response begins right where the format begins.

    Prompt version 2: a role, an exact sentence count stated twice, explicit guidelines and a trailing output cuepython
    prompt_v2 = mlflow.genai.register_prompt(
        name=PROMPT_NAME,
        template="""You are an expert summarizer. Create a summary of the following content in *exactly* 2 sentences (no more, no less - be very careful about the number of sentences).
    
    Guidelines:
    - Include ALL core facts and key findings
    - Use clear, concise language
    - Maintain factual accuracy
    - Cover all main points mentioned
    - Write for a general audience
    - Use exactly 2 sentences
    
    Content: {{content}}
    
    Summary:""",
        commit_message="v2: Added comprehensive fact coverage with 2-sentence requirement"
    )

    The format can also be a template variable. Another registry example asks the model to condense content "into exactly {{ num_sentences }} clear and informative sentences", so one registered template can serve callers who need different lengths. Databricks' guidance for Genie Code follows the same rule: say how much detail you want and what structure the answer should have, for example "Provide instructions in numbered steps".

    A chat request also has a system message, which is the natural place for the role and the conversion task. The structured-outputs examples put the extraction role in the system message and the raw text in the user message:

    A system message that sets the role and the conversion task, with the user message carrying the inputpython
    messages = [{
            "role": "system",
            "content": "You are an expert at structured data extraction. You will be given unstructured text from a research paper and should convert it into the given structure."
          },
          {
            "role": "user",
            "content": "..."
          }]

    Describing a format in words has limits. Few-shot examples show the format instead. Databricks' RAG quality guidance recommends putting well-formed example queries and their ideal responses inside the prompt template, because the examples teach the model the format, style and content you expect. Choose examples from the query types your application struggles with, and make them representative of real traffic.

    Checkpoint 1 of 8· Check yourself

    Your support-bot answers are accurate, but their layout changes from one answer to the next. You want to add something to the prompt template that shows the model exactly what a finished answer looks like. What does Databricks' guidance recommend?

    Checkpoint 2 of 8· Exam question

    A team is building a pipeline that sends customer support tickets to an LLM and expects a JSON object with the keys `customer_name`, `issue_category`, and `urgency` so a downstream Python script can parse the result reliably. The zero-shot instruction "Extract the customer name, issue category, and urgency as JSON" sometimes produces valid JSON and sometimes produces prose with an embedded JSON snippet. Which change to the prompt design would most reliably fix the downstream parsing failures?

    Sources123

    2.Enforce JSON with response_format

    Prompt instructions make the format likely. They do not make it certain. When a program parses the output, as in batch extraction, a data pipeline or any downstream code, you need machine-readable JSON every time. That is the job of structured outputs. You set a response_format field on the chat request, and the request looks the same whichever provider serves the model. Databricks translates it into each provider's own mechanism, so you never write a provider-specific format. Structured outputs work on Foundation Model APIs pay-per-token and provisioned throughput endpoints.

    Databricks recommends structured outputs for three kinds of work: extracting data from large numbers of documents (for example, labelling product reviews as negative, positive or neutral), batch inference that needs outputs in a specified format, and turning unstructured data into structured data.

    You can choose between three output modes:

    The three output modes you can choose with response_format
    SettingWhat the model returnsWhen to use it
    "type": "json_schema" (with a schema and "strict": True)A JSON object that follows the schema you supplyYou know the fields in advance, e.g. title, authors, abstract, keywords
    "type": "json_object"A valid JSON object with no fixed schemaYou need JSON but the schema is not known before hand
    response_format omittedText, with no format enforcedFree-form output; the only way to get unconstrained output from Claude models

    In the json_object example from the docs, the prompt and the request setting do separate jobs. The user message names the fields and asks for "a JSON object", and wraps the input in <description> tags so the model can tell instructions from data. The response_format then guarantees the result is JSON.

    json_object mode: the prompt names the fields and the request guarantees JSONpython
    response_format = {
          "type": "json_object",
        }
    
    messages = [
          {
            "role": "user",
            "content": "Extract the name, size, price, and color from this product description as a JSON object:\n<description>\nThe SmartHome Mini is a compact smart home assistant available in black or white for only $49.99. It's 5 inches wide.\n</description>"
          }]

    When you know the fields in advance, use json_schema. You give the schema a name (the docs use research_paper_extraction), describe it as an object with typed properties, and set "strict": True. To run the same request on a Claude model, the docs change only the model value, from databricks-gpt-oss-20b to databricks-claude-sonnet-4-5.

    Checkpoint 3 of 8· Match them up

    Match each response_format choice to the output it produces

    Tap a term, then the definition that fits it.

    Checkpoint 4 of 8· Exam question

    An engineer is assembling a LangChain-style chain: a prompt template feeds an LLM, and the raw text response must become a validated Python object with fields `title`, `author`, and `year` before being written to a table. Which chain component should the engineer add immediately after the LLM call to reliably convert the model's text response into that structured object?

    Sources4

    3.Design a schema the model can actually follow

    Foundation Model APIs broadly accept the structured outputs format that OpenAI accepts. They deliberately support only part of JSON Schema, though, because simpler schemas produce better JSON. Plan around these limits when you design the schema:

    - Not supported: regular expressions using pattern; composition and validation using anyOf, oneOf, allOf, prefixItems or $ref; and lists of types, except for the special case [type, "null"]. - Not enforced: length and size keywords such as maxProperties, minProperties and maxLength. - Capped: a schema can specify at most 64 keys. - Quality: deeply nested schemas produce worse output, so flatten the schema where you can.

    Structured outputs also have a cost. Databricks uses prompt injection and other techniques to improve them, which changes how many input and output tokens you are billed for.

    Claude models have further limits. They support only json_schema, not json_object; for unconstrained output you leave response_format out. Structured outputs do not work with streaming, so stream must be false. And on Claude, response_format cannot be combined with tools or tool_choice.

    Checkpoint 5 of 8· Check yourself

    A pipeline runs on databricks-gpt-oss-20b with response_format {"type": "json_object"} and streaming turned off. You switch the model to databricks-claude-sonnet-4-5 and change nothing else. What must you change?

    Sources4

    4.The same enforcement from SQL, and when not to use function calling

    In batch work the prompt often runs from SQL. Databricks recommends starting with a task-specific AI Function when one fits your goal. Use ai_query when you need more precise control of the prompt, the model parameters or the output format. ai_query takes a responseFormat argument that enforces a schema so downstream processing gets predictable output. You can write that argument as a JSON schema like the one above, or more compactly as a DDL-style type string:

    Checkpoint 6 of 8· Fill the gap

    Which named argument makes this ai_query call enforce the DDL-style schema?

    SELECT ai_query(
        "system.ai.gpt-oss-20b",
        "Extract research paper details from the following abstract: " || abstract,
         ?  => 'STRUCT<research_paper_extraction:STRUCT<title:STRING, authors:ARRAY<STRING>, abstract:STRING, keywords:ARRAY<STRING>>>'
    )
    FROM research_papers;

    With either form, the docs show output like { "title": "Understanding AI Functions in Databricks", "authors": ["Alice Smith", "Bob Jones"], ... }. That is one JSON object per row, ready to parse.

    Function calling also produces structured JSON, and this confuses some candidates. With function calling, you describe functions in the tools parameter, and each function's arguments are a JSON schema. The model does not run those functions. It returns a JSON object of arguments, and your code calls the function. Function calling is for a model that has to decide whether to use an API, such as get_customers(min_revenue: int, created_before: string, limit: int). For batch inference or for turning unstructured data into structured data, Databricks recommends structured outputs instead.

    Checkpoint 7 of 8· Put it in order

    Put the basic function-calling sequence on Databricks in order

    1. 1.Call the model with the query and a set of functions defined in the tools parameter
    2. 2.Call the model again with the structured response appended as a new message, so it can summarize the results for the user
    3. 3.The model decides whether to call a function; if it does, it returns a JSON object that follows your schema
    4. 4.Parse the JSON in your code and call your function with the arguments it provides

    Sources567

    5.Measure format compliance across prompt versions

    A schema can be enforced by the API. A rule like "exactly 2 sentences" or "answer in Markdown" cannot, so you measure it. The prompt-evaluation tutorial registers both summary prompts, runs each version against an evaluation dataset, and scores the outputs with two scorers: the built-in Correctness scorer and a custom judge that checks only the format rule.

    A pass/fail judge for one format requirementpython
    sentence_count_judge = make_judge(
        name="sentence_count_compliance",
        instructions="""Evaluate if this summary follows the 2-sentence requirement.
    
    Summary: {{ outputs }}
    
    Count the sentences carefully. Return true if the summary has exactly 2 sentences, and false otherwise.""",
        feedback_value_type=bool,
    )

    MLflow averages boolean feedback, so each prompt version ends up with a pass rate (sentence_count_compliance/mean) that you can compare with its correctness score. A related notebook takes this a step further. It starts with the prompt Answer this question: {{question}}, which does not produce Markdown, and uses a judge that rates Markdown quality as high, medium or low. Those ratings become scores of 1.0, 0.5 and 0.0, and a prompt optimizer uses them to rewrite the prompt until the output follows the format.

    Checkpoint 8 of 8· Check yourself

    You are writing a make_judge scorer that checks whether each response follows a required format, and you want a single pass rate to compare prompt versions. What does Databricks recommend?

    Sources89

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Putting maxLength or minProperties in the JSON schema makes the model's output respect those limits.Why is that wrong?

      Foundation Model APIs do not enforce length or size keywords. Put such limits in the prompt and check them afterwards.

      Covered in Design a schema the model can actually follow

    2. 2.Because Databricks translates response_format for every provider, a Claude model can stream structured output and use tools in the same request.Why is that wrong?

      On Claude, structured outputs need stream set to false, and response_format cannot be combined with tools or tool_choice.

      Covered in Design a schema the model can actually follow

    3. 3.Function calling is the best way to turn a large batch of unstructured documents into structured JSON.Why is that wrong?

      Function calling is for a model that decides whether to call your APIs. For batch inference and converting unstructured data into structured data, Databricks recommends structured outputs.

      Covered in The same enforcement from SQL, and when not to use function calling

    4. 4.A richer schema, with anyOf branches, $ref reuse and deep nesting, gives the model more guidance and produces better JSON.Why is that wrong?

      anyOf, oneOf, allOf, prefixItems and $ref are not supported, and heavily nested schemas lower output quality. Flatten the schema where you can.

      Covered in Design a schema the model can actually follow

    Practise it for real

    Run ai_query from SQL and enforce an output schema on the result, first as a DDL type string and then as a JSON schema

    1. 1.On a SQL warehouse or cluster (not a Classic SQL warehouse) running Databricks Runtime 15.4 LTS or above, run the documented ai_query against system.ai.gpt-oss-20b with responseFormat => 'STRUCT<research_paper_extraction:STRUCT<title:STRING, authors:ARRAY<STRING>, abstract:STRING, keywords:ARRAY<STRING>>>' over a table that has an abstract column.

      Why: The DDL form is the most compact way to declare an output schema from SQL.

      You should see: One JSON string per row containing title, authors, abstract and keywords, like the example output in the ai_query docs.

    2. 2.Replace the DDL string with the documented JSON schema responseFormat: type json_schema, name research_paper_extraction, the same four properties, and strict true.

      Why: This is the same schema in the json_schema form that the chat API uses, so you can see the two forms are interchangeable.

      You should see: The output has the same shape as in step 1.

    3. 3.Add failOnError => false to the call.

      Why: On large workloads, one failed row should not stop the whole query.

      You should see: The query finishes and returns error messages for failed rows next to the successful results.

    Stuck? Get a nudge

    If you are working in Python instead of SQL, remember the parameter there is response_format, not responseFormat.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Include examples of well-formed queries and their corresponding ideal responses within the prompt template itself (few-shot learning).”
      ↩︎ Start with the prompt: state the format you want
      “This helps the model understand the desired format, style, and content of the responses.”
      ↩︎ Checkpoint
    2. 3.
      “Condense the following content into exactly {{ num_sentences }} clear and informative sentences that capture the key points.”
      ↩︎ Start with the prompt: state the format you want
    3. 4.
      “Databricks handles any provider-specific translation for you, so you do not need to use a provider's native structured outputs format.”
      ↩︎ Enforce JSON with response_format
      “Batch inference tasks that require outputs to be in a specified format.”
      ↩︎ Enforce JSON with response_format
      “The following is an example of JSON extraction, but the JSON schema is not known before hand.”
      ↩︎ Enforce JSON with response_format
      “using a simpler JSON schema for JSON schema definitions results in higher quality JSON generation.”
      ↩︎ Design a schema the model can actually follow
      “The maximum number of keys specified in the JSON schema is 64.”
      ↩︎ Design a schema the model can actually follow
      “Structured outputs are not supported with streaming. Set stream to false when you specify a response_format.”
      ↩︎ Design a schema the model can actually follow
      “Structured outputs on Databricks let you generate responses in a defined JSON format as part of your AI application workflows.”
      ↩︎ Key concept
      “Foundation Model APIs does not enforce length or size constraints for objects and arrays.”
      ↩︎ Exam trap 1
      “The response_format parameter for Claude structured outputs cannot be combined with tools or tool_choice.”
      ↩︎ Exam trap 2
      “Heavily nested JSON schemas result in lower quality generation. If possible, try flattening the JSON schema for better results.”
      ↩︎ Exam trap 4
      “You can choose to generate text, unstructured JSON objects, and JSON objects that adhere to a specific JSON schema.”
      ↩︎ Checkpoint
      “Foundation Model APIs does not enforce length or size constraints for objects and arrays.”
      ↩︎ Prediction
      “Only the json_schema structured output type is supported. json_object is not supported. For unconstrained output, omit response_format.”
      ↩︎ Checkpoint
    4. 5.
      “Ensure that the output conforms to a specific schema for easier downstream processing using responseFormat.”
      ↩︎ The same enforcement from SQL, and when not to use function calling
    5. 7.
      “it creates a JSON object that users can use to call the functions in their code”
      ↩︎ The same enforcement from SQL, and when not to use function calling
      “For batch inference or data processing tasks, like converting unstructured data into structured data. Databricks recommends using structured outputs.”
      ↩︎ Exam trap 3
      “Call the model using the submitted query and a set of functions defined in the tools parameter.”
      ↩︎ Checkpoint
    6. 9.
      “The notebook walks you through a markdown judge that optimizes a prompt to output in a more markdown format.”
      ↩︎ Measure format compliance across prompt versions

    Spotted a mistake, or was something unclear? Tell us.