CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 4 · Lesson 36/56

    Applying ai_query: arguments, error handling and batch pipelines

    Identify batch inference workloads and apply ai_query() appropriately

    11 min read
    1.79% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Write ai_query calls for foundation-model and custom-model endpoints
    • Use failOnError, modelParameters and responseFormat correctly in batch jobs
    • Deploy ai_query in a Lakeflow pipeline, a scheduled workflow or Structured Streaming
    • Monitor completed and failed inferences with the query profile

    1.ai_query call forms and arguments

    ai_query calls a Databricks Model Serving endpoint and returns the parsed response. Which arguments you pass depends on what the endpoint serves. For a foundation model endpoint, and for a custom model endpoint that has a model schema, the call is ai_query(endpoint, request). For a custom model endpoint without a model schema, the full form is ai_query(endpoint, request, returnType, failOnError).

    Core ai_query arguments and what they accept
    ArgumentMeaning
    endpointSTRING literal naming a Foundation Model, external model, or custom model endpoint in the same workspace
    requestA STRING for external or Foundation Model APIs endpoints; a single column or STRUCT for custom model endpoints
    returnTypeExpected return type; optional from Databricks Runtime 15.2, when it is inferred from the custom endpoint's model schema
    failOnErrorBoolean, defaults to true; when false, returns a STRUCT of response and error message
    modelParametersStruct of chat, completion, or embedding parameters such as max_tokens and temperature
    responseFormatOutput schema for chat foundation models: DDL style or JSON string

    A custom model is usually a traditional ML model that expects named features. You pass it a STRUCT whose field names match those features, and you set returnType so Spark knows the result type.

    Batch scoring a custom classification endpoint with a named STRUCT request and a BOOLEAN returnTypesql
    SELECT text, ai_query(
        endpoint => 'spam-classification-endpoint',
        request => named_struct(
          'timestamp', timestamp,
          'sender', from_number,
          'text', text),
        returnType => 'BOOLEAN') AS is_spam
      FROM messages
      LIMIT 10

    Checkpoint 1 of 5· Match them up

    Match each endpoint situation to the request form ai_query accepts

    Tap a term, then the definition that fits it.

    Sources1

    2.failOnError, modelParameters and responseFormat

    For large workloads, set failOnError => false. The column then holds a STRUCT containing the parsed response and an errorMessage string. A successful row has a null errorMessage. A row that failed because of a model endpoint error has a null response. The guarantee is narrower than it sounds, though: an error from any other cause still fails the whole query.

    Keeping a summarization job running when individual rows failsql
    SELECT text, ai_query(
        "system.ai.meta-llama-3-3-70b-instruct",
        "Summarize the given text comprehensively, covering key points and main ideas concisely while retaining relevant details and examples. Ensure clarity and accuracy without unnecessary repetition or omissions: " || text,
    failOnError => false
    ) AS summary
    FROM uc_catalog.schema.table;

    modelParameters controls generation settings such as max_tokens and temperature. The values must be constants, so you cannot compute a different temperature for each row. Any parameter you leave out takes its default. In ai_query, temperature defaults to 0.0; every other default matches the Foundation Model REST API reference.

    Checkpoint 2 of 5· Fill the gap

    Which argument name completes this call that caps output length and sets temperature?

    SELECT text, ai_query(
        "system.ai.meta-llama-3-3-70b-instruct",
        "Please summarize the following article: " || text,
         ?  => named_struct('max_tokens', 100, 'temperature', 0.7)
    ) AS summary
    FROM uc_catalog.schema.table;

    responseFormat makes a chat foundation model return output in a fixed schema, which makes downstream parsing much simpler. It works only with chat foundation models and accepts either a DDL-style string or a JSON string of type text, json_object or json_schema.

    Enforcing a DDL-style output schema for extractionsql
    SELECT ai_query(
        "system.ai.gpt-oss-20b",
        "Extract research paper details from the following abstract: " || abstract,
        responseFormat => 'STRUCT<research_paper_extraction:STRUCT<title:STRING, authors:ARRAY<STRING>, abstract:STRING, keywords:ARRAY<STRING>>>'
    )
    FROM research_papers;

    Checkpoint 3 of 5· Check yourself

    An engineer wants each row's temperature computed from a column called creativity_score. Why will this not work as intended?

    Sources1

    3.Putting ai_query into production pipelines

    A one-off SELECT is fine for exploration. Production batch inference wraps ai_query in a pipeline that handles ingestion, preprocessing, inference and post-processing. You can write these pipelines in SQL or Python and deploy them in three ways. Lakeflow pipelines suit data that keeps arriving and needs incremental processing. Scheduled Databricks workflows suit recurring jobs, such as a nightly opinion-mining query over product reviews. Structured Streaming suits near-real-time or micro-batch inference.

    The Lakeflow example in the documentation chains four steps. A streaming table news_raw ingests CSV files from a volume. A materialized view news_categorized calls ai_query with a json_schema responseFormat to extract title and category. news_validated applies expectations: the title must have at least three words, and the category must be one of the allowed values. Finally, news_summarized calls ai_query again to summarize the validated articles.

    Checkpoint 4 of 5· Put it in order

    Put the stages of the Lakeflow news pipeline in order

    1. 1.Ingest raw CSV articles into the news_raw streaming table
    2. 2.Summarize the validated articles with ai_query into news_summarized
    3. 3.Validate title length and category with expectations in news_validated
    4. 4.Extract category and title with ai_query and responseFormat into news_categorized

    With Structured Streaming you read a Delta table as a stream and apply ai_query inside F.expr. For a static table, the first micro-batch processes every existing row exactly once, and later triggers process only rows added after that. The output is written to another Delta table, with a checkpoint location.

    Applying ai_query to a Delta stream with Structured Streamingpython
    df_transformed = df_stream.select(
        "document_text",
        F.expr(f"""
          ai_query(
            'system.ai.meta-llama-3-1-8b-instruct',
            CONCAT('{prompt}', document_text)
          )
        """).alias("summary")
    )

    There is also a governance pattern. You can wrap an ai_query call in a SQL UDF such as correct_grammar and grant EXECUTE on it to a team. That team can then fix grammar across a whole table in batch without handling the endpoint name or the prompt themselves.

    Sources21

    4.Monitoring batch inference progress

    Once a large job is running, the query profile shows how many inferences have completed or failed. In Databricks Runtime 16.1 ML and above, open it from the SQL editor: select the Running link below the raw results, click See query profile, then click AI Query. This shows the number of completed and failed inferences and the total request time. Combined with failOnError => false, it tells you how many rows to reprocess and spares you from rerunning the whole dataset.

    Checkpoint 5 of 5· Check yourself

    A nightly ai_query job seems slow and you suspect many rows are failing. Where do you see counts of completed and failed inferences?

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.With failOnError => false, a batch query can never fail because of a bad row.Why is that wrong?

      Only model endpoint errors are captured per row. A row that fails for any other reason still fails the whole query.

      Covered in failOnError, modelParameters and responseFormat

    2. 2.Omitting temperature in ai_query gives the same default as the Foundation Model REST API.Why is that wrong?

      ai_query sets temperature to 0.0 by default. Only the other parameters follow the REST API defaults.

      Covered in failOnError, modelParameters and responseFormat

    Practise it for real

    Run a fault-tolerant ai_query batch job with a controlled output length and inspect its results and metrics

    1. 1.In the SQL editor, run ai_query against a Unity Catalog table using system.ai.meta-llama-3-3-70b-instruct, a prompt concatenated with a text column, and modelParameters => named_struct('max_tokens', 100, 'temperature', 0.7).

      Why: This confirms you can reach a Databricks-hosted model with no endpoint setup and pass constant generation parameters.

      You should see: A summary column holding one model response per input row.

    2. 2.Rerun the query with failOnError => false added.

      Why: Large jobs should keep successful rows when individual rows fail.

      You should see: The summary column becomes a STRUCT of response and errorMessage, with errorMessage null on rows that succeeded.

    3. 3.Open the query profile from the Running link, click See query profile, then AI Query.

      Why: This is where you check batch progress and failure counts.

      You should see: Counts of completed and failed inferences and the total request time.

    Stuck? Get a nudge

    If ai_query is rejected outright, check whether you are on a Classic SQL warehouse or a runtime below 15.4 LTS.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “For Databricks Runtime 15.1 and below, this expression is required for querying a custom model serving endpoint.”
      ↩︎ ai_query call forms and arguments
      “If failOnError => false, the function returns a STRUCT object that contains the parsed response and the error message string.”
      ↩︎ failOnError, modelParameters and responseFormat
      “Only available for querying chat foundation models.”
      ↩︎ failOnError, modelParameters and responseFormat
      “Optionally, you can also wrap a call to ai_query() in a UDF for function calling as follows:”
      ↩︎ Putting ai_query into production pipelines
      “If the inference of the row fails due to other errors, the whole query fails.”
      ↩︎ Exam trap 1
      “With the exception of temperature which has a default value of 0.0”
      ↩︎ Exam trap 2
      “If the endpoint is a custom model serving endpoint, the request can be a single column or a STRUCT expression.”
      ↩︎ Checkpoint
      “A boolean literal that defaults to true.”
      ↩︎ Prediction
      “These model parameters must be constant parameters and not data dependent.”
      ↩︎ Checkpoint
    2. 2.
      “These pipelines can perform end-to-end workflows that include ingestion, preprocessing, inference, and post-processing.”
      ↩︎ Putting ai_query into production pipelines
      “Apply AI inference in near real-time or micro-batch scenarios using ai_query and Structured Streaming.”
      ↩︎ Putting ai_query into production pipelines
      “Summarized political news articles after validation.”
      ↩︎ Checkpoint
    3. 3.
      “you can monitor the progress of AI Functions using the query profile feature.”
      ↩︎ Monitoring batch inference progress
      “Click AI Query to see metrics for that particular query including the number of completed and failed inferences”
      ↩︎ Checkpoint

    Spotted a mistake, or was something unclear? Tell us.