CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 4 · Lesson 36/56

    Batch inference workloads: when to use ai_query on Databricks

    Identify batch inference workloads and apply ai_query() appropriately

    9 min read
    1.79% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Recognize a batch inference workload and say which Databricks AI Functions apply to it
    • Decide between a task-specific AI Function and the general-purpose ai_query
    • Match each model type ai_query supports to the endpoint it needs
    • Apply the Databricks best practices for running ai_query at scale

    Key concept

    ai_query for batch inference — ai_query is the general-purpose AI Function. It sends every row of a table or query result to a supported model and returns the responses as a column. You use it instead of a purpose-built AI Function when you need control over the model, prompt, parameters or output format.

    1.Recognizing a batch inference workload

    A batch inference workload applies a model to data you already hold, such as a table of news summaries, product reviews or support tickets. You process the whole set and write the results back. A user waiting on one answer is real-time serving, which is a different deployment path. On Databricks, the recommended way to run batch inference is with AI Functions. These are built-in functions that apply LLMs to data stored on Databricks, and you can call them from Databricks SQL, notebooks, Lakeflow pipelines and Workflows.

    AI Functions come in two kinds. Task-specific functions such as ai_translate, ai_summarize, ai_classify and ai_extract each do one job and need no customization. ai_query is the single general-purpose function, and you pick the model and write the prompt yourself. Either kind can run batch inference over a table. Databricks' own batch example uses a task-specific function and notes that you drop the LIMIT to process the whole table.

    Batch inference with a task-specific AI Function: translate a column of summaries row by rowsql
    SELECT
    writer_summary,
      ai_translate(writer_summary, "cn") as cn_translation
    from user.batch.news_summaries
    limit 500
    ;

    Two practical points. First, the query runs on whatever compute you submit it from: a SQL warehouse, notebook, cluster or pipeline. Depending on the function, the model inference itself may run on separate Databricks-managed infrastructure that is billed on top of that compute. Second, AI Functions are not available on Classic SQL warehouses, and they need Databricks Runtime 15.4 LTS or above.

    Checkpoint 1 of 4· Check yourself

    An analyst tries to run an ai_query batch job against a Delta table from a Classic SQL warehouse. What happens?

    Sources12

    2.Task-specific function or ai_query?

    Task-specific functions run research techniques that Databricks maintains, so you don't have to tune a prompt. ai_query trades that convenience for control. The documentation names three situations where ai_query is the right choice.

    When the general-purpose ai_query is justified over a task-specific function
    NeedWhy a task-specific function falls short
    Control the prompt, model parameters, or output format more preciselyTask-specific functions are purpose-built for one task and need no customization
    Query a custom, fine-tuned, or external modelTask-specific functions are powered by Databricks-managed systems you don't choose
    Flexibility to further optimize for throughput or qualityai_query gives full control over the model, prompt, and parameters

    With ai_query, the prompt is usually built by concatenating an instruction with a column, and each row of the result gets its own model response. The example below asks a Databricks-hosted Llama model a question about each taxi trip's pickup ZIP code.

    ai_query with a custom prompt against a Databricks-hosted foundation model, one call per rowsql
    SELECT *,
      ai_query(
        'system.ai.meta-llama-3-3-70b-instruct',
        "Can you tell me the name of the US state that serves the provided ZIP code? zip code: " || pickup_zip
        )
      FROM samples.nyctaxi.trips
      LIMIT 10

    Checkpoint 2 of 4· Exam question

    A team needs to classify the sentiment of 3 million historical support tickets stored in a Delta table. The job runs once nightly, has no user waiting on a response, and results are written back to a new Delta table. Which approach best fits this workload and how should it be implemented?

    Sources3

    3.Which models ai_query can call, and what each needs

    ai_query supports Databricks-hosted models, provisioned throughput models, custom models and external models. The exam tests what you have to set up for each type before the query will run.

    Model types supported by ai_query and their endpoint requirements
    Model typeExamplesWhat you must set up
    Databricks-hosted modelssystem.ai.meta-llama-3-3-70b-instruct, system.ai.gpt-oss-120b, system.ai.gte-large-enNothing: no endpoint provisioning or configuration; Databricks Runtime 15.4 LTS or above
    Provisioned throughput modelsFine-tuned foundation models deployed on Model ServingA provisioned throughput endpoint in Model Serving (AI Functions manages its own scaling for batch)
    External modelsFoundation models hosted outside of DatabricksAn external model serving endpoint
    Custom modelsCustom traditional ML and DL modelsA custom model serving endpoint

    The provisioned throughput row is the one people misread. When ai_query calls a provisioned throughput endpoint in batch, it does not use the compute you provisioned for that endpoint. AI Functions handles batch scaling itself. Also note that not every Databricks-hosted model is a good batch choice. A specific list, including the Llama, GPT, Gemini, Claude, Qwen and gte models, is optimized for batch inference. The other hosted models exist for real-time pay-per-token use and are not recommended for batch at scale. Whatever the type, whoever defines the query needs CAN QUERY permission on the endpoint.

    Checkpoint 3 of 4· Check yourself

    A team has a fine-tuned foundation model and wants to run nightly ai_query batch jobs against it. What is required?

    Sources34

    4.Running ai_query well at scale

    Submit the whole table in one query. AI Functions automatically handle parallelization, retries and scaling, and splitting the data into small batches yourself can reduce throughput. The full set of recommendations for large ai_query jobs:

    Databricks best practices for ai_query batch workloads
    PracticeReason
    Use system.ai model services instead of provisioned throughput endpointsFully managed and scale automatically without provisioning or configuration
    Select a model optimized for batch inferenceA non-optimized model can reduce throughput and lengthen job completion
    Submit your full dataset in a single queryAI Functions handle parallelization, retries, and scaling; manual splitting can reduce throughput
    Set failOnError to false for large workloadsThe job completes, failed rows return error messages, successful results are kept

    Requests to system.ai models are routed through Unity Gateway automatically, but only some gateway features apply. Usage tracking, budget integration with hard spend caps, and permission checks on the model service are applied. Guardrails, inference tables, tracing tables, rate limits, fallbacks and external model providers are not. Model services you create yourself in Unity Gateway are not yet supported by ai_query.

    Checkpoint 4 of 4· Check yourself

    Which feature DOES apply to ai_query requests routed through Unity Gateway to a system.ai model?

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Batch ai_query jobs against a provisioned throughput endpoint run on the throughput you provisioned, so you must size the endpoint for the batch.Why is that wrong?

      AI Functions manages batch scaling itself and does not use the endpoint's provisioned compute.

      Covered in Which models ai_query can call, and what each needs

    2. 2.Chunking a large table into small batches and calling ai_query on each gives better throughput.Why is that wrong?

      AI Functions parallelize, retry and scale automatically, so submit the full dataset in one query. Manual splitting can reduce throughput.

      Covered in Running ai_query well at scale

    3. 3.Because system.ai requests go through Unity Gateway, ai_query batch calls get the gateway's guardrails and inference tables.Why is that wrong?

      Only usage tracking, budget integration and permission checks apply. Guardrails, inference tables, tracing tables, rate limits and fallbacks do not.

      Covered in Running ai_query well at scale

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “You can run batch inference using task-specific AI functions or the general purpose function, ai_query.”
      ↩︎ Recognizing a batch inference workload
    2. 2.
      “the model inference may run on separate Databricks-managed infrastructure that is billed in addition to that compute.”
      ↩︎ Recognizing a batch inference workload
    3. 3.
      “Use ai_query when a task-specific function doesn't meet your needs.”
      ↩︎ Task-specific function or ai_query?
      “Other Databricks-hosted models are available for use with AI Functions, but are not recommended for batch inference production workflows at scale.”
      ↩︎ Which models ai_query can call, and what each needs
      “Submit your full dataset in a single query.”
      ↩︎ Running ai_query well at scale
      “Use system.ai model services instead of provisioned throughput endpoints.”
      ↩︎ Running ai_query well at scale
      “AI Functions does not use the compute provisioned to the endpoint.”
      ↩︎ Exam trap 1
      “Manually splitting data into small batches can reduce throughput.”
      ↩︎ Exam trap 2
      “Not supported: guardrails (service policies), inference tables, tracing tables, rate limits, fallbacks, and external model providers.”
      ↩︎ Exam trap 3
      “This function is not available on Classic SQL warehouses.”
      ↩︎ Checkpoint
      “Databricks recommends starting with a task-specific AI Function when one matches your objective.”
      ↩︎ Prediction
      “AI Functions does not use the compute provisioned to the endpoint.”
      ↩︎ Checkpoint
      “Supported: usage tracking, budget integration (including hard spend caps), and permission checks on the model service.”
      ↩︎ Checkpoint
    4. 4.
      “The definer must have CAN QUERY permission on the endpoint.”
      ↩︎ Which models ai_query can call, and what each needs
      “ai_query is a general purpose AI Function that enables you to query existing endpoints for real-time inference or batch inference workloads.”
      ↩︎ Key concept

    Continue to page 2 of 2

    Applying ai_query: arguments, error handling and batch pipelines

    Spotted a mistake, or was something unclear? Tell us.