What you will be able to do
- Recognize a batch inference workload and say which Databricks AI Functions apply to it
- Decide between a task-specific AI Function and the general-purpose ai_query
- Match each model type ai_query supports to the endpoint it needs
- Apply the Databricks best practices for running ai_query at scale
Key concept
ai_query for batch inference — ai_query is the general-purpose AI Function. It sends every row of a table or query result to a supported model and returns the responses as a column. You use it instead of a purpose-built AI Function when you need control over the model, prompt, parameters or output format.
1.Recognizing a batch inference workload
A batch inference workload applies a model to data you already hold, such as a table of news summaries, product reviews or support tickets. You process the whole set and write the results back. A user waiting on one answer is real-time serving, which is a different deployment path. On Databricks, the recommended way to run batch inference is with AI Functions. These are built-in functions that apply LLMs to data stored on Databricks, and you can call them from Databricks SQL, notebooks, Lakeflow pipelines and Workflows.
AI Functions come in two kinds. Task-specific functions such as ai_translate, ai_summarize, ai_classify and ai_extract each do one job and need no customization. ai_query is the single general-purpose function, and you pick the model and write the prompt yourself. Either kind can run batch inference over a table. Databricks' own batch example uses a task-specific function and notes that you drop the LIMIT to process the whole table.
SELECT
writer_summary,
ai_translate(writer_summary, "cn") as cn_translation
from user.batch.news_summaries
limit 500
;Two practical points. First, the query runs on whatever compute you submit it from: a SQL warehouse, notebook, cluster or pipeline. Depending on the function, the model inference itself may run on separate Databricks-managed infrastructure that is billed on top of that compute. Second, AI Functions are not available on Classic SQL warehouses, and they need Databricks Runtime 15.4 LTS or above.
Checkpoint 1 of 4· Check yourself
An analyst tries to run an ai_query batch job against a Delta table from a Classic SQL warehouse. What happens?
Classic SQL warehouses are explicitly excluded. Use a supported warehouse type or Databricks Runtime 15.4 LTS or above.
“This function is not available on Classic SQL warehouses.”Source: docs.databricks.com
2.Task-specific function or ai_query?
Task-specific functions run research techniques that Databricks maintains, so you don't have to tune a prompt. ai_query trades that convenience for control. The documentation names three situations where ai_query is the right choice.
| Need | Why a task-specific function falls short |
|---|---|
| Control the prompt, model parameters, or output format more precisely | Task-specific functions are purpose-built for one task and need no customization |
| Query a custom, fine-tuned, or external model | Task-specific functions are powered by Databricks-managed systems you don't choose |
| Flexibility to further optimize for throughput or quality | ai_query gives full control over the model, prompt, and parameters |
With ai_query, the prompt is usually built by concatenating an instruction with a column, and each row of the result gets its own model response. The example below asks a Databricks-hosted Llama model a question about each taxi trip's pickup ZIP code.
SELECT *,
ai_query(
'system.ai.meta-llama-3-3-70b-instruct',
"Can you tell me the name of the US state that serves the provided ZIP code? zip code: " || pickup_zip
)
FROM samples.nyctaxi.trips
LIMIT 10Checkpoint 2 of 4· Exam question
A team needs to classify the sentiment of 3 million historical support tickets stored in a Delta table. The job runs once nightly, has no user waiting on a response, and results are written back to a new Delta table. Which approach best fits this workload and how should it be implemented?
Correct answer: A — This is a batch inference workload; call `ai_query()` inside a `SELECT` statement over the Delta table so each row is scored against a model serving endpoint in one distributed job.
- A. Correct: a large, scheduled, non-interactive scoring pass over an existing Delta table is the canonical batch inference pattern, and `ai_query()` lets you invoke a model serving endpoint directly from SQL for every row in a single distributed query.
- B. Provisioned throughput and low-latency REST calls target interactive, per-request serving where a caller is waiting on an immediate response, not a nightly bulk scoring pass over millions of stored rows.
- C. A streaming table reacting to new files is meant for continuously arriving data, but this scenario describes a fixed, already-landed historical table scored on a schedule, not an incremental stream.
- D. Feature Store publishes reusable feature values for training or serving; it does not perform model inference on its own, so it cannot produce the sentiment classifications the ticket asks for.
Sources3
3.Which models ai_query can call, and what each needs
ai_query supports Databricks-hosted models, provisioned throughput models, custom models and external models. The exam tests what you have to set up for each type before the query will run.
| Model type | Examples | What you must set up |
|---|---|---|
| Databricks-hosted models | system.ai.meta-llama-3-3-70b-instruct, system.ai.gpt-oss-120b, system.ai.gte-large-en | Nothing: no endpoint provisioning or configuration; Databricks Runtime 15.4 LTS or above |
| Provisioned throughput models | Fine-tuned foundation models deployed on Model Serving | A provisioned throughput endpoint in Model Serving (AI Functions manages its own scaling for batch) |
| External models | Foundation models hosted outside of Databricks | An external model serving endpoint |
| Custom models | Custom traditional ML and DL models | A custom model serving endpoint |
The provisioned throughput row is the one people misread. When ai_query calls a provisioned throughput endpoint in batch, it does not use the compute you provisioned for that endpoint. AI Functions handles batch scaling itself. Also note that not every Databricks-hosted model is a good batch choice. A specific list, including the Llama, GPT, Gemini, Claude, Qwen and gte models, is optimized for batch inference. The other hosted models exist for real-time pay-per-token use and are not recommended for batch at scale. Whatever the type, whoever defines the query needs CAN QUERY permission on the endpoint.
Checkpoint 3 of 4· Check yourself
A team has a fine-tuned foundation model and wants to run nightly ai_query batch jobs against it. What is required?
Fine-tuned foundation models need a provisioned throughput endpoint, but AI Functions does not run batch work on that endpoint's provisioned compute.
“AI Functions does not use the compute provisioned to the endpoint.”Source: docs.databricks.com
4.Running ai_query well at scale
Submit the whole table in one query. AI Functions automatically handle parallelization, retries and scaling, and splitting the data into small batches yourself can reduce throughput. The full set of recommendations for large ai_query jobs:
| Practice | Reason |
|---|---|
| Use system.ai model services instead of provisioned throughput endpoints | Fully managed and scale automatically without provisioning or configuration |
| Select a model optimized for batch inference | A non-optimized model can reduce throughput and lengthen job completion |
| Submit your full dataset in a single query | AI Functions handle parallelization, retries, and scaling; manual splitting can reduce throughput |
| Set failOnError to false for large workloads | The job completes, failed rows return error messages, successful results are kept |
Requests to system.ai models are routed through Unity Gateway automatically, but only some gateway features apply. Usage tracking, budget integration with hard spend caps, and permission checks on the model service are applied. Guardrails, inference tables, tracing tables, rate limits, fallbacks and external model providers are not. Model services you create yourself in Unity Gateway are not yet supported by ai_query.
Checkpoint 4 of 4· Check yourself
Which feature DOES apply to ai_query requests routed through Unity Gateway to a system.ai model?
Of the Unity Gateway features, only usage tracking, budget integration and permission checks apply to routed ai_query requests.
“Supported: usage tracking, budget integration (including hard spend caps), and permission checks on the model service.”Source: docs.databricks.com
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Batch ai_query jobs against a provisioned throughput endpoint run on the throughput you provisioned, so you must size the endpoint for the batch.Why is that wrong?
AI Functions manages batch scaling itself and does not use the endpoint's provisioned compute.
Covered in Which models ai_query can call, and what each needs
2.Chunking a large table into small batches and calling ai_query on each gives better throughput.Why is that wrong?
AI Functions parallelize, retry and scale automatically, so submit the full dataset in one query. Manual splitting can reduce throughput.
Covered in Running ai_query well at scale
3.Because system.ai requests go through Unity Gateway, ai_query batch calls get the gateway's guardrails and inference tables.Why is that wrong?
Only usage tracking, budget integration and permission checks apply. Guardrails, inference tables, tracing tables, rate limits and fallbacks do not.
Covered in Running ai_query well at scale
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“You can run batch inference using task-specific AI functions or the general purpose function, ai_query.”
↩︎ Recognizing a batch inference workload - 2.
“the model inference may run on separate Databricks-managed infrastructure that is billed in addition to that compute.”
↩︎ Recognizing a batch inference workload - 3.
“Use ai_query when a task-specific function doesn't meet your needs.”
↩︎ Task-specific function or ai_query?“Other Databricks-hosted models are available for use with AI Functions, but are not recommended for batch inference production workflows at scale.”
↩︎ Which models ai_query can call, and what each needs“Submit your full dataset in a single query.”
↩︎ Running ai_query well at scale“Use system.ai model services instead of provisioned throughput endpoints.”
↩︎ Running ai_query well at scale“AI Functions does not use the compute provisioned to the endpoint.”
↩︎ Exam trap 1“Manually splitting data into small batches can reduce throughput.”
↩︎ Exam trap 2“Not supported: guardrails (service policies), inference tables, tracing tables, rate limits, fallbacks, and external model providers.”
↩︎ Exam trap 3“This function is not available on Classic SQL warehouses.”
↩︎ Checkpoint“Databricks recommends starting with a task-specific AI Function when one matches your objective.”
↩︎ Prediction“AI Functions does not use the compute provisioned to the endpoint.”
↩︎ Checkpoint“Supported: usage tracking, budget integration (including hard spend caps), and permission checks on the model service.”
↩︎ Checkpoint - 4.
“The definer must have CAN QUERY permission on the endpoint.”
↩︎ Which models ai_query can call, and what each needs“ai_query is a general purpose AI Function that enables you to query existing endpoints for real-time inference or batch inference workloads.”
↩︎ Key concept