What you will be able to do
- Describe Databricks SQL as a lakehouse data warehouse that runs on SQL warehouses
- Recognize Delta Live Tables as the former name of Lakeflow pipelines and distinguish streaming tables, materialized views and views
- Describe Lakeflow Jobs in terms of jobs, tasks and triggers, and how it differs from a pipeline
- Describe the AI capabilities reached through the Mosaic AI area of the workspace, including AI functions for SQL analysts
1.Databricks SQL: the warehouse on the lakehouse
Databricks runs several kinds of workload on the same governed Delta tables. For a data analyst, the most important one is Databricks SQL, which the documentation describes as "a cloud data warehouse built on lakehouse architecture." It "runs directly on your data lake, supports ANSI SQL with Delta Lake extensions", so you build a warehouse without moving data into a separate system.
Databricks SQL runs on SQL warehouses. Administrators configure scalable compute as SQL warehouses so that end users can run queries without managing the complexities of the cloud. The reference architecture also names serverless SQL warehouses as part of how the platform serves data warehousing and BI. You then reach the same warehouse from many interfaces: the SQL editor, notebooks attached to a SQL warehouse, scheduled jobs, AI/BI dashboards, metric views, alerts and a REST API. Databricks SQL also covers ETL, because you can define and refresh streaming tables and materialized views directly in it. Query history, query profile and query performance insights help you work out why a query is slow.
Don't confuse SQL warehouses with Lakebase. Lakebase is a different kind of system: an online transactional processing (OLTP) database, fully managed and Postgres-based, that is integrated with the platform. It is not what Databricks SQL runs on.
| Interface | Purpose |
|---|---|
| SQL editor | Write and run SQL queries with integrated AI assistance, code comments, and version history |
| Dashboards | Create interactive AI/BI dashboards to share insights |
| Metric views | Define business metrics with consistent calculations using a semantic layer |
| Alerts | Monitor query results, evaluate conditions, and deliver notifications |
| Query profile | Inspect the execution plan for a query to identify bottlenecks |
Checkpoint 1 of 7· Check yourself
Which compute resource does Databricks SQL run on?
Databricks SQL runs on SQL warehouses, which administrators configure so analysts can query without managing cloud infrastructure. Lakebase is an OLTP Postgres database, not the compute for Databricks SQL.
“Databricks SQL runs on SQL warehouses and is accessible from multiple interfaces for querying, visualization, pipeline management, and automation.”Source: docs.databricks.com
2.Delta Live Tables, now Lakeflow pipelines
The product formerly known as Delta Live Tables (DLT) has been updated to Lakeflow pipelines, and the glossary now lists DLT as "the deprecated name for Lakeflow pipelines." Don't confuse it with Delta Lake. Delta Lake is the storage layer, while Lakeflow pipelines are a way of building ETL that writes Delta tables.
Lakeflow pipelines "provide a declarative framework for building batch and streaming data pipelines in SQL and Python." With hand-written Spark code you would order and retry each step yourself. A pipeline handles this automatically: it runs its processing steps (flows) in the correct order with maximum parallelism and retries transient failures. Its incremental engine reprocesses only new or changed source data when possible.
You can add expectations, which are SQL boolean constraints that validate records as they flow through. For each expectation you choose what happens on failure: warn, drop the record, or fail the update. All tables created and managed by pipelines are Delta tables, so they keep ACID transactions, time travel and schema enforcement.
| Dataset type | How records are processed | Typical use |
|---|---|---|
| Streaming table | Each record is processed exactly one time, assuming an append-only source | Ingestion and incremental processing of continuously growing data |
| Materialized view | Results are recomputed as needed to reflect the current state of the data | Transformations, aggregations, pre-computed results for many consumers |
| View | Evaluated on demand, not persisted | Intermediate transformations and checks inside the pipeline |
Checkpoint 2 of 7· Match them up
Match each pipeline dataset type to how it processes records
Tap a term, then the definition that fits it.
Streaming tables suit ingestion, materialized views suit aggregations that many queries read, and views suit intermediate steps that don't need publishing to a catalog.
“Each record is processed exactly one time, assuming an append-only source.”Source: docs.databricks.com
3.Lakeflow Jobs: orchestrating the work
A pipeline handles the ordering inside a single ETL flow. To coordinate several kinds of work, such as a pipeline, a notebook and a SQL query, you need Lakeflow Jobs. It is "workflow automation for Databricks, providing orchestration for data processing workloads", and it lets you schedule repeatable tasks and manage complex workflows.
The docs name three main concepts for Lakeflow Jobs: jobs, tasks and triggers. A job is the main resource for coordinating, scheduling and running your operations. It can be anything from one notebook to hundreds of tasks with conditional logic, and its tasks are visually represented by a Directed Acyclic Graph (DAG). A task is one unit of work. A notebook task runs a notebook, a Python script task runs a file, and a pipeline task runs a Lakeflow pipeline. Tasks can depend on one another and can branch with if/else or loop with for each. A trigger decides when the job runs. Triggers can be time-based, such as every day at 2 AM, or event-based, such as when new data arrives in cloud storage.
Analysts use this too, because Databricks SQL lets you schedule SQL queries as jobs for automated data processing and reporting. In the workspace, the Jobs & Pipelines UI is the entry point to Jobs, Lakeflow pipelines and Lakeflow Connect.
Checkpoint 3 of 7· Match them up
Match each Lakeflow Jobs concept to what it is
Tap a term, then the definition that fits it.
A job is the container, tasks are the units of work inside it, and a trigger decides when the job starts, either on a time schedule or when an event occurs.
“A trigger is a mechanism that initiates running a job based on specific conditions or events.”Source: docs.databricks.com
Checkpoint 4 of 7· Check yourself
A team wants one nightly workflow that first refreshes a Lakeflow pipeline and then runs a notebook. What coordinates the two steps?
A pipeline only orders the flows inside itself. A job coordinates different kinds of tasks, and running a pipeline is just one task type, so a pipeline task and a notebook task can sit in the same job.
“A pipeline task runs a pipeline.”Source: docs.databricks.com
Model answer: the materialized view is a pipeline dataset, because a pipeline is the declarative ETL framework that produces materialized views, streaming tables and views. The nightly schedule and the ordering between the refresh and the dashboard query are Lakeflow Jobs concerns. You would build a job with a pipeline task that refreshes the view and a later task that runs the query, with a time-based trigger such as every day at 2 AM.
4.Mosaic AI: the AI capabilities on the same data
The last component is the one the documentation names least directly. In the current docs, "Mosaic AI" appears as a section of the workspace homepage, and the docs say: "This section provides quick access to AI-related assets and features." The sources do not give a formal product definition. They describe the features behind that entry point as Databricks AI capabilities, and that is what this section covers.
Databricks "provides a platform for building, evaluating, deploying, and monitoring AI applications (AI apps)." It hosts AI models from top providers as Foundation Models, and you can query these models, external models and your own models through UI, API/SDK and SQL interfaces. Model Serving is the serving capability: the reference architecture describes it as scalable, real-time and enterprise-grade, and the glossary says Databricks supports real-time and batch inference through it. For machine learning work, the Feature Store is a central repository for storing, managing and serving features, and managed MLflow handles tracing, evaluation and monitoring. The tools also include AI Playground for prototyping, Knowledge Assistant for guided agent building, custom agents in Python and AI Search for indexing unstructured data.
This connects back to the lakehouse in two ways. First, AI uses the same governed data: all data is managed under Unity Catalog, and a Unity Catalog table can be served for AI through a vector index or a feature table. Second, analysts don't need to leave SQL, because Databricks provides AI functions that SQL data analysts can use to call LLMs directly within their pipelines and workflows.
Checkpoint 5 of 7· Check yourself
A SQL analyst wants to classify the text of customer reviews with an LLM, without writing Python. Which capability fits?
AI functions let SQL analysts call LLMs from inside their queries and pipelines. The other options orchestrate, store or govern data but don't call models.
“Databricks provides AI functions that SQL data analysts can use to access LLMs, including from OpenAI, directly within their data pipelines and workflows.”Source: docs.databricks.com
5.Unity Catalog: names and managed tables in the warehouse
Every table you query from a SQL warehouse or reach from a pipeline or job is governed by Unity Catalog. Data and AI assets such as tables, views, volumes, functions and models "follow a three-level namespace (catalog.schema.object)", so a table is addressed as catalog.schema.table.
Tables and volumes can be managed or external. With a managed table, Unity Catalog handles both governance and the underlying file storage lifecycle. With an external table, Unity Catalog handles governance only, and the glossary adds that the data lifecycle is managed outside of Databricks. Keep that difference in mind when you answer the questions below.
Checkpoint 6 of 7· Exam question
An analyst is asked to write a query against a table named `transactions` that lives inside the `sales` schema of the `east_region` catalog, in a workspace where Unity Catalog is enabled. Which fully qualified reference correctly identifies this table?
Correct answer: A — east_region.sales.transactions
- A. Unity Catalog's three-level namespace orders identifiers as catalog, then schema, then table, so naming the catalog first, the schema second, and the table last is the correct fully qualified reference.
- B. This swaps the catalog and schema positions, which points the query at a schema named `east_region` inside a catalog named `sales` rather than the intended location.
- C. This puts the table name first, which Unity Catalog does not accept as the leading segment of the three-level namespace, so the reference does not resolve to the intended table.
- D. This puts the schema and table segments out of order relative to the catalog, so it does not match the catalog.schema.table structure Unity Catalog expects.
Checkpoint 7 of 7· Exam question
A data analyst creates a managed table in Unity Catalog by running `CREATE TABLE` without a `LOCATION` clause, then later drops the table with `DROP TABLE`. What happens to the underlying data files as a direct result of that drop?
Correct answer: A — Unity Catalog deletes the underlying data files from the managed location it provisioned, because a managed table's physical files are fully controlled by the catalog itself.
- A. For a managed table, Unity Catalog owns both the metadata and the storage location it provisioned, so dropping the table removes the catalog entry and deletes the underlying files together.
- B. Leaving files untouched describes the behavior of dropping an external table, where Unity Catalog only governs metadata; a managed table's files are deleted along with its metadata, so this does not match the scenario.
- C. There is no built-in ninety-day archive volume that Unity Catalog moves managed table files into on drop, so this describes a recovery mechanism that does not exist for this operation.
- D. A manual cleanup step is unnecessary here because dropping a managed table already deletes its files automatically; a separate admin job is only relevant for external tables whose storage the catalog does not own.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Delta Live Tables and Delta Lake are the same thing, or Delta Live Tables was discontinued and replaced with something incompatible.Why is that wrong?
Delta Lake is the storage layer. Delta Live Tables is the former name of Lakeflow pipelines, a declarative ETL framework that writes Delta tables, and existing DLT code still runs.
Covered in Delta Live Tables, now Lakeflow pipelines
2.Lakeflow Jobs and Lakeflow pipelines are interchangeable names for the same orchestration tool.Why is that wrong?
A pipeline is declarative ETL that orders its own flows. Lakeflow Jobs orchestrates tasks of many types, and running a pipeline is just one kind of task inside a job.
Covered in Lakeflow Jobs: orchestrating the work
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/sqlOfficial docs
“Databricks SQL is a cloud data warehouse built on lakehouse architecture.”
↩︎ Databricks SQL: the warehouse on the lakehouse“It runs directly on your data lake, supports ANSI SQL with Delta Lake extensions”
↩︎ Databricks SQL: the warehouse on the lakehouse“Define and refresh streaming tables and materialized views directly in Databricks SQL for incremental ETL pipelines.”
↩︎ Databricks SQL: the warehouse on the lakehouse“Schedule SQL queries as jobs for automated data processing and reporting workflows.”
↩︎ Lakeflow Jobs: orchestrating the work“Databricks SQL runs on SQL warehouses and is accessible from multiple interfaces for querying, visualization, pipeline management, and automation.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/introductionOfficial docs
“Administrators configure scalable compute clusters as SQL warehouses, so end users can run queries without managing the complexities of working in the cloud.”
↩︎ Databricks SQL: the warehouse on the lakehouse“Lakebase is an online transactional processing (OLTP) database that is fully integrated with the Databricks Data + AI Platform.”
↩︎ Databricks SQL: the warehouse on the lakehouse“Databricks provides AI functions that SQL data analysts can use to access LLMs, including from OpenAI, directly within their data pipelines and workflows.”
↩︎ Checkpoint - 3.
“the data warehouse powered by SQL warehouses, and serverless SQL warehouses”
↩︎ Databricks SQL: the warehouse on the lakehouse“Model Serving is a scalable, real-time, enterprise-grade model serving capability hosted in the Databricks control plane.”
↩︎ Mosaic AI: the AI capabilities on the same data - 4.
“The deprecated name for Lakeflow pipelines.”
↩︎ Delta Live Tables, now Lakeflow pipelines“A central repository for storing, managing, and serving features for machine learning models.”
↩︎ Mosaic AI: the AI capabilities on the same data“Databricks supports real-time and batch inference through Model Serving.”
↩︎ Mosaic AI: the AI capabilities on the same data“the data lifecycle is managed outside of Databricks”
↩︎ Unity Catalog: names and managed tables in the warehouse - 5.https://docs.databricks.com/aws/en/ldp/conceptsOfficial docs
“Lakeflow pipelines provide a declarative framework for building batch and streaming data pipelines in SQL and Python.”
↩︎ Delta Live Tables, now Lakeflow pipelines“specify what happens when a record fails: warn, drop the record, or fail the update.”
↩︎ Delta Live Tables, now Lakeflow pipelines“All tables created and managed by pipelines are Delta tables.”
↩︎ Delta Live Tables, now Lakeflow pipelines“Each record is processed exactly one time, assuming an append-only source.”
↩︎ Checkpoint - 6.https://docs.databricks.com/aws/en/jobsOfficial docs
“Lakeflow Jobs is workflow automation for Databricks, providing orchestration for data processing workloads”
↩︎ Lakeflow Jobs: orchestrating the work“There are three main concepts when using Lakeflow Jobs for orchestration in Databricks: jobs, tasks, and triggers.”
↩︎ Lakeflow Jobs: orchestrating the work“The tasks in a job are visually represented by a Directed Acyclic Graph (DAG).”
↩︎ Lakeflow Jobs: orchestrating the work“A pipeline task runs a pipeline.”
↩︎ Exam trap 2“A trigger is a mechanism that initiates running a job based on specific conditions or events.”
↩︎ Checkpoint“A pipeline task runs a pipeline.”
↩︎ Checkpoint - 7.
“The Jobs & Pipelines workspace UI provides entry to the Jobs, Lakeflow pipelines, and Lakeflow Connect UIs”
↩︎ Lakeflow Jobs: orchestrating the work - 8.
“This section provides quick access to AI-related assets and features.”
↩︎ Mosaic AI: the AI capabilities on the same data - 9.
“Databricks provides a platform for building, evaluating, deploying, and monitoring AI applications (AI apps).”
↩︎ Mosaic AI: the AI capabilities on the same data“All of these models can be queried through UI, API/SDK, and SQL interfaces.”
↩︎ Mosaic AI: the AI capabilities on the same data“A table in Unity Catalog can be served for AI, using a vector index for unstructured data or a feature table for structured data.”
↩︎ Mosaic AI: the AI capabilities on the same data - 10.
“follow a three-level namespace (catalog.schema.object)”
↩︎ Unity Catalog: names and managed tables in the warehouse“Unity Catalog handles both governance and the underlying file storage lifecycle”
↩︎ Unity Catalog: names and managed tables in the warehouse“or external, where Unity Catalog handles governance only”
↩︎ Unity Catalog: names and managed tables in the warehouse
Also cited
“The product formerly known as Delta Live Tables (DLT) has been updated to Lakeflow pipelines.”
↩︎ Exam trap 1“If you have previously used DLT, there is no migration required to use Lakeflow pipelines: your code will still work.”
↩︎ Prediction