What you will be able to do
- Explain what the lakehouse architecture unifies, and match the platform's current and former names
- Describe what the Data Intelligence Engine does for users of the platform
- Describe Delta Lake as the default storage layer and name the guarantees it adds to Parquet files
- Describe Unity Catalog's role as the governance layer, its three-level namespace, and the difference between managed and external tables
Key concept
Lakehouse architecture — A single platform that combines the open, low-cost storage of a data lake with the reliability and governance of a data warehouse. Every Databricks component (storage, governance, SQL, pipelines, orchestration, AI) is built on this one shared copy of data.
1.One platform on the lakehouse
Before looking at the individual components, it helps to see how the platform describes itself. The Databricks documentation now calls it the Databricks Data + AI Platform. The glossary says this name was "formerly called the Databricks Lakehouse Platform and the Data Intelligence Platform". Your exam guide uses the older name, "Data Intelligence Platform", and all three names refer to the same product.
The platform-scope page describes it in a single sentence: it is "built on the lakehouse architecture and powered by a data intelligence engine that understands the unique qualities of your data." The same page calls it an open and unified foundation for ETL, ML/AI, and DWH/BI workloads, with Unity Catalog as the central governance solution.
The word to focus on is *unified*. A data lakehouse combines enterprise data warehouses and data lakes. Data engineers, data scientists, analysts and production systems can all treat it as their single source of truth. That means less work building, maintaining and syncing several separate data systems.
| Layer | What the scope page says |
|---|---|
| Architecture | Built on the lakehouse architecture |
| Intelligence | Powered by a data intelligence engine that understands the unique qualities of your data |
| Workloads | An open and unified foundation for ETL, ML/AI, and DWH/BI workloads |
| Governance | Unity Catalog as the central data and AI governance solution |
Checkpoint 1 of 6· Check yourself
A colleague says the lakehouse is "basically a data lake with a SQL tool bolted on". Which description matches the Databricks documentation?
The lakehouse combines warehouses and lakes. Engineers, scientists, analysts and production systems share it as one source of truth instead of syncing separate systems.
“The data lakehouse combines enterprise data warehouses and data lakes to accelerate, simplify, and unify enterprise data solutions.”Source: docs.databricks.com
2.The Data Intelligence Engine
The glossary describes the engine under its "Databricks AI features" entry. That entry is a glossary heading, not a separate product: it defines "The data intelligence engine powering the Databricks Platform." The engine is not a storage format or a single query optimizer. The glossary calls it "a compound AI system that combines the use of AI models, retrieval, ranking, and personalization systems", which work "to understand the semantics of your organization's data and usage patterns".
The introduction to Databricks explains what this means in practice. Databricks "uses AI with the data lakehouse to understand the unique semantics of your data". It then automatically optimizes performance and manages infrastructure to match your business needs. As an analyst, you notice the engine when you search for data: natural language processing learns your business's language, so you can search and discover data by asking a question in your own words. Natural language assistance also helps you write code, troubleshoot errors and find answers in the documentation.
The introduction says Databricks "uses AI with the data lakehouse to understand the unique semantics of your data", and the glossary says the engine learns the semantics of your data and usage patterns. The documentation states the link but does not give a further reason for it.
Checkpoint 2 of 6· Check yourself
Which statement best describes the Data Intelligence Engine?
The glossary defines the data intelligence engine as a compound AI system. The other options describe Delta Lake, SQL warehouses and Unity Catalog.
“a compound AI system that combines the use of AI models, retrieval, ranking, and personalization systems”Source: docs.databricks.com
3.Delta Lake: the storage layer under every table
The lakehouse needs reliable tables on top of cheap object storage, and Delta Lake provides them. Delta Lake is "the optimized storage layer that provides the foundation for tables in a lakehouse on Databricks." It is open-source software that "extends Parquet data files with a file-based transaction log for ACID transactions and scalable metadata handling."
The Parquet guess in the prediction is understandable. A Delta table does keep its data in Parquet files, and Databricks stores all data and metadata for Delta Lake tables in cloud object storage. The difference is the transaction log. It turns a folder of files into a table with warehouse-style guarantees. Because you get Delta automatically by saving data with default settings, these guarantees apply whether you use SQL or DataFrames.
The transaction log also gives you table history. Each write creates a new table version, and you can use the log to review modifications to your table and query previous table versions. The log has a well-defined open protocol that any system can use to read it. To inspect a table's configuration and metadata, the documentation points to the DESCRIBE DETAIL command.
Delta Lake was also built for tight integration with Structured Streaming, so one copy of data can serve both batch and streaming work.
| Capability | What it means for you |
|---|---|
| ACID transactions | Provided by the file-based transaction log that Delta Lake adds to Parquet files |
| Table versions | Each write creates a new table version; the log lets you review modifications and query previous versions |
| Schema enforcement | Delta Lake validates schema on write against the requirements you set |
| Batch + streaming | A single copy of data serves both batch and streaming operations |
| Data skipping | Column statistics and data layout reduce the number of files scanned per query |
Checkpoint 3 of 6· Check yourself
What does Delta Lake add to ordinary Parquet data files?
Delta Lake keeps Parquet and adds a transaction log. The log's protocol is open, so any system can read it, which rules out the proprietary option.
“Delta Lake is open source software that extends Parquet data files with a file-based transaction log for ACID transactions and scalable metadata handling.”Source: docs.databricks.com
Sources4
4.Unity Catalog: governance for data and AI
Delta Lake makes the data reliable, and Unity Catalog controls who can reach it. Unity Catalog is "the unified governance layer for data and AI built into Databricks." When it is enabled, it works underneath every interaction automatically: "enforcing access control when you query a table or call a model, tracking lineage as data and AI assets are used", and logging activity for auditing. It is automatically enabled for all Databricks workspaces created after November 8, 2023.
Every governed asset is a securable object that you can grant permissions on. Data and AI assets (tables, views, volumes, functions, models and services) use a three-level namespace, catalog.schema.object. Catalogs are the highest-level container for organizing and isolating data. Schemas, also called databases, sit inside catalogs and provide a finer level of organization. A view is a read-only object derived from one or more tables and views. Volumes represent a logical volume of storage in a cloud object storage location and organize and govern access to non-tabular data.
Tables come in two kinds. A managed table is one where Unity Catalog handles both governance and the underlying file storage lifecycle. An external table is one where Unity Catalog handles governance only. In both cases, access goes through Unity Catalog.
You browse these objects in Catalog Explorer, where you can find data objects and owners, understand data relationships across tables, and manage permissions and sharing. Unity Catalog's other capabilities include lineage that follows assets from source data through to models and dashboards, an audit log system table, data classification, data quality monitoring, and secure sharing across organizations through the open OpenSharing protocol.
Checkpoint 4 of 6· Match them up
Match each Unity Catalog object to its description
Tap a term, then the definition that fits it.
Catalogs contain schemas, and schemas contain tables, views, volumes, functions and models, which together form the catalog.schema.object namespace.
“Catalogs are the highest level container for organizing and isolating data on Databricks.”Source: docs.databricks.com
Checkpoint 5 of 6· Check yourself
A team registers an external table in Unity Catalog over files in their own storage location. What does Unity Catalog manage for that table?
For an external table, Unity Catalog still governs access but does not manage the files. For a managed table, it does both.
“or external, where Unity Catalog handles governance only”Source: docs.databricks.com
Checkpoint 6 of 6· Exam question
A data engineering team wants a single platform that combines the low-cost flexible storage of a data lake with the transaction support and BI performance of a data warehouse, while also using AI that understands the structure and semantics of their own data to power search and governance automatically. Which underlying architecture and capability of the Databricks platform provides this combination?
Correct answer: A — The lakehouse architecture combined with the Data Intelligence Engine, which unifies warehouse transactions with lake storage while using AI to understand the organization's data semantics.
- A. This correctly describes the lakehouse architecture as the foundation that unifies warehouse and lake workloads, with the Data Intelligence Engine layering AI-driven understanding of the organization's specific data on top of that unified storage and compute layer.
- B. Bolting a warehouse onto a separate lake with manual synchronization is exactly the two-system problem the lakehouse architecture was designed to eliminate, so this describes the older pattern rather than the platform's actual approach.
- C. Model serving and a plain object store address hosting predictions, not the combined transactional-and-analytical storage layer or the semantic understanding the scenario asks about, so this option targets the wrong layer of the platform.
- D. Scheduled batch orchestration between two separate systems still leaves data duplicated across a lake and a warehouse and adds no AI-driven semantic layer, so it does not match what the scenario is asking for.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A table on Databricks is a Delta table only if you explicitly ask for Delta format.Why is that wrong?
Delta Lake is the default for all operations on Databricks. A table you create without specifying a format is a Delta Lake table.
2.External tables sit outside Unity Catalog, so it doesn't govern them.Why is that wrong?
Unity Catalog governs both kinds of table. The only difference is that it also manages the file storage lifecycle for managed tables.
Covered in Unity Catalog: governance for data and AI
3."Data Intelligence Platform" and "Data + AI Platform" are two different Databricks products.Why is that wrong?
They are the same platform. The glossary lists Data Intelligence Platform, along with Lakehouse Platform, as former names of the Data + AI Platform.
Covered in One platform on the lakehouse
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“It is built on the lakehouse architecture and powered by a data intelligence engine that understands the unique qualities of your data.”
↩︎ One platform on the lakehouse“It is an open and unified foundation for ETL, ML/AI, and DWH/BI workloads, and has Unity Catalog as the central data and AI governance solution.”
↩︎ One platform on the lakehouse - 2.
“Formerly called the Databricks Lakehouse Platform and the Data Intelligence Platform.”
↩︎ One platform on the lakehouse“The data intelligence engine powering the Databricks Platform.”
↩︎ The Data Intelligence Engine“to understand the semantics of your organization's data and usage patterns”
↩︎ The Data Intelligence Engine“Formerly called the Databricks Lakehouse Platform and the Data Intelligence Platform.”
↩︎ Exam trap 3“a compound AI system that combines the use of AI models, retrieval, ranking, and personalization systems”
↩︎ Checkpoint - 3.https://docs.databricks.com/aws/en/introductionOfficial docs
“Databricks uses AI with the data lakehouse to understand the unique semantics of your data.”
↩︎ The Data Intelligence Engine“Natural language processing learns your business's language, so you can search and discover data by asking a question in your own words.”
↩︎ The Data Intelligence Engine“The data lakehouse combines enterprise data warehouses and data lakes to accelerate, simplify, and unify enterprise data solutions.”
↩︎ Key concept“The data lakehouse combines enterprise data warehouses and data lakes to accelerate, simplify, and unify enterprise data solutions.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/deltaOfficial docs
“Delta Lake is the optimized storage layer that provides the foundation for tables in a lakehouse on Databricks.”
↩︎ Delta Lake: the storage layer under every table“Each write to a Delta Lake table creates a new table version.”
↩︎ Delta Lake: the storage layer under every table“You can use the transaction log to review modifications to your table and query previous table versions.”
↩︎ Delta Lake: the storage layer under every table“The Delta Lake transaction log has a well-defined open protocol that can be used by any system to read the log.”
↩︎ Delta Lake: the storage layer under every table“View table configurations and metadata using the DESCRIBE DETAIL command.”
↩︎ Delta Lake: the storage layer under every table“Databricks stores all data and metadata for Delta Lake tables in cloud object storage.”
↩︎ Delta Lake: the storage layer under every table“Unless otherwise specified, all tables on Databricks are Delta Lake tables.”
↩︎ Exam trap 1“Unless otherwise specified, all tables on Databricks are Delta Lake tables.”
↩︎ Prediction“Delta Lake is open source software that extends Parquet data files with a file-based transaction log for ACID transactions and scalable metadata handling.”
↩︎ Checkpoint - 5.
“Unity Catalog is the unified governance layer for data and AI built into Databricks.”
↩︎ Unity Catalog: governance for data and AI“Data and AI assets such as tables, views, volumes, functions, models, and services (model services and MCP services) follow a three-level namespace (catalog.schema.object).”
↩︎ Unity Catalog: governance for data and AI“managed, where Unity Catalog handles both governance and the underlying file storage lifecycle”
↩︎ Unity Catalog: governance for data and AI“Securely share live data and AI assets across organizations and clouds using the open OpenSharing protocol.”
↩︎ Unity Catalog: governance for data and AI“or external, where Unity Catalog handles governance only”
↩︎ Exam trap 2“or external, where Unity Catalog handles governance only”
↩︎ Checkpoint - 6.
“Volumes represent a logical volume of storage in a cloud object storage location and organize and govern access to non-tabular data.”
↩︎ Unity Catalog: governance for data and AI“You can use it to find data objects and owners, understand data relationships across tables, and manage permissions and sharing.”
↩︎ Unity Catalog: governance for data and AI“Catalogs are the highest level container for organizing and isolating data on Databricks.”
↩︎ Checkpoint