CertSafari
    Databricks Certified Data Analyst Associate· Lessons

    Domain 1 · Lesson 1/39

    Databricks Platform Foundations: Lakehouse, Data Intelligence Engine, Delta Lake and Unity Catalog

    Describe the core components of the Databricks Intelligence Platform, including Mosaic AI, DeltaLive tables, Lakeflow Jobs, Data Intelligence Engine, Delta Lake, Unity Catalog, and Databricks SQL.

    11 min read
    2.56% of exam
    6 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain what the lakehouse architecture unifies, and match the platform's current and former names
    • Describe what the Data Intelligence Engine does for users of the platform
    • Describe Delta Lake as the default storage layer and name the guarantees it adds to Parquet files
    • Describe Unity Catalog's role as the governance layer, its three-level namespace, and the difference between managed and external tables

    Key concept

    Lakehouse architecture — A single platform that combines the open, low-cost storage of a data lake with the reliability and governance of a data warehouse. Every Databricks component (storage, governance, SQL, pipelines, orchestration, AI) is built on this one shared copy of data.

    1.One platform on the lakehouse

    Before looking at the individual components, it helps to see how the platform describes itself. The Databricks documentation now calls it the Databricks Data + AI Platform. The glossary says this name was "formerly called the Databricks Lakehouse Platform and the Data Intelligence Platform". Your exam guide uses the older name, "Data Intelligence Platform", and all three names refer to the same product.

    The platform-scope page describes it in a single sentence: it is "built on the lakehouse architecture and powered by a data intelligence engine that understands the unique qualities of your data." The same page calls it an open and unified foundation for ETL, ML/AI, and DWH/BI workloads, with Unity Catalog as the central governance solution.

    The word to focus on is *unified*. A data lakehouse combines enterprise data warehouses and data lakes. Data engineers, data scientists, analysts and production systems can all treat it as their single source of truth. That means less work building, maintaining and syncing several separate data systems.

    How the platform's pieces map to its workloads, according to the platform-scope page
    LayerWhat the scope page says
    ArchitectureBuilt on the lakehouse architecture
    IntelligencePowered by a data intelligence engine that understands the unique qualities of your data
    WorkloadsAn open and unified foundation for ETL, ML/AI, and DWH/BI workloads
    GovernanceUnity Catalog as the central data and AI governance solution

    Checkpoint 1 of 6· Check yourself

    A colleague says the lakehouse is "basically a data lake with a SQL tool bolted on". Which description matches the Databricks documentation?

    Sources12

    2.The Data Intelligence Engine

    The glossary describes the engine under its "Databricks AI features" entry. That entry is a glossary heading, not a separate product: it defines "The data intelligence engine powering the Databricks Platform." The engine is not a storage format or a single query optimizer. The glossary calls it "a compound AI system that combines the use of AI models, retrieval, ranking, and personalization systems", which work "to understand the semantics of your organization's data and usage patterns".

    The introduction to Databricks explains what this means in practice. Databricks "uses AI with the data lakehouse to understand the unique semantics of your data". It then automatically optimizes performance and manages infrastructure to match your business needs. As an analyst, you notice the engine when you search for data: natural language processing learns your business's language, so you can search and discover data by asking a question in your own words. Natural language assistance also helps you write code, troubleshoot errors and find answers in the documentation.

    Checkpoint 2 of 6· Check yourself

    Which statement best describes the Data Intelligence Engine?

    Sources23

    3.Delta Lake: the storage layer under every table

    The lakehouse needs reliable tables on top of cheap object storage, and Delta Lake provides them. Delta Lake is "the optimized storage layer that provides the foundation for tables in a lakehouse on Databricks." It is open-source software that "extends Parquet data files with a file-based transaction log for ACID transactions and scalable metadata handling."

    The Parquet guess in the prediction is understandable. A Delta table does keep its data in Parquet files, and Databricks stores all data and metadata for Delta Lake tables in cloud object storage. The difference is the transaction log. It turns a folder of files into a table with warehouse-style guarantees. Because you get Delta automatically by saving data with default settings, these guarantees apply whether you use SQL or DataFrames.

    The transaction log also gives you table history. Each write creates a new table version, and you can use the log to review modifications to your table and query previous table versions. The log has a well-defined open protocol that any system can use to read it. To inspect a table's configuration and metadata, the documentation points to the DESCRIBE DETAIL command.

    Delta Lake was also built for tight integration with Structured Streaming, so one copy of data can serve both batch and streaming work.

    What the transaction log gives a Delta Lake table
    CapabilityWhat it means for you
    ACID transactionsProvided by the file-based transaction log that Delta Lake adds to Parquet files
    Table versionsEach write creates a new table version; the log lets you review modifications and query previous versions
    Schema enforcementDelta Lake validates schema on write against the requirements you set
    Batch + streamingA single copy of data serves both batch and streaming operations
    Data skippingColumn statistics and data layout reduce the number of files scanned per query

    Checkpoint 3 of 6· Check yourself

    What does Delta Lake add to ordinary Parquet data files?

    Sources4

    4.Unity Catalog: governance for data and AI

    Delta Lake makes the data reliable, and Unity Catalog controls who can reach it. Unity Catalog is "the unified governance layer for data and AI built into Databricks." When it is enabled, it works underneath every interaction automatically: "enforcing access control when you query a table or call a model, tracking lineage as data and AI assets are used", and logging activity for auditing. It is automatically enabled for all Databricks workspaces created after November 8, 2023.

    Every governed asset is a securable object that you can grant permissions on. Data and AI assets (tables, views, volumes, functions, models and services) use a three-level namespace, catalog.schema.object. Catalogs are the highest-level container for organizing and isolating data. Schemas, also called databases, sit inside catalogs and provide a finer level of organization. A view is a read-only object derived from one or more tables and views. Volumes represent a logical volume of storage in a cloud object storage location and organize and govern access to non-tabular data.

    Tables come in two kinds. A managed table is one where Unity Catalog handles both governance and the underlying file storage lifecycle. An external table is one where Unity Catalog handles governance only. In both cases, access goes through Unity Catalog.

    You browse these objects in Catalog Explorer, where you can find data objects and owners, understand data relationships across tables, and manage permissions and sharing. Unity Catalog's other capabilities include lineage that follows assets from source data through to models and dashboards, an audit log system table, data classification, data quality monitoring, and secure sharing across organizations through the open OpenSharing protocol.

    Checkpoint 4 of 6· Match them up

    Match each Unity Catalog object to its description

    Tap a term, then the definition that fits it.

    Checkpoint 5 of 6· Check yourself

    A team registers an external table in Unity Catalog over files in their own storage location. What does Unity Catalog manage for that table?

    Checkpoint 6 of 6· Exam question

    A data engineering team wants a single platform that combines the low-cost flexible storage of a data lake with the transaction support and BI performance of a data warehouse, while also using AI that understands the structure and semantics of their own data to power search and governance automatically. Which underlying architecture and capability of the Databricks platform provides this combination?

    Sources56

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A table on Databricks is a Delta table only if you explicitly ask for Delta format.Why is that wrong?

      Delta Lake is the default for all operations on Databricks. A table you create without specifying a format is a Delta Lake table.

      Covered in Delta Lake: the storage layer under every table

    2. 2.External tables sit outside Unity Catalog, so it doesn't govern them.Why is that wrong?

      Unity Catalog governs both kinds of table. The only difference is that it also manages the file storage lifecycle for managed tables.

      Covered in Unity Catalog: governance for data and AI

    3. 3."Data Intelligence Platform" and "Data + AI Platform" are two different Databricks products.Why is that wrong?

      They are the same platform. The glossary lists Data Intelligence Platform, along with Lakehouse Platform, as former names of the Data + AI Platform.

      Covered in One platform on the lakehouse

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “It is built on the lakehouse architecture and powered by a data intelligence engine that understands the unique qualities of your data.”
      ↩︎ One platform on the lakehouse
      “It is an open and unified foundation for ETL, ML/AI, and DWH/BI workloads, and has Unity Catalog as the central data and AI governance solution.”
      ↩︎ One platform on the lakehouse
    2. 2.
      “Formerly called the Databricks Lakehouse Platform and the Data Intelligence Platform.”
      ↩︎ One platform on the lakehouse
      “The data intelligence engine powering the Databricks Platform.”
      ↩︎ The Data Intelligence Engine
      “to understand the semantics of your organization's data and usage patterns”
      ↩︎ The Data Intelligence Engine
      “Formerly called the Databricks Lakehouse Platform and the Data Intelligence Platform.”
      ↩︎ Exam trap 3
      “a compound AI system that combines the use of AI models, retrieval, ranking, and personalization systems”
      ↩︎ Checkpoint
    3. 3.
      “Databricks uses AI with the data lakehouse to understand the unique semantics of your data.”
      ↩︎ The Data Intelligence Engine
      “Natural language processing learns your business's language, so you can search and discover data by asking a question in your own words.”
      ↩︎ The Data Intelligence Engine
      “The data lakehouse combines enterprise data warehouses and data lakes to accelerate, simplify, and unify enterprise data solutions.”
      ↩︎ Key concept
      “The data lakehouse combines enterprise data warehouses and data lakes to accelerate, simplify, and unify enterprise data solutions.”
      ↩︎ Checkpoint
    4. 4.
      “Delta Lake is the optimized storage layer that provides the foundation for tables in a lakehouse on Databricks.”
      ↩︎ Delta Lake: the storage layer under every table
      “Each write to a Delta Lake table creates a new table version.”
      ↩︎ Delta Lake: the storage layer under every table
      “You can use the transaction log to review modifications to your table and query previous table versions.”
      ↩︎ Delta Lake: the storage layer under every table
      “The Delta Lake transaction log has a well-defined open protocol that can be used by any system to read the log.”
      ↩︎ Delta Lake: the storage layer under every table
      “View table configurations and metadata using the DESCRIBE DETAIL command.”
      ↩︎ Delta Lake: the storage layer under every table
      “Databricks stores all data and metadata for Delta Lake tables in cloud object storage.”
      ↩︎ Delta Lake: the storage layer under every table
      “Unless otherwise specified, all tables on Databricks are Delta Lake tables.”
      ↩︎ Exam trap 1
      “Unless otherwise specified, all tables on Databricks are Delta Lake tables.”
      ↩︎ Prediction
      “Delta Lake is open source software that extends Parquet data files with a file-based transaction log for ACID transactions and scalable metadata handling.”
      ↩︎ Checkpoint
    5. 5.
      “Unity Catalog is the unified governance layer for data and AI built into Databricks.”
      ↩︎ Unity Catalog: governance for data and AI
      “Data and AI assets such as tables, views, volumes, functions, models, and services (model services and MCP services) follow a three-level namespace (catalog.schema.object).”
      ↩︎ Unity Catalog: governance for data and AI
      “managed, where Unity Catalog handles both governance and the underlying file storage lifecycle”
      ↩︎ Unity Catalog: governance for data and AI
      “Securely share live data and AI assets across organizations and clouds using the open OpenSharing protocol.”
      ↩︎ Unity Catalog: governance for data and AI
      “or external, where Unity Catalog handles governance only”
      ↩︎ Exam trap 2
      “or external, where Unity Catalog handles governance only”
      ↩︎ Checkpoint
    6. 6.
      “Volumes represent a logical volume of storage in a cloud object storage location and organize and govern access to non-tabular data.”
      ↩︎ Unity Catalog: governance for data and AI
      “You can use it to find data objects and owners, understand data relationships across tables, and manage permissions and sharing.”
      ↩︎ Unity Catalog: governance for data and AI
      “Catalogs are the highest level container for organizing and isolating data on Databricks.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Databricks Workloads: Databricks SQL, Delta Live Tables (Lakeflow Pipelines), Lakeflow Jobs and Mosaic AI

    Spotted a mistake, or was something unclear? Tell us.