What you will be able to do
- Find any table or view in Catalog Explorer using Unity Catalog's three-level namespace
- Recognise the catalogs Databricks creates for you and explain what the default catalog does
- Explain what a schema is, including why it is also called a database
- Tell managed tables from external tables by storage location, supported formats and what happens on DROP
- Describe what a view stores and why views have no managed or external variant
Key concept
Three-level namespace (catalog.schema.object) — Every data object in Unity Catalog has a three-part address: a catalog, then a schema inside it, then the table, view, volume, model or function inside that. Catalog Explorer is organised the same way, and permissions set at a higher level pass down to the levels below.
1.Catalogs: the top of the hierarchy
When you click Catalog in a Databricks workspace, you open Catalog Explorer. Its left-hand tree follows Unity Catalog's hierarchy, and the top level of that tree is the catalog. A catalog is the main way Unity Catalog organises data. It holds schemas, and those schemas hold the actual data objects. Catalogs are registered in a Unity Catalog metastore. Your account has one metastore per region, so catalogs are separated by region automatically.
Because a catalog sits at the top, it should mark a logical boundary for data isolation and access. Organisations often create catalogs to match business units or development stages, such as one catalog for production data and another for development data. Each catalog usually has its own managed storage location, which keeps its managed data physically separate from other catalogs.
The tree will not be empty, because Databricks creates several catalogs for you:
- Workspace catalog: created by default in new workspaces and usually named after the workspace. Every user in that workspace, and only that workspace, can access it by default, so it's a good place to experiment. - hive_metastore: holds objects from the legacy Hive metastore. The Hive metastore is deprecated, and Databricks says all workspaces should migrate to Unity Catalog. - __databricks_internal: stores internal state for features such as Lakeflow pipelines and AI/BI dashboards. Seeing it in Catalog Explorer is expected, but you should not query, modify or delete anything in it.
Each Unity Catalog-enabled workspace also has a default catalog. If you leave the catalog name out of a query, Databricks uses the default catalog. A workspace admin can change which catalog is the default.
When you create a catalog yourself, you choose between two types. A standard catalog is the normal kind. A foreign catalog is used only with Lakehouse Federation: it mirrors a database in an external system and supports read-only queries against it.
Checkpoint 1 of 5· Check yourself
An analyst runs SELECT * FROM sales.orders and leaves out the catalog name. What happens?
Every Unity Catalog-enabled workspace has a default catalog, and Databricks uses it whenever the catalog name is left out. A workspace admin can change which catalog that is.
“If you omit the top-level catalog name when you perform data operations, the default catalog is assumed.”Source: docs.databricks.com
Sources1
2.Schemas: the second level
Expand a catalog in Catalog Explorer and you see its schemas. A schema belongs to a catalog and can hold tables, views, volumes, models and functions. It groups data more finely than a catalog does. A schema typically represents one use case, project or team sandbox.
In Databricks, a schema is sometimes called a database, and CREATE DATABASE is an alias for CREATE SCHEMA. Some relational systems use these words differently: there, a database contains several schemas. Don't carry that meaning over to Databricks.
Unity Catalog adds it automatically to every catalog. It is a reserved schema of read-only views that describe the objects in that catalog, and it is separate from any schemas users create.
You can optionally give a schema its own managed storage location. This keeps the files for all managed tables and volumes in that schema separate from other schemas in the same catalog. If the schema has no managed location, its data goes to the catalog's location. If the catalog has none either, the data goes to the metastore's location. This only affects managed data. How external tables are isolated depends on how you manage your own cloud storage.
Checkpoint 2 of 5· Check yourself
A schema has no managed storage location of its own, but its parent catalog does. Where are files for a new managed table in that schema stored?
Unity Catalog uses the nearest managed location it finds, checking the schema first, then the catalog, then the metastore. Here the catalog has one, so the files go there.
“If you don't specify a managed storage location for the schema, data resides in the catalog's managed storage location”Source: docs.databricks.com
Sources2
3.Managed versus external tables
Unity Catalog governs every table registered in it, managed or external. It controls who can access the table and records auditing and lineage for it. Managed and external tables differ in one thing: who controls the underlying data files.
- Managed table: Unity Catalog decides where the files are stored, using the schema, catalog or metastore managed location. It also manages the files' lifecycle, including optimisation, organisation and deletion. Managed tables use the Delta or Apache Iceberg format. - External table: you choose the storage location, and you or another system manage the files. External tables support Delta, CSV, JSON, Avro, Parquet and ORC.
In both cases the data stays in your own cloud account. Databricks does not take ownership of it. Both kinds can also be read, written and created by external engines through open APIs, so using managed tables does not lock you in.
| Property | Managed table | External table |
|---|---|---|
| Storage location | Set by Unity Catalog (in your cloud account) | Set by you |
| File lifecycle management | Managed by Unity Catalog (optimization, organization, deletion) | Managed by you |
| Drop behavior | Data files are permanently deleted after an 8-day retention period | Data files remain in place |
| Formats | Delta or Apache Iceberg | Delta, CSV, JSON, Avro, Parquet, and ORC |
| Governed by Unity Catalog | Yes | Yes |
The word *manage* has several meanings in Unity Catalog, and exam questions can take advantage of that. An object managed by Unity Catalog is simply one whose access Unity Catalog governs, and that includes external tables. A managed table is a table whose storage location and file lifecycle Unity Catalog also controls. MANAGE is a privilege: it lets someone grant or revoke privileges on an object, transfer its ownership and delete it, without being its owner.
Checkpoint 3 of 5· Match them up
Match each use of manage to its meaning
Tap a term, then the definition that fits it.
Only a managed table (or volume) means Unity Catalog controls the files. Being managed by Unity Catalog means it governs access, and MANAGE is a privilege that can be granted on any securable object.
“MANAGE allows a user to assign or revoke privileges on, transfer ownership of, and delete an object without being the owner.”Source: docs.databricks.com
Checkpoint 4 of 5· Exam question
A data analyst runs `DROP TABLE retail.sales.orders_managed` against a managed table that Unity Catalog created with its own storage location. What happens to the Parquet/Delta data files after the statement finishes?
Correct answer: A — The data files are deleted along with the catalog metadata, because Unity Catalog owns the full storage lifecycle of a managed table by default.
- A. Correct. Unity Catalog manages both the governance metadata and the underlying storage lifecycle of a managed table, so dropping it removes the data files together with the catalog entry.
- B. Incorrect. This describes external table behavior, where Unity Catalog only tracks metadata; a managed table's files are removed together with the drop, not left behind.
- C. Incorrect. Dropping a managed table does not move its files into any quarantine volume; there is no such automatic relocation step in Unity Catalog.
- D. Incorrect. A drop statement does not re-register the files as a new external table; the storage and metadata for a managed table are removed together.
Sources3
4.Views: saved queries, not stored data
Views appear in Catalog Explorer next to tables, at the third level of the namespace (catalog.schema.view). A view is a read-only object defined by a query over one or more tables or other views, and those can come from different schemas and catalogs. Creating a view does not process or write any data. Only the text of the query is saved, in the schema where the view lives. After that, anyone with permission can query the view from anywhere in Databricks.
A view doesn't store data files, so the managed-versus-external distinction doesn't apply to it. That distinction covers tables and volumes only. Databricks also recommends defining views by referring to tables or views by name, not by file path or URI, because path-based definitions can make governance requirements confusing.
Checkpoint 5 of 5· Check yourself
A colleague says the team should make its new view an external view, so the data survives if the view is dropped. What's wrong with that suggestion?
A view saves only its query text, so there is no data to keep. The managed-versus-external distinction applies only to tables and volumes.
“The distinction between managed and external applies to tables and volumes only.”Source: docs.databricks.com
Sources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Dropping any Unity Catalog table deletes its data files.Why is that wrong?
Only dropping a managed table deletes the data files, and that happens after an 8-day retention period. Dropping an external table removes only its metadata; the files stay where they are.
Covered in Managed versus external tables
2.External tables are not governed by Unity Catalog because they are not managed.Why is that wrong?
Unity Catalog governs external tables too. The only thing it doesn't control for them is the storage and lifecycle of the files.
Covered in Managed versus external tables
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/catalogsOfficial docs
“They contain schemas, which in turn can contain tables, views, volumes, models, and functions.”
↩︎ Catalogs: the top of the hierarchy“The Hive metastore is deprecated, and all Databricks workspaces should migrate to Unity Catalog.”
↩︎ Catalogs: the top of the hierarchy“Do not query, modify, or delete objects in the __databricks_internal catalog.”
↩︎ Catalogs: the top of the hierarchy“Catalogs are the first layer in Unity Catalog's three-level namespace (catalog.schema.table-etc).”
↩︎ Key concept“If you omit the top-level catalog name when you perform data operations, the default catalog is assumed.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/schemasOfficial docs
“In Databricks, schemas are sometimes called databases. For example, CREATE DATABASE is an alias for CREATE SCHEMA.”
↩︎ Schemas: the second level“If you don't specify a managed storage location for the schema, data resides in the catalog's managed storage location”
↩︎ Checkpoint - 3.https://docs.databricks.com/aws/en/data-governance/unity-catalog/managed-versus-externalOfficial docs
“Managed tables use the Delta or Apache Iceberg format.”
↩︎ Managed versus external tables“When you drop an external table, Unity Catalog removes the table metadata from the metastore, but the underlying data files remain in place.”
↩︎ Exam trap 1“External assets: Unity Catalog controls governance only.”
↩︎ Exam trap 2“When you drop an external table, Unity Catalog removes the table metadata from the metastore, but the underlying data files remain in place.”
↩︎ Prediction“MANAGE allows a user to assign or revoke privileges on, transfer ownership of, and delete an object without being the owner.”
↩︎ Checkpoint“The distinction between managed and external applies to tables and volumes only.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/viewsOfficial docs
“Creating a view does not process or write any data.”
↩︎ Views: saved queries, not stored data“Databricks recommends that you always define views by referencing data sources using a table or view name.”
↩︎ Views: saved queries, not stored data