What you will be able to do
- Explain what an external volume is and when to use one external volume or several
- Verify an external volume and set a default EXTERNAL_VOLUME at the account, database or schema level
- Distinguish internal from external catalogs in Snowflake Open Catalog, the managed service for Apache Polaris
- Describe how service principals, principal roles and catalog roles give query engines access to Open Catalog
1.External volumes: the storage connection for Iceberg tables
Apache Iceberg tables keep their data and metadata in your own cloud storage, so Snowflake needs a way to reach that storage. For stages, that is a storage integration. For Iceberg tables, it is an external volume: a named, account-level object that stores an IAM entity for a storage location. Snowflake uses that entity to read and write table data, Iceberg metadata and manifest files. An external volume must exist before you can create an Iceberg table, and one volume can serve many tables.
How many volumes you need depends on how you want to organize and secure the data. A single volume can hold the data and metadata of all your tables in subdirectories of one location. You might instead separate them by access: a read-only volume for externally managed Iceberg tables, and a read-write volume for Snowflake-managed tables. Each storage location has these settings when you add it in Snowsight:
| Field | Meaning |
|---|---|
| Region type | Standard (public AWS outside China) or Government (GovCloud) |
| S3 role ARN | Case-sensitive ARN of the IAM role with privileges on the bucket |
| Encryption | None (default), SSE-S3, or SSE-KMS with a key |
| Connectivity | Public (default) or Private (AWS PrivateLink) outbound connectivity |
| Storage base URL | Base URL of the storage location |
To check that Snowflake can authenticate to the storage, verify the volume. In Snowsight, go to Catalog » External data » External volumes and choose … » Verify connection. In SQL, call the system function:
SELECT SYSTEM$VERIFY_EXTERNAL_VOLUME('my_s3_external_volume');Checkpoint 1 of 6· Check yourself
A team wants Spark-written Iceberg tables that Snowflake only reads, plus Snowflake-managed Iceberg tables that Snowflake writes. They want different security on the two. What does Snowflake's guidance suggest?
Multiple external volumes let you secure storage locations differently. The documentation gives exactly this read-only versus read-write split as its example.
“A read-only external volume for externally managed Iceberg tables.”Source: docs.snowflake.com
Sources1
2.Default volumes, USAGE and adding storage locations
Instead of naming a volume in every CREATE ICEBERG TABLE, set the EXTERNAL_VOLUME parameter. An account administrator can set it with ALTER ACCOUNT, which makes it the default for every Iceberg table in the account. Users can override it on a database or schema, and the narrowest setting wins: schema, then database, then account. Setting it on a database or schema requires the usual ALTER privileges plus USAGE on the external volume. USAGE also lets a role reference the volume and view its details. The owner can grant it in Snowsight with + Privilege.
Checkpoint 2 of 6· Fill the gap
Which parameter makes my_s3_vol the default for new Iceberg tables in this database?
ALTER DATABASE my_database_1
SET ? = 'my_s3_vol';EXTERNAL_VOLUME is the parameter that sets the default volume. BASE_LOCATION is a per-table path inside that volume.
Source: docs.snowflake.comWith a database default in place, a table that uses Snowflake as its catalog needs only a BASE_LOCATION:
CREATE ICEBERG TABLE iceberg_reviews_table (
id STRING,
product_name STRING,
product_id STRING,
reviewer_name STRING,
review_date DATE,
review STRING
)
CATALOG = 'SNOWFLAKE'
BASE_LOCATION = 'my/product_reviews/';To add another storage location to an existing volume, use the ADD STORAGE_LOCATION parameter of ALTER EXTERNAL VOLUME. In Snowsight, a role with OWNERSHIP on the volume can choose … » Add storage location.
Checkpoint 3 of 6· Check yourself
The account default is vol_acct, database DW sets vol_db, and schema DW.STAGE sets vol_sch. Which volume does a new Iceberg table in DW.STAGE use if CREATE ICEBERG TABLE names none?
The narrowest declaration wins. The schema setting overrides the database and account defaults.
“The lowest-scoped declaration is used: schema > database > account.”Source: docs.snowflake.com
Sources1
3.Snowflake Open Catalog: catalogs, namespaces and storage
The exam guide calls it the Polaris Catalog, and Snowflake's product name is Snowflake Open Catalog: a managed service for Apache Polaris, built on the open Iceberg REST protocol. It gives REST-compatible engines centralized, secure read and write access to Iceberg tables. Note the status: customers without an existing Open Catalog account cannot sign up, and new customers are directed to Snowflake Horizon Catalog. Existing customers can keep using it and create more Open Catalog accounts.
An Open Catalog account holds one or more catalogs. Each catalog stores the current metadata pointer for its tables and updates it atomically. A catalog is one of two types. An internal catalog is managed by Open Catalog, and both third-party engines and Snowflake can read and write its tables. An external catalog is managed by another provider; currently only Snowflake is offered. Its tables are synced into Open Catalog and are read-only there. Namespaces, which can be nested, group the tables inside a catalog.
When you create a catalog you also create its storage configuration. For S3, that means a default base location, the allowed locations, an S3 role ARN and an optional external ID. Azure needs a tenant ID. Open Catalog generates an IAM entity that the cloud provider is told to trust, which is the same pattern as a storage integration. Access privileges are only enforced correctly if each directory holds one table's files and the directory hierarchy matches the namespace hierarchy. Snowflake-managed tables in an external catalog can overlap directories, so give each one a unique BASE_LOCATION. Dropping a table without purging it leaves its data in cloud storage. For that reason, never create a new table with the same name and location as the dropped one.
Checkpoint 4 of 6· Match them up
Match each Open Catalog term to its meaning
Tap a term, then the definition that fits it.
An internal catalog is managed by Open Catalog and is writable by engines. An external catalog is a read-only mirror of another provider's catalog. Namespaces organize tables, and the storage configuration provides cloud access.
“External: The catalog is externally managed by another Iceberg catalog provider (for example, Snowflake, Glue, Dremio Arctic).”Source: docs.snowflake.com
Sources2
4.Connecting engines: service principals and roles
Engines connect to Open Catalog through service principals. Each one holds credentials, a Client ID and Client Secret pair generated by Open Catalog. A service connection represents an engine such as Spark, Flink or Trino. When an administrator creates one, they give its service principal a new or existing principal role. Privileges are granted to catalog roles, catalog roles are granted to principal roles, and through them the service principal gets its access. A new principal role starts with no privileges until catalog roles are granted to it. Snowflake itself reads Open Catalog tables through a catalog integration, by creating an externally managed Iceberg table.
| Object | Role in access control |
|---|---|
| Service principal | Holds the Client ID and Client Secret an engine uses to connect |
| Principal role | Groups service principals; receives catalog roles |
| Catalog role | Receives privileges on securable objects in a catalog |
Checkpoint 5 of 6· Check yourself
Snowflake-managed Iceberg tables are synced to Open Catalog as an external catalog. A Spark job connects through a service principal that has every privilege on that catalog. What can Spark do with these tables?
Tables in an external catalog belong to their source catalog. Other engines can read them, but only Snowflake writes them.
“The Snowflake query engine can read from or write to these tables. However, the other query engines can only read from these tables.”Source: docs.snowflake.com
Checkpoint 6 of 6· Exam question
A storage integration permits `s3://corp-data/` but the compliance team now requires that the `s3://corp-data/hr/` prefix be unreachable from any stage built on this integration, while all other prefixes keep working. What is the most appropriate change?
Correct answer: B — Run `ALTER STORAGE INTEGRATION` to set `STORAGE_BLOCKED_LOCATIONS = ('s3://corp-data/hr/')` while leaving the allowed location list as it is.
- A. Incorrect. Revoking usage removes access to every prefix, not just the HR one, and it pushes teams back toward embedded credentials, which weakens the control.
- B. Correct. Blocked locations are evaluated alongside allowed ones, so the HR prefix is excluded while the rest of the bucket stays reachable through the integration.
- C. Incorrect. That would permit only the HR prefix, which is the opposite of the requirement, and it would break every other prefix currently in use.
- D. Incorrect. A `PATTERN` is chosen by the caller of each statement and does not enforce anything at integration level, so users could simply omit it.
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Changing a database's EXTERNAL_VOLUME moves that database's existing Iceberg tables to the new volume.Why is that wrong?
The change only applies to tables created afterwards. Existing tables keep the volume they were created with.
Covered in Default volumes, USAGE and adding storage locations
2.Any new Snowflake customer can sign up for Snowflake Open Catalog to get a managed Polaris catalog.Why is that wrong?
First-time sign-ups are closed. New customers are directed to Snowflake Horizon Catalog, and only existing Open Catalog customers can create more Open Catalog accounts.
Covered in Snowflake Open Catalog: catalogs, namespaces and storage
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“An external volume is a named, account-level Snowflake object that you use to connect Snowflake to your external cloud storage for Iceberg tables.”
↩︎ External volumes: the storage connection for Iceberg tables“You must create an external volume before you can create an Apache Iceberg™ table in Snowflake.”
↩︎ External volumes: the storage connection for Iceberg tables“The USAGE privilege grants the ability to reference the external volume and view details for the external volume.”
↩︎ Default volumes, USAGE and adding storage locations“To add a storage location to an external volume by using SQL, use the ADD STORAGE_LOCATION parameter of the ALTER EXTERNAL VOLUME command.”
↩︎ Default volumes, USAGE and adding storage locations“Existing tables continue to use the external volume specified when they were created.”
↩︎ Exam trap 1“A read-only external volume for externally managed Iceberg tables.”
↩︎ Checkpoint“Existing tables continue to use the external volume specified when they were created.”
↩︎ Prediction“The lowest-scoped declaration is used: schema > database > account.”
↩︎ Checkpoint - 2.
“Snowflake Open Catalog is a managed service for Apache Polaris™.”
↩︎ Snowflake Open Catalog: catalogs, namespaces and storage“These tables are read-only in Open Catalog.”
↩︎ Snowflake Open Catalog: catalogs, namespaces and storage“A directory hierarchy matches the namespace hierarchy for the catalog.”
↩︎ Snowflake Open Catalog: catalogs, namespaces and storage“an IAM entity is generated and used to create a trust relationship between the cloud storage provider and Open Catalog.”
↩︎ Snowflake Open Catalog: catalogs, namespaces and storage“When you drop a table without purging it, its data is retained in the external cloud storage.”
↩︎ Snowflake Open Catalog: catalogs, namespaces and storage“Open Catalog generates a Client ID and Client Secret pair for each service principal.”
↩︎ Connecting engines: service principals and roles“the Open Catalog administrator grants privileges to catalog roles and then grants these catalog roles to the new principal role.”
↩︎ Connecting engines: service principals and roles“New customers should use Snowflake Horizon Catalog for Apache Iceberg™ tables and multi-engine interoperability with Iceberg.”
↩︎ Exam trap 2“External: The catalog is externally managed by another Iceberg catalog provider (for example, Snowflake, Glue, Dremio Arctic).”
↩︎ Checkpoint“The Snowflake query engine can read from or write to these tables. However, the other query engines can only read from these tables.”
↩︎ Checkpoint