What you will be able to do
- Explain why a storage integration replaces explicit cloud credentials and what the AWS administrator has to configure for it
- Write CREATE STORAGE INTEGRATION with allowed and blocked locations, and predict what ENABLED = FALSE does
- Create a named external stage that references a storage integration
- Choose the account parameters that block ad hoc unloads and stages created with credentials
Key concept
Storage integration — A storage integration is an account-level Snowflake object that holds a Snowflake-generated cloud identity and, optionally, a list of allowed and blocked storage locations. Stages use it to authenticate, so nobody types secret keys into SQL. The cloud administrator grants permissions to that identity, not to individual users.
1.Why stages authenticate through an integration
Before Snowflake can load files from a private S3 bucket, or unload files to one, it has to authenticate to that bucket. One way is to paste an access key and secret into a stage definition. The other is a storage integration. Integrations are named, first-class objects that remove the need to pass explicit cloud credentials. Each integration stores an AWS IAM user ID, and an administrator in your organization grants that IAM user permissions in the AWS account. The relationship runs in one direction: an external stage references the integration, and the integration carries the identity. Many stages can point at different buckets and paths while sharing a single integration for authentication.
On the AWS side, Snowflake needs s3:GetBucketLocation, s3:GetObject, s3:GetObjectVersion and s3:ListBucket on the bucket and folder in order to read files. Other SQL actions need extra permissions, shown in the table below. Snowflake recommends putting these permissions in an IAM policy and attaching that policy to an IAM role. When you create the role, choose AWS account as the trusted entity and enter your own account ID for now. Select the Require external ID option and enter a placeholder such as 0000. In a later step you edit the role's trust relationship so that it trusts the Snowflake integration's IAM user and external ID. Record the role ARN, because the integration references it.
| S3 permission | Needed for |
|---|---|
| s3:GetBucketLocation, s3:GetObject, s3:GetObjectVersion, s3:ListBucket | Any access to files in the folder (loading) |
| s3:PutObject | Unloading files to the bucket |
| s3:DeleteObject | Purging files after a successful load (PURGE) or running REMOVE |
Checkpoint 1 of 7· Put it in order
Put the AWS-side and Snowflake-side setup steps for an S3 storage integration in order
- 1.Run CREATE STORAGE INTEGRATION with STORAGE_AWS_ROLE_ARN set to that ARN
- 2.Create an IAM role with Require external ID set to a placeholder, and attach the policy
- 3.Record the role ARN from the role summary page
- 4.Create an IAM policy that grants the required S3 permissions on the bucket and prefix
First the policy is created, then a role that uses it. The role's ARN is what the integration's STORAGE_AWS_ROLE_ARN parameter points at.
“You have now created an IAM policy for a bucket, created an IAM role, and attached the policy to the role.”Source: docs.snowflake.com
Revoking access does not take effect immediately. Snowflake caches the temporary credentials it obtains for up to 60 minutes. If you revoke Snowflake's access, users may still be able to list files and read data until the cache expires.
Sources1
2.Creating and managing the integration object
Only users with the ACCOUNTADMIN role, or a role that holds the global CREATE INTEGRATION privilege, can run CREATE STORAGE INTEGRATION. For S3, the statement names the IAM role ARN you recorded and lists the locations stages are allowed to use. It can also list locations they are blocked from using.
CREATE STORAGE INTEGRATION <integration_name>
TYPE = EXTERNAL_STAGE
STORAGE_PROVIDER = 'S3'
ENABLED = TRUE
STORAGE_AWS_ROLE_ARN = '<iam_role>'
STORAGE_ALLOWED_LOCATIONS = ('<protocol>://<bucket>/<path>/', '<protocol>://<bucket>/<path>/')
[ STORAGE_BLOCKED_LOCATIONS = ('<protocol>://<bucket>/<path>/', '<protocol>://<bucket>/<path>/') ]STORAGE_ALLOWED_LOCATIONS limits the stages that use the integration to the buckets and optional paths you list. Paths are case-sensitive. The parameter also accepts the * wildcard, which allows every bucket and path. For S3, the protocol is s3 for public AWS regions outside China, s3china for public regions in China, and s3gov for government regions. GCS locations use gcs:// and Azure locations use azure://account.blob.core.windows.net/container/path/. A storage integration only works for a government-region or China-region bucket from a Snowflake account hosted in that same region. In other cases, use the CREDENTIALS parameter of CREATE STAGE instead.
| ENABLED value | New stages | Existing stages |
|---|---|---|
| TRUE (default) | Can be created referencing the integration | Function normally |
| FALSE | Cannot be created referencing the integration | Cannot access the storage location in their definition |
Checkpoint 2 of 7· Check yourself
A security engineer runs ALTER STORAGE INTEGRATION s3_int SET ENABLED = FALSE. What happens to the five stages that already reference s3_int?
Setting ENABLED to FALSE blocks new stages from referencing the integration. It also stops existing stages from reaching their storage location.
“Existing stages that reference this integration cannot access the storage location in the stage definition.”Source: docs.snowflake.com
Checkpoint 3 of 7· Exam question
An administrator creates a storage integration for an S3 bucket and builds an IAM role whose trust policy names the customer's own AWS account. `CREATE STAGE` using the integration then fails with an access denied error. What should the administrator do to fix the trust relationship?
Correct answer: C — Run `DESC STORAGE INTEGRATION`, then trust `STORAGE_AWS_IAM_USER_ARN` as principal and require `STORAGE_AWS_EXTERNAL_ID` as the `sts:ExternalId` in the trust policy.
- A. Incorrect. Storage integrations are designed to avoid stored credentials, and they have no access-key properties; the delegation works through role assumption only.
- B. Incorrect. Snowflake roles are not AWS principals, and the account locator is not the external ID; the generated IAM user ARN and external ID come from the integration description.
- C. Correct. Snowflake generates an IAM user and external ID for each integration, and the role must trust exactly that principal with that external ID before Snowflake can assume it.
- D. Incorrect. Storage integrations authenticate by Snowflake assuming the IAM role, so the role trust policy has to name Snowflake's IAM user; a bucket policy on a VPC endpoint does not replace that.
3.Named external stages that use the integration
You can load directly from files in an S3 bucket. If the folder path ends with /, every object in that folder is loaded. A named external stage packages everything needed to reach the files: the bucket, the storage integration (or S3 credentials) and, if the files are encrypted, an encryption key. Named stages are optional, but they are recommended when you load from the same location regularly. If the stage also names a file format, later COPY commands do not have to repeat the format options.
Checkpoint 4 of 7· Fill the gap
Which parameter connects this stage to the s3_int integration?
Copy codeCREATE STAGE my_s3_stage ? = s3_int URL = 's3://mybucket/encrypted_files/' FILE_FORMAT = my_csv_format;STORAGE_INTEGRATION names the integration whose IAM identity the stage uses. CREDENTIALS would put explicit keys into the stage definition instead.
Source: docs.snowflake.comYou can also create the stage in Snowsight, under Create » Stage » External Stage. Directory tables are enabled by default in that dialog and can be deselected. A directory table lists the files on the stage, but it needs a warehouse and therefore costs money. Snowflake uses multipart uploads to S3 and GCS, which can leave incomplete uploads in the stage location. A lifecycle rule on the bucket stops these from accumulating.
Checkpoint 5 of 7· Check yourself
Integration s3_int has STORAGE_ALLOWED_LOCATIONS = ('s3://mybucket/loads/'). A developer creates a stage on s3_int with URL = 's3://otherbucket/loads/'. What is the outcome?
The allowed-locations list limits which URLs a stage built on the integration can reference. Granting privileges does not change that list.
“The URL in the stage definition must align with the S3 buckets (and optional paths) specified for the STORAGE_ALLOWED_LOCATIONS parameter.”Source: docs.snowflake.com
Sources3
4.Preventing data exfiltration through stages and unloads
Allowed and blocked locations on an integration set the boundary for stages that use it. Without further settings, though, a user could create a stage with their own access keys, or unload straight to a URL in a COPY statement. Snowflake provides account parameters that close these gaps. All of them default to FALSE, so they do nothing until an administrator sets them.
| Parameter | Effect when TRUE |
|---|---|
| REQUIRE_STORAGE_INTEGRATION_FOR_STAGE_CREATION | CREATE STAGE for a private cloud location must reference a storage integration rather than secret keys or tokens |
| REQUIRE_STORAGE_INTEGRATION_FOR_STAGE_OPERATION | Loads from and unloads to private cloud storage must go through a named external stage that references a storage integration |
| PREVENT_UNLOAD_TO_INLINE_URL | COPY INTO <location> must target a named stage or an internal user or table stage, not an inline URL with access settings |
| PREVENT_UNLOAD_TO_INTERNAL_STAGES | Unloading table data to any internal stage, including user and table stages, is blocked |
Used together, these parameters send every load and unload through a stage whose location is fixed in its definition and whose identity comes from a storage integration. That integration's allowed locations are under administrator control. When PREVENT_UNLOAD_TO_INLINE_URL is TRUE, a named external stage must store the cloud storage URL and access settings in its definition.
Checkpoint 6 of 7· Check yourself
Auditors find that analysts have been running COPY INTO 's3://personal-bucket/...' with credentials typed into the statement. Which single parameter blocks exactly this?
No stage is created, so stage-creation rules and an integration's blocked list never apply. PREVENT_UNLOAD_TO_INLINE_URL blocks ad hoc unloads to external URLs.
“Specifies whether to prevent ad hoc data unload operations to external cloud storage locations”Source: docs.snowflake.com
Checkpoint 7 of 7· Exam question
A security policy states that nobody may create a stage with an inline credentials string and that unloads must never target a raw cloud storage URL typed into a `COPY INTO` statement. Select TWO account parameters that enforce this.(Select 2)
Correct answers: A, E — Set `PREVENT_UNLOAD_TO_INLINE_URL = TRUE` so that `COPY INTO <location>` statements cannot write to a cloud storage URL given directly in the command.; Set `REQUIRE_STORAGE_INTEGRATION_FOR_STAGE_CREATION = TRUE` so that new external stages must reference a storage integration rather than embed credentials.
- A. Correct. This parameter blocks unloads that pass an inline URL, forcing unloads to go through a named stage, which can be governed by an integration.
- B. Incorrect. This parameter only influences the physical column types in unloaded Parquet files; it has no effect on destination or authentication rules.
- C. Incorrect. Rekeying rotates encryption keys for data at rest in Snowflake storage; it does not constrain how stages are created or where unloads can write.
- D. Incorrect. This parameter changes what the login page shows for SSO; it does not restrict stage creation or unload destinations.
- E. Correct. With this enabled, `CREATE STAGE` on an external location fails unless a storage integration is specified, so credentials can no longer be embedded in stage definitions.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Revoking the IAM role's access in AWS cuts Snowflake off immediately.Why is that wrong?
Snowflake caches temporary credentials for up to 60 minutes. Users may keep listing and reading files until the cache expires.
2.A storage integration can reach S3 buckets in a government region from any Snowflake account.Why is that wrong?
Only accounts hosted in the same government (or China) region can do this. Other accounts must use the CREDENTIALS parameter of CREATE STAGE.
Covered in Creating and managing the integration object
3.Requiring storage integrations for stage creation also stops users from unloading data to arbitrary S3 URLs.Why is that wrong?
Ad hoc COPY INTO <location> statements put the URL and credentials inline and never create a stage. Blocking them needs PREVENT_UNLOAD_TO_INLINE_URL.
Covered in Preventing data exfiltration through stages and unloads
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Integrations are named, first-class Snowflake objects that avoid the need for passing explicit cloud provider credentials such as secret keys or access tokens.”
↩︎ Why stages authenticate through an integration“many external stage objects can reference different buckets and paths and use the same storage integration for authentication.”
↩︎ Why stages authenticate through an integration“Select the Require external ID option.”
↩︎ Why stages authenticate through an integration“Only account administrators (users with the ACCOUNTADMIN role) or a role with the global CREATE INTEGRATION privilege can execute this SQL command.”
↩︎ Creating and managing the integration object“Snowflake creates a single IAM user that is referenced by all S3 storage integrations in your Snowflake account.”
↩︎ Key concept“Snowflake caches the temporary credentials for a period that cannot exceed the 60-minute expiration time.”
↩︎ Exam trap 1“You have now created an IAM policy for a bucket, created an IAM role, and attached the policy to the role.”
↩︎ Checkpoint“The URL in the stage definition must align with the S3 buckets (and optional paths) specified for the STORAGE_ALLOWED_LOCATIONS parameter.”
↩︎ Checkpoint - 2.
“FALSE prevents users from creating new stages that reference this integration.”
↩︎ Creating and managing the integration object“Alternatively supports the * wildcard”
↩︎ Creating and managing the integration object“use the CREDENTIALS parameter in the CREATE STAGE command (rather than using a storage integration) to provide the credentials for authentication.”
↩︎ Exam trap 2“Existing stages that reference this integration cannot access the storage location in the stage definition.”
↩︎ Checkpoint - 3.
“Named external stages are optional, but recommended when you plan to load data regularly from the same location.”
↩︎ Named external stages that use the integration“Directory tables let you see files on the stage, but require a warehouse and thus incur a cost.”
↩︎ Named external stages that use the integration - 4.
“Specifies whether to require a storage integration object as cloud credentials when creating a named external stage”
↩︎ Preventing data exfiltration through stages and unloads“Specifies whether to require using a named external stage that references a storage integration object as cloud credentials when loading data”
↩︎ Preventing data exfiltration through stages and unloads“Specifies whether to prevent data unload operations to internal (Snowflake) stages using COPY INTO <location> statements.”
↩︎ Preventing data exfiltration through stages and unloads“TRUE: COPY INTO <location> statements must reference either a named internal (Snowflake) or external stage or an internal user or table stage.”
↩︎ Exam trap 3“COPY INTO <location> statements that specify the cloud storage URL and access settings directly in the statement”
↩︎ Prediction“Specifies whether to prevent ad hoc data unload operations to external cloud storage locations”
↩︎ Checkpoint - 5.
“Snowflake provides a set of parameters to further restrict data unloading operations”
↩︎ Preventing data exfiltration through stages and unloads