What you will be able to do
- Explain how a storage integration lets external stages reach cloud storage without stored credentials
- Configure the scope of an API integration with API_PROVIDER and API_ALLOWED_PREFIXES
- Set up a Git repository clone from an API integration and a secret, then run SQL from it with EXECUTE IMMEDIATE FROM
- Troubleshoot common Git and storage integration failures from their documented rules
1.Storage integration: delegated access to cloud storage
When Snowflake has to read from or write to your own cloud storage, it needs an identity that storage provider accepts. Without a storage integration, someone would paste secret keys into each stage definition. With one, Snowflake keeps a generated identity and access management (IAM) entity for your cloud storage. A cloud administrator grants that entity permissions on the buckets or containers, and users no longer supply credentials when they create stages or load and unload data.
The integration also controls *where* stages can point. STORAGE_ALLOWED_LOCATIONS lists the buckets and optional paths an external stage may reference, and STORAGE_BLOCKED_LOCATIONS can carve out exceptions. One integration can serve many external stages, but each stage URL must fall inside the allowed locations.
CREATE [ OR REPLACE ] STORAGE INTEGRATION [IF NOT EXISTS]
<name>
TYPE = { EXTERNAL_STAGE | POSTGRES_EXTERNAL_STORAGE | POSTGRES_INTERNAL_STORAGE }
cloudProviderParams
ENABLED = { TRUE | FALSE }
STORAGE_ALLOWED_LOCATIONS = ('<cloud>://<bucket>/<path>/' [ , '<cloud>://<bucket>/<path>/' ... ] )
[ STORAGE_BLOCKED_LOCATIONS = ('<cloud>://<bucket>/<path>/' [ , '<cloud>://<bucket>/<path>/' ... ] ) ]
[ COMMENT = '<string_literal>' ]The provider-specific parameters show how much of the work Snowflake handles. Amazon S3 needs STORAGE_AWS_ROLE_ARN, the IAM role your AWS administrator sets up. Microsoft Azure needs AZURE_TENANT_ID. Google Cloud Storage needs only STORAGE_PROVIDER = 'GCS': no key file is exchanged or stored, because the identity is generated on the Snowflake side.
Here is how it works for S3. The external stage references the storage integration. Snowflake associates the integration with the IAM user it created for your account, and your AWS administrator grants that user access to the buckets. Each time someone loads or unloads through the stage, Snowflake checks the IAM user's permissions on the bucket before it allows or denies access.
Two edge cases come up in exam questions. First, setting ENABLED = FALSE does more than stop new stages from using the integration. Existing stages that reference it lose access to their storage location too. Second, an integration cannot always be used. For cloud storage in a government region or a China region, it only works from a Snowflake account hosted in that same region. In other cases you supply credentials with the CREDENTIALS parameter of CREATE STAGE.
Checkpoint 1 of 8· Check yourself
An administrator runs ALTER on a storage integration to set ENABLED = FALSE. What happens to the external stages that already reference it?
With ENABLED = FALSE, nobody can create new stages that reference the integration, and existing stages that reference it lose access to their storage location.
“Existing stages that reference this integration cannot access the storage location in the stage definition.”Source: docs.snowflake.com
Checkpoint 2 of 8· Exam question
A platform team is building a Java-based ETL scheduler that must connect to Snowflake through a broad set of existing client tools already standardized on a common database connectivity API. Which driver best fits this requirement?
Correct answer: A — The JDBC driver, because it implements the widely adopted Java Database Connectivity standard that most Java tools already expect.
- A. The JDBC driver implements the standard Java Database Connectivity API, which is exactly the interface Java-based schedulers and most client tools already expect, giving the broadest compatibility described.
- B. The Node.js driver targets JavaScript runtimes and does not provide a JDBC interface, so it would not satisfy a requirement built around Java-based tooling and JDBC compatibility.
- C. Go applications use the Go Snowflake driver's own native interface; Go does not natively expose a JDBC-compatible surface, so this does not meet the stated requirement.
- D. PDO is PHP's own database abstraction layer and is unrelated to JDBC, so Java-based tooling cannot consume it without an entirely different integration layer.
2.API integration: reaching services over HTTPS
A storage integration is for object storage. An API integration is for any service Snowflake calls over HTTPS. The object records which kind of service it is, the identifier and credentials Snowflake uses to call it, and which endpoints are allowed. The service might be a cloud provider's HTTPS proxy, such as an Amazon API Gateway behind an external function, a Git repository API, or an external Model Context Protocol (MCP) server. The API_PROVIDER parameter sets the type of service, and each provider has its own syntax.
| API_PROVIDER | Service reached | Provider-specific identity parameters |
|---|---|---|
| aws_api_gateway / aws_private_api_gateway | Amazon API Gateway with regional or private endpoints | API_AWS_ROLE_ARN |
| aws_gov_api_gateway / aws_gov_private_api_gateway | Amazon API Gateway with U.S. government GovCloud endpoints | API_AWS_ROLE_ARN |
| azure_api_management | Azure API Management services | AZURE_TENANT_ID, AZURE_AD_APPLICATION_ID |
| google_api_gateway | Google API Gateway | GOOGLE_AUDIENCE |
| git_https_api | A Git repository API | ALLOWED_AUTHENTICATION_SECRETS, or API_USER_AUTHENTICATION |
| external_mcp | An external MCP server used by MCP Connectors | API_USER_AUTHENTICATION (OAUTH2 or OAUTH_DYNAMIC_CLIENT) |
For external functions, the API integration is the object Snowflake checks before any outbound call leaves the account. API_ALLOWED_PREFIXES limits which proxy endpoints and resources the functions using the integration may call. Each URL in the list works as a prefix: allowing https://xyz.amazonaws.com/production/ allows every resource under that path. API_BLOCKED_PREFIXES (Azure, Google and Git syntax) excludes paths that would otherwise be allowed. The prefix list is enforced by Snowflake regardless of what the attached cloud role could reach, so it is the place to tighten scope. Two smaller rules: an Azure API integration can authenticate to only one tenant, and setting ENABLED = FALSE breaks every external function that relies on the integration.
Checkpoint 3 of 8· Check yourself
A review requires that an external function may call only two production paths on an Amazon API Gateway, even though its IAM role has broader permissions. Which setting enforces this in Snowflake?
API_ALLOWED_PREFIXES limits which proxy endpoints functions using the integration may reference, and the docs say to keep it as narrow as practical. STORAGE_ALLOWED_LOCATIONS belongs to storage integrations.
“To maximize security, you should restrict allowed locations as narrowly as practical.”Source: docs.snowflake.com
Checkpoint 4 of 8· Fill the gap
This API integration will be used for a Git repository that authenticates with a personal access token stored in a secret. Which value completes the API_PROVIDER line?
CREATE [ OR REPLACE ] API INTEGRATION [ IF NOT EXISTS ] <integration_name>
API_PROVIDER = ?
API_ALLOWED_PREFIXES = ('<...>')
[ API_BLOCKED_PREFIXES = ('<...>') ]
[ ALLOWED_AUTHENTICATION_SECRETS = ( { <secret_name> [, <secret_name>, ... ] } ) | all | none ]
ENABLED = { TRUE | FALSE }
[ COMMENT = '<string_literal>' ]
;Git repository integrations use API_PROVIDER = git_https_api, written without quotation marks. ALLOWED_AUTHENTICATION_SECRETS then lists the secrets that may hold the token.
Source: docs.snowflake.comCheckpoint 5 of 8· Exam question
A data science team writes Python automation scripts that run parameterized SQL, fetch results into pandas DataFrames, and manage sessions as part of a CI pipeline. Which connectivity option is purpose-built for this workflow?
Correct answer: A — The Python connector, since it offers a native session and query API with built-in helpers for loading results into DataFrames.
- A. The Python connector is Snowflake's native Python library, offering session management, parameterized query execution, and built-in helpers to load results into pandas DataFrames, matching every part of the scenario.
- B. ODBC can be called from Python but is not Python's native interface and is not the only scripted option; the Python connector avoids the extra wrapper layer and is purpose-built for this case.
- C. SnowSQL is a command-line client for interactive or scripted SQL execution, but it does not return DataFrame objects to a Python process, so it does not satisfy the DataFrame integration requirement.
- D. .NET is Microsoft's framework for building C#/.NET applications and has no native pandas DataFrame support, so it cannot serve this Python-based CI automation requirement.
Sources3
3.Git integration: a repository clone built on an API integration
Git integration is not a third kind of integration object. It is a Git repository clone, a schema-level object (GIT REPOSITORY) that sits on top of an API integration whose API_PROVIDER is git_https_api. Once set up, Snowflake keeps a clone of your remote repository that includes all branches, tags and commits. Snowflake then acts as one more client of that repository, so your local development tools carry on working as before.
Supported platforms are GitHub, GitLab, BitBucket, Azure DevOps and AWS CodeCommit, including instances at custom URLs, such as a corporate Git server. The origin URL must use HTTPS. Repositories larger than 2 GB are not supported.
Creating a clone takes three things: the ORIGIN URL, the API_INTEGRATION, and optionally GIT_CREDENTIALS. GIT_CREDENTIALS names a Snowflake secret, ideally holding a personal access token. If the remote needs authentication, the secret must already exist and must be listed in the API integration's ALLOWED_AUTHENTICATION_SECRETS. That rule explains a common failure: the token and the integration can each be correct and authentication still fails, because the two objects were never linked. The creating role needs CREATE GIT REPOSITORY on the schema plus USAGE on both the integration and the secret.
CREATE OR REPLACE GIT REPOSITORY snowflake_extensions
API_INTEGRATION = git_api_integration
GIT_CREDENTIALS = git_secret
ORIGIN = 'https://github.com/my-account/snowflake-extensions.git';Checkpoint 6 of 8· Check yourself
A Git repository clone for a private Azure DevOps repo fails to fetch with an authentication error. The token in the secret is valid and the API integration's prefixes are correct. What should be checked next?
GIT_CREDENTIALS only works with a secret that the API integration explicitly allows. Custom URLs are supported, and Azure DevOps repositories still use git_https_api.
“The secret you specify here must be a secret specified by the ALLOWED_AUTHENTICATION_SECRETS parameter of the API integration”Source: docs.snowflake.com
After the clone exists, you work with it much like a stage. The command family is CREATE, ALTER, DESCRIBE and DROP GIT REPOSITORY, plus SHOW GIT REPOSITORIES, SHOW GIT BRANCHES and SHOW GIT TAGS. ALTER GIT REPOSITORY … FETCH pulls the latest branches, tags and commits from the remote. You list files with the stage path syntax, for example LS @repository_name/branches/branch_name, and the same path pattern works for tags and commits. Paths into the clone can also point to handler files for procedures and UDFs.
To run a SQL script kept in the repository, use EXECUTE IMMEDIATE FROM with a path into the clone. It can be run from any Snowflake session, so you can keep a script such as a new-account setup in Git. The role running it needs the privileges for every statement in the file. In the other direction, Workspaces, Streamlit apps and Snowflake notebooks can commit and push changes back to the remote repository.
EXECUTE IMMEDIATE FROM @snowflake_extensions/branches/main/sql/create-database.sql;Checkpoint 7 of 8· Put it in order
Put these steps in order to run a new setup.sql script from a remote GitHub repository in Snowflake
- 1.Create the git_https_api API integration and the secret that holds the access token
- 2.Run EXECUTE IMMEDIATE FROM @repo/branches/main/scripts/setup.sql
- 3.After pushing setup.sql to the remote, run ALTER GIT REPOSITORY … FETCH
- 4.Run CREATE GIT REPOSITORY with ORIGIN, API_INTEGRATION and GIT_CREDENTIALS
The clone references the integration and the secret, so those come first. FETCH brings the newly pushed file into the clone, and only then can EXECUTE IMMEDIATE FROM find and run it.
“Before creating a Git repository clone, you’ll need to create a secret (if the remote repository requires authentication) and an API integration.”Source: docs.snowflake.com
Checkpoint 8 of 8· Exam question
Which of the following is a Snowflake-provided native connector for a specific data source, rather than a general-purpose driver library for building custom client applications?
Correct answer: A — Kafka connector
- A. The Kafka connector is a native connector that reads from Kafka topics and loads data into Snowflake tables automatically, rather than a generic library used to build custom client applications.
- B. JDBC is a general-purpose driver library that lets developers build or connect custom Java-based client applications to Snowflake, not a purpose-built source connector.
- C. The Python connector is a general-purpose driver library for writing custom Python applications against Snowflake, not a dedicated source-system connector.
- D. ODBC is a general-purpose driver library used by many client tools to build database connectivity, not a native connector tied to a specific data source.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Every S3 storage integration gets its own Snowflake IAM user, so separate integrations are isolated at the AWS identity level.Why is that wrong?
Snowflake creates one IAM user per account, and all S3 storage integrations reference it. Separation comes from bucket permissions and STORAGE_ALLOWED_LOCATIONS.
Covered in Storage integration: delegated access to cloud storage
2.Git integration is its own integration type, created with a dedicated CREATE ... INTEGRATION statement separate from API integrations.Why is that wrong?
A Git repository clone references an ordinary API integration whose API_PROVIDER is git_https_api.
Covered in Git integration: a repository clone built on an API integration
3.Any secret containing a valid token can be passed as GIT_CREDENTIALS.Why is that wrong?
The secret must also be listed in the API integration's ALLOWED_AUTHENTICATION_SECRETS. A valid token in a secret that isn't listed still fails.
Covered in Git integration: a repository clone built on an API integration
Practise it for real
Connect a GitHub repository to Snowflake and run a SQL script from it
1.Create an API integration with API_PROVIDER = git_https_api. Set API_ALLOWED_PREFIXES to your organization's GitHub URL and list your secret's name in ALLOWED_AUTHENTICATION_SECRETS.
Why: The Git repository clone needs a git_https_api integration, and the integration decides which secrets and URLs are allowed.
You should see: The integration appears in SHOW INTEGRATIONS.
2.Create a secret of TYPE password that holds your GitHub username and a personal access token, and make sure its name matches the one listed in ALLOWED_AUTHENTICATION_SECRETS.
Why: GIT_CREDENTIALS must name a secret the integration allows, and the docs recommend a personal access token as the password.
You should see: The secret exists in a schema your role can use.
3.Run CREATE GIT REPOSITORY with ORIGIN set to the repository's HTTPS URL, API_INTEGRATION set to your integration, and GIT_CREDENTIALS set to your secret.
Why: This creates the clone in Snowflake, including its branches, tags and commits.
You should see: SHOW GIT REPOSITORIES lists the new clone, and SHOW GIT BRANCHES lists your branches.
4.Push a setup.sql file to the remote, then run ALTER GIT REPOSITORY <name> FETCH.
Why: The clone only sees remote changes after a fetch.
You should see: LS @<name>/branches/main; lists the new file.
5.Run EXECUTE IMMEDIATE FROM @<name>/branches/main/setup.sql; using a role that can run every statement in the file.
Why: This runs the script stored in Git from a Snowflake session.
You should see: The result of the file's last statement is returned.
Stuck? Get a nudge
If CREATE GIT REPOSITORY or FETCH fails with an authentication error, check ALLOWED_AUTHENTICATION_SECRETS on the integration before you check the token.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“This option allows users to avoid supplying credentials when creating stages or when loading or unloading data.”
↩︎ Storage integration: delegated access to cloud storage“The URL in the stage definition must align with the storage location specified for the STORAGE_ALLOWED_LOCATIONS parameter.”
↩︎ Storage integration: delegated access to cloud storage“use the CREDENTIALS parameter in the CREATE STAGE command (rather than using a storage integration)”
↩︎ Storage integration: delegated access to cloud storage“Existing stages that reference this integration cannot access the storage location in the stage definition.”
↩︎ Checkpoint - 2.
“Snowflake verifies the permissions granted to the IAM user on the bucket before allowing or denying access.”
↩︎ Storage integration: delegated access to cloud storage“Snowflake creates a single IAM user that is referenced by all S3 storage integrations in your Snowflake account.”
↩︎ Exam trap 1“Snowflake creates a single IAM user that is referenced by all S3 storage integrations in your Snowflake account.”
↩︎ Prediction - 3.
“An API integration object stores information about a service reached via HTTPS API”
↩︎ API integration: reaching services over HTTPS“Explicitly limits external functions that use the integration to reference one or more HTTPS proxy service endpoints”
↩︎ API integration: reaching services over HTTPS“If the API integration is disabled, any external function that relies on it will not work.”
↩︎ API integration: reaching services over HTTPS“To maximize security, you should restrict allowed locations as narrowly as practical.”
↩︎ Checkpoint - 4.
“The clone includes all branches, tags, and commits from the remote repository.”
↩︎ Git integration: a repository clone built on an API integration“Commit and push changes to the remote repository from Workspaces, Streamlit apps, and Snowflake notebooks.”
↩︎ Git integration: a repository clone built on an API integration - 5.
“Git repositories larger than 2 GB aren’t supported.”
↩︎ Git integration: a repository clone built on an API integration“The API integration you specify here must have an API_PROVIDER parameter whose value is set to git_https_api.”
↩︎ Exam trap 2“The secret you specify here must be a secret specified by the ALLOWED_AUTHENTICATION_SECRETS parameter of the API integration”
↩︎ Exam trap 3“The secret you specify here must be a secret specified by the ALLOWED_AUTHENTICATION_SECRETS parameter of the API integration”
↩︎ Checkpoint - 6.
“You can use EXECUTE IMMEDIATE FROM to execute code in a Git repository clone.”
↩︎ Git integration: a repository clone built on an API integration
Also cited
“Before creating a Git repository clone, you’ll need to create a secret (if the remote repository requires authentication) and an API integration.”
↩︎ Checkpoint