What you will be able to do
- Promote failover groups and connections safely when refreshes are in progress
- Monitor refresh phases during the transition and know what the documentation does not cover
- Re-establish cloud trust for replicated storage integrations
- Validate network policies, security integrations, roles, and client redirection on the new primary
1.Promoting the target account
Failover always runs from the target account. In Snowsight, the recommended path is a bulk failover: select all the relevant failover groups and connections and promote them together. After promotion, the objects in those groups become writable in the new primary and read-only in the former primary. In SQL, a role with the FAILOVER privilege runs ALTER FAILOVER GROUP … PRIMARY for each group, and an ACCOUNTADMIN runs ALTER CONNECTION … PRIMARY to move the clients.
The usual obstacle is a refresh that is still running. Snowflake blocks promotion while a secondary group is being refreshed, and the command returns an error. To clear it, you can run ALTER FAILOVER GROUP myfg SUSPEND, which stops future refreshes but still waits for the current one. You can also run SUSPEND IMMEDIATE, which also cancels a scheduled refresh that is in progress. Then confirm that nothing is running before you promote:
SELECT phase_name, start_time, job_uuid FROM TABLE(INFORMATION_SCHEMA.REPLICATION_GROUP_REFRESH_HISTORY('myfg')) WHERE phase_name <> 'COMPLETED' and phase_name <> 'CANCELED';Checkpoint 1 of 5· Put it in order
Order the SQL failover steps for failover group myfg when a refresh might be running
- 1.Query REPLICATION_GROUP_REFRESH_HISTORY to confirm no refresh is in progress
- 2.ALTER FAILOVER GROUP myfg PRIMARY;
- 3.ALTER FAILOVER GROUP myfg RESUME; in each target account with a secondary
- 4.ALTER FAILOVER GROUP myfg SUSPEND IMMEDIATE;
Promotion fails while a refresh runs, so suspend and verify first. Refreshes are suspended by the failover and have to be resumed afterwards.
“ALTER FAILOVER GROUP … RESUME must be executed in each target account with a secondary failover group to resume automatic refreshes.”Source: docs.snowflake.com
Sources1
2.Monitoring the transition
While groups are switching over, what you watch most closely is the phase of each refresh. Canceling a refresh in most phases is safe, but not in the two download phases. A refresh that has reached SECONDARY_DOWNLOADING_METADATA or SECONDARY_DOWNLOADING_DATA runs to completion even if the source is unavailable. Waiting for it leaves the replicas consistent, and afterwards you can resume or replay ingest pipelines.
SELECT phase_name, start_time, end_time
FROM TABLE(
INFORMATION_SCHEMA.REPLICATION_GROUP_REFRESH_PROGRESS('myfg')
);Two more records are worth reviewing. Snowsight's replication history can be filtered by status (Refresh Cancelled, Refresh Failed, Refresh In Progress, Refresh Successful). The refresh history function can also list canceled operations, and each of those may have left the target in an inconsistent state. Treat an unexpected failure or cancellation during the transition as an anomaly to investigate.
The exam guide also asks you to monitor audit logs for anomalies such as unusual logins or access. The documentation provided for this lesson does not cover that kind of monitoring. It describes monitoring only through refresh progress and history, so this lesson goes no further.
Checkpoint 2 of 5· Check yourself
A scheduled refresh for myfg is in the SECONDARY_DOWNLOADING_DATA phase when the source region goes down. What is the documented safe action?
Canceling during either download phase risks an inconsistent target. The phase finishes even without the source, so wait for it.
“canceling a refresh operation in the SECONDARY_DOWNLOADING_METADATA or SECONDARY_DOWNLOADING_DATA phase might result in an inconsistent state on the target account.”Source: docs.snowflake.com
Sources2
3.Re-establishing trust for external resources
Some security configuration lives outside Snowflake, so replication cannot copy it. A replicated storage integration is the clearest example, and it is what external stages depend on. The replica has its own identity and access management (IAM) entity, separate from the primary integration's. The cloud provider has only ever trusted the primary's identity. Until you update your cloud provider permissions to grant the replicated integration access, external stages in the new primary cannot reach the storage.
The work is similar to granting access in the source account, using the same S3, Google Cloud Storage, or Azure procedures, and you only do it once per target account. Do it before the outage. If you replicate an external stage with a directory table that auto-refreshes, you also have to set up automated refresh for the secondary directory table.
Checkpoint 3 of 5· Check yourself
After failover, COPY from an external stage fails with an access error in the new primary, although the stage and its storage integration both replicated. What is the most likely cause?
The replicated integration has a different IAM identity, so the cloud storage trust must be configured for it separately.
“you must update your cloud provider permissions to grant the replicated integration access to your cloud storage.”Source: docs.snowflake.com
Sources3
4.Post-failover validation audit
Once the new primary is serving traffic, audit it and do not assume anything carried over.
Network policies. Replicating a network policy carries both the policy object and its references and assignments. A policy assigned to a user lands on that user in the target, provided the user exists in both accounts. Confirm that the policies are present and still assigned to the users and the account you expect.
Security integrations. Each type behaves differently. SAML2 integrations that specify the connection URL work on the new primary connection after failover. Custom OIDC integrations (OIDC_PROVIDER='CUSTOM') are not Client Redirect-aware. You must register the target account's redirect URI with your IdP, which you get from DESC INTEGRATION. After you promote, sign in through SSO again to confirm it works. For OAuth, reconnect with the client to confirm.
Roles and permissions. Compare SHOW USERS with your baseline. Confirm that the FAILOVER and REPLICATE grants exist in this account, because neither privilege replicates.
Client redirection. Run SHOW CONNECTIONS in the promoted account. is_primary should be true there, and failover_allowed_to_accounts should list the accounts the connection can still redirect to. With private connectivity, also confirm that the DNS CNAME record was updated.
Checkpoint 4 of 5· Match them up
Match each replicated object to its post-failover validation point
Tap a term, then the definition that fits it.
Each object needs its own check. Only the custom OIDC integration requires work at the IdP after it replicates.
“OIDC integrations using OIDC_PROVIDER='CUSTOM' are not Client Redirect-aware.”Source: docs.snowflake.com
Checkpoint 5 of 5· Exam question
A company wants to add network policies, users, roles, and security integrations to an existing failover group. What account requirement applies?
Correct answer: A — Every account in the group must run Business Critical Edition or higher, because security object replication is limited to that tier.
- A. Correct. Replicating users, roles, network policies, and security integrations in failover groups requires Business Critical Edition or higher.
- B. Incorrect. A mixed Enterprise and Standard setup does not satisfy the requirement; the higher edition is needed for the accounts that take part in the group.
- C. Incorrect. The requirement is not confined to the targets, because a failover can promote any member to source and that account must support the same object types.
- D. Incorrect. Organization-level replication enablement is separate; it does not lift the edition requirement for security object types.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.In an outage you can promote a secondary failover group at any moment, even mid-refresh.Why is that wrong?
Promotion is blocked while a refresh is running. You must wait, or suspend the group, before ALTER FAILOVER GROUP … PRIMARY succeeds.
Covered in Promoting the target account
2.A replicated storage integration reuses the primary's cloud identity, so existing bucket trust policies keep working.Why is that wrong?
The replica has its own IAM entity, which needs its own one-time trust configuration with the cloud provider.
Covered in Re-establishing trust for external resources
3.Every replicated SSO integration follows Client Redirect automatically.Why is that wrong?
Custom-provider OIDC integrations are tied to the account where they were created, so the target account's redirect URI has to be registered with the IdP.
Covered in Post-failover validation audit
Practise it for real
Run a controlled failover test of a replicated security integration and confirm that clients follow the connection URL
1.In the target account, run SHOW USERS and SHOW INTEGRATIONS and note the counts.
Why: You need a baseline to prove the refresh brought the security objects across.
You should see: Counts recorded before any replication.
2.Run ALTER FAILOVER GROUP fg REFRESH; in the target account, then SHOW INTEGRATIONS, SHOW USERS and DESCRIBE INTEGRATION on the replicated integration.
Why: This confirms that the integration and new users arrived through the failover group.
You should see: One new integration and the added users appear.
3.Query REPLICATION_GROUP_REFRESH_HISTORY for fg, filtering out COMPLETED and CANCELED phases.
Why: Promotion fails while a refresh is in progress.
You should see: No rows returned.
4.Run ALTER FAILOVER GROUP fg PRIMARY; and then ALTER CONNECTION global PRIMARY;
Why: This promotes the objects and moves clients to the target account.
You should see: SHOW CONNECTIONS in the target shows is_primary true.
5.Sign in through SSO using the connection URL, then fail back by promoting the original account's group and connection.
Why: The test is only complete once both directions have worked.
You should see: SSO succeeds in both directions.
Stuck? Get a nudge
If promotion errors that the group is being refreshed, suspend the group, wait for the in-progress refresh to finish, and retry.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“select all the applicable failover groups and connections and promote them all at the same time.”
↩︎ Promoting the target account“Those objects become read-only on the account that formerly was the primary and is now a secondary account.”
↩︎ Promoting the target account“The example in this section must be executed by a role with the FAILOVER privilege.”
↩︎ Promoting the target account“Snowflake prevents failover if a refresh operation is in progress.”
↩︎ Exam trap 1“On failover, scheduled refreshes on all secondary failover groups are suspended.”
↩︎ Prediction“ALTER FAILOVER GROUP … RESUME must be executed in each target account with a secondary failover group to resume automatic refreshes.”
↩︎ Checkpoint“canceling a refresh operation in the SECONDARY_DOWNLOADING_METADATA or SECONDARY_DOWNLOADING_DATA phase might result in an inconsistent state on the target account.”
↩︎ Checkpoint - 2.
“To monitor the progress of a replication or failover group refresh, query the REPLICATION_GROUP_REFRESH_PROGRESS”
↩︎ Monitoring the transition“Refresh Cancelled Refresh Failed Refresh In Progress Refresh Successful”
↩︎ Monitoring the transition - 3.
“You only need to configure this trust relationship on target accounts one time.”
↩︎ Re-establishing trust for external resources“The replicated integration has its own identity and access management (IAM) entity”
↩︎ Exam trap 2“you must update your cloud provider permissions to grant the replicated integration access to your cloud storage.”
↩︎ Checkpoint - 4.
“Replicating a network policy replicates the network policy object and any network policy references/assignments.”
↩︎ Post-failover validation audit“After failover, SAML SSO works on the new primary connection.”
↩︎ Post-failover validation audit“OIDC integrations using OIDC_PROVIDER='CUSTOM' are not Client Redirect-aware.”
↩︎ Exam trap 3“OIDC integrations using OIDC_PROVIDER='CUSTOM' are not Client Redirect-aware.”
↩︎ Checkpoint - 5.
“Indicates whether the connection is a primary connection.”
↩︎ Post-failover validation audit“A list of any accounts that the primary connection can redirect to.”
↩︎ Post-failover validation audit