What you will be able to do
- Create and configure tags with allowed values and propagation
- Implement a tag-based masking policy and predict which policy wins on a column
- Classify a table with EXTRACT_SEMANTIC_CATEGORIES and apply the results as tags
- Audit reads, writes and policy changes with the ACCESS_HISTORY view
- Enable and manage Trust Center scanner packages in Snowsight
1.Implementing and managing object tags
A tag is a schema-level object that holds a string value. You set it on a database, schema, table or column with CREATE or ALTER on that object. The sources show tags used for several governance jobs. You can mark data with a label like pii = 'sensitive'. You can attach masking policies to a tag so that tagging a column protects it. Classification writes its results as system tags. Every tag change is recorded for audit in ACCESS_HISTORY. Before adding protections, you can check which tags are already set by querying the Account Usage TAG_REFERENCES view.
alter table hr.tables.empl_info alter column email set tag governance.tags.data_category = 'sensitive';You manage a tag with ALTER TAG. ALLOWED_VALUES limits the strings that can be assigned, up to 5,000 values. MULTI_VALUE = TRUE lets a tag hold more than one value, and once set it can't be undone. PROPAGATE makes Snowflake copy the tag automatically from source objects to target objects. When propagated values conflict, ON_CONFLICT decides the result: a fixed string, the first match in the allowed-values order (ALLOWED_VALUES_SEQUENCE), or MERGE, which needs MULTI_VALUE. If ON_CONFLICT isn't set, the value becomes the string CONFLICT. Changes to a tag apply from the next query that uses it.
ALTER TAG [ IF EXISTS ] <name> SET
[ ALLOWED_VALUES '<val_1>' [ , '<val_2>' [ , ... ] ] ]
[ MULTI_VALUE = TRUE ]
[ PROPAGATE = { ON_DEPENDENCY_AND_DATA_MOVEMENT | ON_DEPENDENCY | ON_DATA_MOVEMENT }
[ ON_CONFLICT = { '<string>' | ALLOWED_VALUES_SEQUENCE | MERGE } ] ]
[ COMMENT = '<string_literal>' ]Checkpoint 1 of 9· Check yourself
A tag is set to propagate, and two source objects carry different values for it. No ON_CONFLICT is configured. What value does the target object get?
Without ON_CONFLICT, a conflict between propagated values sets the tag to the string CONFLICT. MERGE and ALLOWED_VALUES_SEQUENCE must be chosen explicitly.
“the value of the tag is set to the string CONFLICT”Source: docs.snowflake.com
2.Tag-based masking policies
With tag-based masking you attach the masking policy to a tag instead of a column. Any column that carries the tag, or sits in a table, schema or database that carries it, is then protected, as long as the column's data type matches the policy signature. A tag can hold only one masking policy per data type, for example one for STRING and one for NUMBER. An External Tokenization policy can be assigned to a tag the same way. Inside the policy, SYSTEM$GET_TAG_ON_CURRENT_COLUMN and SYSTEM$GET_TAG_ON_CURRENT_TABLE let the conditions read the tag value.
Checkpoint 2 of 9· Put it in order
Put the steps for setting up a tag-based masking policy in order
- 1.Create a masking policy with CREATE MASKING POLICY
- 2.Set the tag on the table, schema or database with ALTER TABLE, ALTER SCHEMA or ALTER DATABASE
- 3.Set the masking policy on the tag with ALTER TAG
- 4.Create a tag with CREATE TAG
The tag and the policy must both exist before ALTER TAG can link them. Setting the tag on an object is what applies the protection.
“Set the masking policy on the tag using an ALTER TAG command.”Source: docs.snowflake.com
Some rules decide which policy applies and what you can change:
- Precedence: if a column has a directly assigned masking policy and also a tag-based one, the direct policy wins. - Dropping: a tag can't be dropped while it has a masking policy, and a masking policy can't be dropped while it is assigned to a tag. - Replacing a policy on a tag: either UNSET then SET in two statements, or use FORCE in one. The two-statement route leaves the column unprotected between the statements. FORCE avoids that gap, but if the new policy is stricter, users may suddenly lose access. - Privileges for databases and schemas: the role needs global APPLY MASKING POLICY plus APPLY TAG, or, as schema owner, APPLY MASKING POLICY plus APPLY on the tag.
ALTER TAG security SET MASKING POLICY ssn_mask_2 FORCE;Checkpoint 3 of 9· Check yourself
Column SSN has masking policy ssn_full_mask attached directly. Its table also carries tag pii, which has STRING policy ssn_partial_mask. Which policy masks SSN at query time?
When both exist on a column, the directly assigned masking policy wins over the tag-based one.
“the directly assigned masking policy takes precedence over the masking policy assigned to the tag”Source: docs.snowflake.com
Checkpoint 4 of 9· Exam question
A production table `hr.emp` has an `ssn` column. Role `HR_ADMIN` must see full values, role `ANALYST` must see only the last four digits, and every other role must see a fixed mask. Which approach is MOST appropriate?
Correct answer: C — Create a masking policy on the VARCHAR column whose CASE expression checks CURRENT_ROLE() and returns the full value, the last four digits, or a fixed mask
- A. Row access policies filter whole rows; they cannot transform a column value into a partial or masked form. Column-level output needs a masking policy.
- B. Matching hard-coded user names is brittle and gives no way to separate analysts from other roles or to apply a fixed mask to everyone else. Role-based logic is the supported pattern.
- C. A masking policy evaluated at query time with a CASE on CURRENT_ROLE() (or IS_ROLE_IN_SESSION) can return three different outputs for three role groups. This is the standard column-level security design.
- D. Object privileges such as SELECT never change what a masking policy returns. The policy body alone decides the output, so HR_ADMIN would still see only four digits.
Checkpoint 5 of 9· Exam question
A masking policy `pii_mask` is already attached to `sales.customers.email`. An engineer runs `ALTER TABLE sales.customers MODIFY COLUMN email SET MASKING POLICY new_mask;` and the statement fails. What is the MOST appropriate way to swap the policy in one statement?
Correct answer: A — Run the same ALTER TABLE statement with the FORCE keyword so that the existing policy is replaced by the new one atomically in a single step
- A. SET MASKING POLICY ... FORCE replaces the existing policy on the column atomically, so there is no window in which the column is unprotected. Without FORCE the statement errors when a policy is already set.
- B. The error is caused by the column already having a policy, not by a missing privilege. Granting more privileges would not change that behaviour.
- C. CREATE OR REPLACE on a policy that is still attached to columns is not allowed, and it would not create the different new_mask policy being requested anyway.
- D. Dropping and re-creating the column destroys the data and is far more disruptive than needed. A column holds only one masking policy, and a supported replacement path exists.
3.Data classification with EXTRACT_SEMANTIC_CATEGORIES
Classification samples a table, external table, view or materialized view and suggests what each column contains. EXTRACT_SEMANTIC_CATEGORIES returns a JSON object per supported column. Columns that are entirely NULL are skipped. Each result includes a semantic_category such as PASSPORT, a privacy_category (IDENTIFIER, QUASI-IDENTIFIER or SENSITIVE), a confidence of HIGH, MEDIUM or LOW, a coverage figure, and a list of alternates. The optional second argument sets how many rows to sample, from 1 to 10,000, with 10,000 as the default. The function needs a running warehouse, which affects cost.
To write the results as tags, call the stored procedure the docs name ASSOCIATE_SEMANTIC_CATEGORY_TAGS. The exam guide calls it ASSOCIATE_SEMANTIC_CATEGORIES. It applies the Classification system tags from the top-level recommendation only. Alternates are not applied. To use an alternate, store the results in a table, edit them, then pass the table in, or set the tag yourself with ALTER TABLE … MODIFY COLUMN … SET TAG. The docs mark both the function and the procedure as legacy: they are no longer updated, and Snowflake recommends other classification methods.
The current approach is automatic sensitive data classification, set up through the Trust Center. Every column identified as sensitive gets two categories: a semantic category (the type of attribute, native or custom) and a privacy category (IDENTIFIER, QUASI_IDENTIFIER or SENSITIVE). Snowflake records them with the system tags SNOWFLAKE.CORE.SEMANTIC_CATEGORY and SNOWFLAKE.CORE.PRIVACY_CATEGORY.
- Classification profile: the settings you choose in the Trust Center are saved as a classification profile, which you can edit later. In the web interface the profile also controls which databases it classifies. You can also create and modify profiles with SQL, but then associating the profile with a database is a separate step.
- Automatic coverage: once a database is associated with a profile, all its tables and views are classified automatically. Views are excluded by default.
- Map to your own tags: you can map user-defined tags to the system tags. For example, whenever SEMANTIC_CATEGORY = 'NAME' is applied, your own tag tag_db.sch.pii = 'Highly confidential' is applied too.
- Tag-based masking: if a masking policy is attached to that user-defined tag, columns are masked as soon as classification applies the tag, including columns in newly added data.
- Custom categories and AI mode: you can create custom categories for data that native categories don't cover. A profile can optionally use AI mode, which uses an LLM to find additional semantic categories.
- Monitoring and cost: in Snowsight, open Governance & security » Trust Center, then the Data Security tab and its Dashboard, and find the Databases monitored by classification tile. SYSTEM$SHOW_SENSITIVE_DATA_MONITORED_ENTITIES('DATABASE') lists the classified databases. Classification consumes serverless compute credits, and AI mode adds token charges.
SELECT EXTRACT_SEMANTIC_CATEGORIES('my_db.my_schema.hr_data', 5000);CALL ASSOCIATE_SEMANTIC_CATEGORY_TAGS('mydb.my_schema.hr_data', EXTRACT_SEMANTIC_CATEGORIES('mydb.my_schema.hr_data'));Checkpoint 6 of 9· Check yourself
EXTRACT_SEMANTIC_CATEGORIES recommends NAME for column LNAME and lists an alternate category. What does ASSOCIATE_SEMANTIC_CATEGORY_TAGS apply when given these results unedited?
The procedure applies only the top-level results. To use an alternate, you edit the stored results first or set the tag yourself.
“Alternate values are not applied.”Source: docs.snowflake.com
4.Auditing with the ACCESS_HISTORY view
After policies and tags are in place, ACCESS_HISTORY records how data was actually used. It is in ACCOUNT_USAGE and covers the last 365 days. An ORGANIZATION_USAGE version also exists. One detail matters on the exam: read queries are logged at two levels. Suppose the chain is base_table » view_1 » view_2 » view_3 and you run select * from view_2. Then view_2 goes in direct_objects_accessed and base_table goes in base_objects_accessed. view_1 and view_3 appear in neither column. Not every QUERY_HISTORY record appears in ACCESS_HISTORY. ACCESS_HISTORY does not include certain short-running transactional queries. The AGGREGATE_ACCESS_HISTORY view covers both analytical and transactional queries, aggregated over time for repeated queries in one-minute intervals.
| Column | What it records |
|---|---|
| direct_objects_accessed | Objects named in the query, explicitly or through * |
| base_objects_accessed | All base data objects needed to run the query |
| objects_modified | Objects written by the query, with baseSources and directSources for column lineage |
| object_modified_by_ddl | DDL, including setting row access or masking policies and tag changes |
| policies_referenced | Enforced masking and row access policies, including those on intermediate objects |
Checkpoint 7 of 9· Check yourself
Objects chain as base_table → view_1 → view_2 → view_3. A user runs SELECT * FROM view_2. What does base_objects_accessed record?
base_objects_accessed records the original source of the data. view_2 goes in direct_objects_accessed, and the intermediate view_1 is not recorded.
“base_table in the base_objects_accessed column because that is the original source of the data in view_2.”Source: docs.snowflake.com
5.Trust Center, Horizon Catalog and lineage in Snowsight
The Trust Center is the Snowsight page for checking account security. To get there, use a role that has the SNOWFLAKE.TRUST_CENTER_ADMIN application role and select Governance & security » Trust Center. Its checks come in scanner packages:
- Security Essentials: the only package on by default. It can't be deactivated, and its regular fixed-schedule runs incur no serverless cost. Any other run does incur charges. - CIS Benchmarks: checks the account against the CIS Snowflake Benchmarks. It runs once a day by default, and you can change the schedule. - Threat Intelligence: flags risky users and logins. - AI Security: monitors AI-related configuration. For Business Critical or VPS accounts with a capacity contract, Snowflake can enable it automatically when AI feature usage is detected, unless it was ever explicitly disabled.
In the Manage scanners tab you can enable packages, enable or disable individual scanners, change schedules and run scans on demand. Results appear on the Violations tab, each with remediation guidance.
Read CIS findings carefully. Benchmark 4.10 raises a violation only if the account has no masking policy at all, and 4.11 does the same for row access policies. Neither checks whether sensitive columns are actually protected.
Checkpoint 8 of 9· Check yourself
A new account has never opened the Trust Center. Which scanner package is already running?
Every package except Security Essentials is deactivated by default. CIS Benchmarks and Threat Intelligence must be enabled by a TRUST_CENTER_ADMIN role.
“Scanner packages are deactivated by default, except for the Security Essentials scanner package.”Source: docs.snowflake.com
For Horizon Catalog, the sources say only two things. Dynamic data masking policies are enforced on Apache Iceberg tables queried from Apache Spark through Horizon Catalog. ACCESS_HISTORY also records events from the Horizon Iceberg REST Catalog API, with event_source = horizon_irc. For data lineage, ACCESS_HISTORY's baseSources and directSources fields record column lineage for writes, and the GET_LINEAGE function returns lineage information upstream or downstream from a Snowflake object. The sources don't describe Universal Search, the Snowsight lineage graph, or Snowsight's tag and policy monitoring pages. Study those from the Snowsight governance documentation rather than relying on this lesson.
Checkpoint 9 of 9· Exam question
A fintech company must keep raw card numbers out of Snowflake entirely, because its compliance team requires that only an external tokenization service can ever reveal them. Which feature MOST appropriately meets this requirement?
Correct answer: A — External tokenization, where tokens are loaded into Snowflake and an external function call to the provider detokenizes values for authorized roles
- A. With external tokenization, data is tokenized before loading and a masking policy calls an external function to detokenize for authorized roles. Plaintext never resides in Snowflake.
- B. Dynamic data masking leaves the real values stored in Snowflake and only changes what is returned, so plaintext still exists in the platform. That fails a rule that raw values must never be stored.
- C. A secure view limiting output still requires the full card number to exist in the base table, so raw values remain stored in Snowflake.
- D. Tri-Secret Secure strengthens at-rest key control but does not stop authorized queries from reading plaintext. It is not a tokenization mechanism.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A tag-based masking policy overrides a masking policy assigned directly to the same column.Why is that wrong?
The directly assigned policy wins. The tag-based policy only protects columns that don't have a policy of their own.
Covered in Tag-based masking policies
2.ASSOCIATE_SEMANTIC_CATEGORY_TAGS writes both the recommended and the alternate categories as tags.Why is that wrong?
It applies only the top-level results. To use an alternate, edit the stored results or set the tag manually.
Covered in Data classification with EXTRACT_SEMANTIC_CATEGORIES
3.If the Trust Center shows no violation for CIS 4.10, sensitive columns are masked.Why is that wrong?
The check only confirms that at least one masking policy exists. It doesn't check whether that policy is attached to anything.
Covered in Trust Center, Horizon Catalog and lineage in Snowsight
Practise it for real
Classify a table, then protect a column with a tag-based masking policy and confirm that the change was audited
1.Run SELECT EXTRACT_SEMANTIC_CATEGORIES('my_db.my_schema.hr_data', 5000); with a warehouse in use
Why: Shows what classification recommends before anything is written
You should see: One JSON entry per supported column, with semantic_category, privacy_category, confidence and alternates
2.Run CALL ASSOCIATE_SEMANTIC_CATEGORY_TAGS('mydb.my_schema.hr_data', EXTRACT_SEMANTIC_CATEGORIES('mydb.my_schema.hr_data'));
Why: Writes the top-level recommendations as Classification system tags
You should see: Columns carry system tags; no alternates are applied
3.Create a tag and a STRING masking policy, run ALTER TAG … SET MASKING POLICY, then set the tag on a table
Why: Links protection to the tag rather than to individual columns
You should see: STRING columns in the table are masked for roles the policy doesn't allow
4.Query ACCOUNT_USAGE.ACCESS_HISTORY and filter on object_modified_by_ddl for your tag
Why: Confirms that the policy and tag changes were recorded for audit
You should see: Records with objectDomain TAG or Table and a maskingPolicies or tags property
Stuck? Get a nudge
If the masked column is still in plain text, check whether it already has a directly assigned masking policy, which takes precedence.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Specifies that the tag will be automatically propagated from source objects to target objects.”
↩︎ Implementing and managing object tags“The maximum number of tag values in this list is 5,000.”
↩︎ Implementing and managing object tags“Replaces a masking policy that is currently set on a tag with a different masking policy in a single statement.”
↩︎ Tag-based masking policies“the value of the tag is set to the string CONFLICT”
↩︎ Checkpoint - 2.
“Query the Account Usage TAG_REFERENCES view to verify the existing tags set on a table or a column in a table.”
↩︎ Implementing and managing object tags“A tag can have only one masking policy per data type.”
↩︎ Tag-based masking policies“could lead to a data leak because the column data is unprotected in the time interval between the UNSET and SET operations”
↩︎ Tag-based masking policies“the directly assigned masking policy takes precedence over the masking policy assigned to the tag”
↩︎ Exam trap 1“Set the masking policy on the tag using an ALTER TAG command.”
↩︎ Checkpoint“the directly assigned masking policy takes precedence over the masking policy assigned to the tag”
↩︎ Checkpoint - 3.
“an external tokenization masking policy can be assigned to a tag to provide tag-based external tokenization”
↩︎ Tag-based masking policies - 4.
“EXTRACT_SEMANTIC_CATEGORIES is a legacy function.”
↩︎ Data classification with EXTRACT_SEMANTIC_CATEGORIES“The possible values are IDENTIFIER, QUASI-IDENTIFIER and SENSITIVE.”
↩︎ Data classification with EXTRACT_SEMANTIC_CATEGORIES - 5.https://docs.snowflake.com/en/sql-reference/stored-procedures/associate_semantic_category_tagsOfficial docs
“Takes the results of the EXTRACT_SEMANTIC_CATEGORIES function on a table/view and applies the results as tags on the supported columns in the table/view.”
↩︎ Data classification with EXTRACT_SEMANTIC_CATEGORIES“Alternate values are not applied.”
↩︎ Exam trap 2“Alternate values are not applied.”
↩︎ Checkpoint - 6.
“those settings are saved as a classification profile.”
↩︎ Data classification with EXTRACT_SEMANTIC_CATEGORIES“You can map user-defined tags to system-defined classification tags.”
↩︎ Data classification with EXTRACT_SEMANTIC_CATEGORIES“all the tables and views in that database are being automatically classified”
↩︎ Data classification with EXTRACT_SEMANTIC_CATEGORIES“Sensitive data classification consumes credits as it uses serverless compute resources”
↩︎ Data classification with EXTRACT_SEMANTIC_CATEGORIES - 7.
“within the last 365 days (1 year).”
↩︎ Auditing with the ACCESS_HISTORY view“horizon_irc — Events generated by calls made to the Horizon Iceberg REST Catalog API.”
↩︎ Trust Center, Horizon Catalog and lineage in Snowsight - 8.
“Records in the Account Usage QUERY_HISTORY view do not always get recorded in the ACCESS_HISTORY view.”
↩︎ Auditing with the ACCESS_HISTORY view“base_table in the base_objects_accessed column because that is the original source of the data in view_2.”
↩︎ Checkpoint - 9.
“does not include certain short-running transactional queries”
↩︎ Auditing with the ACCESS_HISTORY view“aggregated over time for repeated queries in one-minute intervals.”
↩︎ Auditing with the ACCESS_HISTORY view“These columns facilitate column lineage.”
↩︎ Auditing with the ACCESS_HISTORY view - 10.
“Switch to a role with the SNOWFLAKE.TRUST_CENTER_ADMIN application role granted to it.”
↩︎ Trust Center, Horizon Catalog and lineage in Snowsight“Snowflake automatically enables the Trust Center AI Security scanner package when AI feature usage is detected”
↩︎ Trust Center, Horizon Catalog and lineage in Snowsight - 11.
“This scanner package runs once a day by default, but you can change the schedule.”
↩︎ Trust Center, Horizon Catalog and lineage in Snowsight“Is enabled by default. You can’t deactivate it.”
↩︎ Trust Center, Horizon Catalog and lineage in Snowsight“Runs regularly on a fixed schedule without incurring any serverless compute cost.”
↩︎ Trust Center, Horizon Catalog and lineage in Snowsight“Trust Center displays a violation if the account does not have at least one masking policy”
↩︎ Trust Center, Horizon Catalog and lineage in Snowsight“Trust Center displays a violation if the account doesn’t have at least one row access policy”
↩︎ Trust Center, Horizon Catalog and lineage in Snowsight“displays a violation if the account does not have at least one masking policy, but does not evaluate whether sensitive data is protected appropriately”
↩︎ Exam trap 3“Scanner packages are deactivated by default, except for the Security Essentials scanner package.”
↩︎ Checkpoint - 12.
“Given a Snowflake object, returns data lineage information upstream or downstream from that object.”
↩︎ Trust Center, Horizon Catalog and lineage in Snowsight - 13.
“Snowflake supports enforcing dynamic data masking policies on Apache Iceberg tables that you query from Apache Spark™ through Snowflake Horizon Catalog.”
↩︎ Trust Center, Horizon Catalog and lineage in Snowsight