What you will be able to do
- Explain the semantic and privacy categories and the SNOWFLAKE.CORE system tags that classification applies
- Set up automatic classification with a classification profile and know its defaults and side effects
- Create a custom classifier with regular expressions and attach it to a profile
- Classify a single object on demand with SYSTEM$CLASSIFY
- Connect classification to governance policy through tag maps and tag-based masking
1.What classification produces: categories and system tags
Sensitive data classification finds sensitive columns for you, so that governance controls such as tags and masking policies can follow them. Each column it identifies as sensitive gets two labels. The semantic category says what kind of personal attribute the column holds. Snowflake ships native categories such as NAME and NATIONAL_IDENTIFIER, and you can add custom ones. The privacy category says how sensitive the attribute is: IDENTIFIER, QUASI_IDENTIFIER, or SENSITIVE. SENSITIVE is a generic, non-identifying category for data such as health records or salary.
Both labels are stored as system-defined tags on the column: SNOWFLAKE.CORE.SEMANTIC_CATEGORY and SNOWFLAKE.CORE.PRIVACY_CATEGORY. Because these are ordinary tag assignments, the tagging rules still apply. A classified value overrides an inherited one, and tag audits report it with APPLY_METHOD = CLASSIFIED.
Checkpoint 1 of 6· Match them up
Match each classification term to its meaning
Tap a term, then the definition that fits it.
Every sensitive column gets both a semantic category and a privacy category, and each is recorded as a system tag.
“every column that is identified as containing sensitive data is assigned two categories: a semantic category and a privacy category.”Source: docs.snowflake.com
Sources1
2.Automatic classification with classification profiles
Automatic classification is driven by a classification profile. Settings you choose in the Trust Center are saved as a profile, and in the web interface that profile also decides which databases are classified. In SQL, you create the profile as an instance of SNOWFLAKE.DATA_PRIVACY.CLASSIFICATION_PROFILE, which needs the CLASSIFICATION_ADMIN database role. Associating it with a database is a separate step.
Once a database is associated with a profile, Snowflake automatically classifies all the tables and views in it using serverless compute, which consumes credits. Views are excluded by default (classify_views), because classifying a view can cost more than classifying a table. The auto_tag setting controls whether the recommended system tags are actually written. ai_mode adds LLM-based detection of extra semantic categories using the openai-gpt-5-mini model, which adds token charges.
Two operational details are easy to miss. First, running CREATE OR REPLACE on a profile detaches it from every database and schema, which switches automatic classification off. Second, a database shows as partially monitored when a profile was set with SQL directly on one of its schemas rather than on the database.
CREATE OR REPLACE SNOWFLAKE.DATA_PRIVACY.CLASSIFICATION_PROFILE
my_classification_profile(
{
'minimum_object_age_for_classification_days': 0,
'maximum_classification_validity_days': 30,
'auto_tag': true,
'classify_views': false
});SELECT SYSTEM$SHOW_SENSITIVE_DATA_MONITORED_ENTITIES('DATABASE');Checkpoint 2 of 6· Check yourself
An engineer runs CREATE OR REPLACE on an existing classification profile to change one setting. What else happens?
Replacing the profile detaches it everywhere, so you have to associate it with the databases again. To change one setting, use a profile method such as SET_AUTO_TAG or SET_CLASSIFY_VIEWS instead.
“Executing a CREATE OR REPLACE command removes the classification profile from all databases and schemas, which turns off automatic classification.”Source: docs.snowflake.com
Checkpoint 3 of 6· Exam question
A data engineering team set `PROPAGATE = ON_DEPENDENCY` on the tag `pii_level`. Source table `RAW.CUSTOMERS` has tagged columns. A pipeline now runs `CREATE TABLE CURATED.CUSTOMERS_COPY AS SELECT * FROM RAW.CUSTOMERS`, and the new table's columns carry no tags, while a view built on the same source does. What change makes the CTAS table receive the tags?
Correct answer: A — Switch the tag to `PROPAGATE = ON_DEPENDENCY_AND_DATA_MOVEMENT`, issued by the tag owner holding account-level `APPLY TAG`.
- A. Correct. ON_DEPENDENCY covers views and similar dependent objects only. Data movement such as CTAS needs ON_DATA_MOVEMENT or the combined mode, set by the tag owner with APPLY TAG.
- B. Incorrect. Time Travel retention has no effect on tag propagation and does not carry column tag metadata into a new table.
- C. Incorrect. DML only propagates tags when the tag's PROPAGATE setting includes data movement, so swapping statement types would not help while the setting stays ON_DEPENDENCY.
- D. Incorrect. The creating role is not the deciding factor; propagation on data movement is governed by the tag's PROPAGATE setting, not by who runs the CTAS.
Sources1
3.Custom classification with CUSTOM_CLASSIFIER
If no native category fits your data, such as internal employee IDs or medical codes, you define a custom classifier. It is an instance of SNOWFLAKE.DATA_PRIVACY.CUSTOM_CLASSIFIER that holds a custom semantic category, regular expressions, and one of the predefined privacy categories. You need the SNOWFLAKE.CLASSIFICATION_ADMIN database role to create an instance. The PRIVACY_USER instance role lets other roles call ADD_REGEX, LIST and DELETE_CATEGORY on it.
ADD_REGEX takes a value regex, an optional column-name regex, and a THRESHOLD. The default threshold is 0.8, meaning 80% of sampled values must match. When the custom classifier is compared with Snowflake's native categories, the custom category wins only if the values meet the threshold and, where a name regex was supplied, the column name matches too.
CALL internal_ids!ADD_REGEX( SEMANTIC_CATEGORY => 'EMPLOYEE_ID', PRIVACY_CATEGORY => 'IDENTIFIER', VALUE_REGEX => '^[0-9]{6}$', COL_NAME_REGEX => 'EMP.*ID.*', DESCRIPTION => 'Add a regex to identify employee IDs in a column', THRESHOLD => 0.8 );| Name regex provided | Values match >= threshold | Column name matches | Recommendation |
|---|---|---|---|
| True | True | True | Custom category |
| True | False | True | Snowflake category |
| True | True | False | Snowflake category |
| True | False | False | Snowflake category |
| False | True | Not applicable | Custom category |
| False | False | Not applicable | Snowflake category |
To use a custom classifier for automatic classification, add it to a profile with the custom_classifiers key. The profile stores a copy of the classifier's definition, not a reference to it. If you edit the classifier later, call SET_CUSTOM_CLASSIFIERS on the profile, or it will keep using the old definition. Custom classifier instances are replicated with their database and cloned with their schema.
CREATE SNOWFLAKE.DATA_PRIVACY.CLASSIFICATION_PROFILE my_classification_profile(
{
'minimum_object_age_for_classification_days':0,
'auto_tag':true,
'custom_classifiers': {
'medical_codes': medical_codes!list(),
'finance_codes': finance_codes!list()
}
}
);Checkpoint 4 of 6· Put it in order
Put the documented steps for classifying a table with a custom classifier in order
- 1.Call SYSTEM$CLASSIFY_SCHEMA to classify the table
- 2.Create a schema to store custom classifier instances
- 3.Call ADD_REGEX to define the semantic category, privacy category and regex
- 4.Grant the SNOWFLAKE.CLASSIFICATION_ADMIN database role to the data owner role
- 5.Call LIST to verify the regular expression
- 6.Create the instance with CREATE SNOWFLAKE.DATA_PRIVACY.CUSTOM_CLASSIFIER
Privileges come first, then the instance and its regex, then verification, and classification runs last.
“Create a custom classifier instance. Add the custom semantic category and regular expressions to the instance. Classify the table.”Source: docs.snowflake.com
4.Manual, on-demand classification with SYSTEM$CLASSIFY
Profiles cover whole databases. SYSTEM$CLASSIFY classifies one table, external table, view or materialized view when you call it, which suits one-off checks and newly loaded tables. Its second argument is either the name of a classification profile, so the call follows that profile's criteria including AI mode, or an options object: - NULL or {}: classify with the default configuration and leave the columns untagged, so you can review the results first. - {'sample_count': n}: sample n rows, where n is from 1 to 10000. - {'auto_tag': true}: write the recommended system tags to the columns when classification finishes. The calling role must own the schema.
This gives you three levels of control: review recommendations only, tag immediately, or reuse a profile's settings.
Checkpoint 5 of 6· Check yourself
A data steward wants SYSTEM$CLASSIFY to write the recommended system tags to a table's columns. Which requirement applies?
auto_tag writes the recommended tags, and it requires the role that owns the schema. NULL sets no tags, and sample_count only controls sampling.
“When you use this argument, call the stored procedure with the role that has the OWNERSHIP privilege on the schema.”Source: docs.snowflake.com
Sources4
5.Wiring classification into governance policies
Classification only protects data once its output drives a policy. The link is a tag map. A profile's tag_map maps your own tags to the SEMANTIC_CATEGORY system tag, so whenever classification tags a column, your tag is applied too.
In column_tag_map, give tag_name alone to apply the user tag to every classified column, with its value copied from SEMANTIC_CATEGORY. Or give tag_value together with semantic_categories to set a fixed value for selected categories, for example 'pii' for EMAIL and NATIONAL_IDENTIFIER. Either supply both of those keys or omit both. If one tag is mapped to different values for the same category, the order of entries decides, so list them from highest to lowest preference.
CREATE SNOWFLAKE.DATA_PRIVACY.CLASSIFICATION_PROFILE my_classification_profile(
{
'minimum_object_age_for_classification_days':0,
'auto_tag':true,
'tag_map':{
'column_tag_map':[
{
'tag_name':'tag_db.sch.pii'
}
]
}
}
);Now attach a masking policy to that user-defined tag with tag-based masking. When classification applies the tag, the column is masked straight away, and as new data lands in the database, newly classified columns pick up the policy without anyone stepping in. If the tag also propagates, the masking follows the data into downstream views and tables, which you can confirm on the Lineage tab.
Checkpoint 6 of 6· Check yourself
You want columns that classification marks as NAME to be masked automatically, including new tables loaded next month. Which design achieves that?
The documented pattern is to map a user-defined tag in the profile and attach the masking policy to that tag with tag-based masking. Classification then applies the tag, and the masking with it, automatically.
“the data will be automatically masked when Snowflake applies the tag as part of the classification process.”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Editing a custom classifier immediately changes how every profile that uses it classifies data.Why is that wrong?
The profile stores a copy of the classifier's definition, not a reference. You must call SET_CUSTOM_CLASSIFIERS to push the new definition into the profile.
Covered in Custom classification with CUSTOM_CLASSIFIER
2.Associating a profile with a database classifies its views as well as its tables by default.Why is that wrong?
Views are excluded unless classify_views is enabled, because classifying a view can cost more than classifying a table.
Covered in Automatic classification with classification profiles
3.Any SYSTEM$CLASSIFY call tags the columns it classifies.Why is that wrong?
With NULL or {} options, no system tags are set. Tags are written only with auto_tag or through a profile.
Covered in Manual, on-demand classification with SYSTEM$CLASSIFY
Practise it for real
Detect six-digit employee IDs with a custom classifier and make them part of automatic classification
1.As ACCOUNTADMIN, run GRANT DATABASE ROLE SNOWFLAKE.CLASSIFICATION_ADMIN TO ROLE data_owner;
Why: Creating a custom classifier instance requires the CLASSIFICATION_ADMIN database role.
You should see: The grant succeeds and data_owner can create CUSTOM_CLASSIFIER instances.
2.As data_owner, create a schema for classifiers and run CREATE OR REPLACE SNOWFLAKE.DATA_PRIVACY.CUSTOM_CLASSIFIER internal_ids();
Why: The instance holds your custom semantic categories and regular expressions.
You should see: SHOW SNOWFLAKE.DATA_PRIVACY.CUSTOM_CLASSIFIER lists INTERNAL_IDS.
3.Call internal_ids!ADD_REGEX with SEMANTIC_CATEGORY 'EMPLOYEE_ID', PRIVACY_CATEGORY 'IDENTIFIER', VALUE_REGEX '^[0-9]{6}$' and THRESHOLD 0.8
Why: This defines what a matching column looks like and how sensitive it is.
You should see: The call returns EMPLOYEE_ID.
4.Run SELECT internal_ids!LIST();
Why: Check the stored regex, threshold and privacy category before relying on them.
You should see: A JSON object with an EMPLOYEE_ID key and its settings.
5.Create a classification profile with 'auto_tag':true and a custom_classifiers entry that uses internal_ids!list()
Why: Automatic classification uses only the custom classifiers stored in the profile.
You should see: The profile is created. Associating it with a database is a separate step.
Stuck? Get a nudge
If you change the regex later, the profile won't notice until you call SET_CUSTOM_CLASSIFIERS on it.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“SNOWFLAKE.CORE.SEMANTIC_CATEGORY: Tag used to identify the native or custom category of the data in a column.”
↩︎ What classification produces: categories and system tags“It can be IDENTIFIER, QUASI_IDENTIFIER, or SENSITIVE”
↩︎ What classification produces: categories and system tags“If you are using SQL, associating the classification profile with a database to start the classification process is a separate step.”
↩︎ Automatic classification with classification profiles“all the tables and views in that database are being automatically classified according to the criteria defined in the profile.”
↩︎ Automatic classification with classification profiles“When AI mode is enabled, Snowflake uses the openai-gpt-5-mini model to augment classification results and improve accuracy.”
↩︎ Automatic classification with classification profiles“You can map user-defined tags to system-defined classification tags.”
↩︎ Wiring classification into governance policies“By default, views are excluded from classification.”
↩︎ Exam trap 2“every column that is identified as containing sensitive data is assigned two categories: a semantic category and a privacy category.”
↩︎ Checkpoint“the data will be automatically masked when Snowflake applies the tag as part of the classification process.”
↩︎ Checkpoint - 2.
“SNOWFLAKE.CLASSIFICATION_ADMIN: database role that enables you to create a custom classifier instance.”
↩︎ Custom classification with CUSTOM_CLASSIFIER“Create a custom classifier instance. Add the custom semantic category and regular expressions to the instance. Classify the table.”
↩︎ Checkpoint - 3.
“Eighty percent of the data in the sample must match the regular expressions that you add to the instance.”
↩︎ Custom classification with CUSTOM_CLASSIFIER“If you change the custom classifier, you must use the SET_CUSTOM_CLASSIFIERS method to update the classification profile with the new definition.”
↩︎ Custom classification with CUSTOM_CLASSIFIER“Sensitive data classification stores the definition of a custom classifier, not a reference.”
↩︎ Exam trap 1 - 4.
“Specifies a classification profile in order to classify based on the criteria specified in the profile”
↩︎ Manual, on-demand classification with SYSTEM$CLASSIFY“Any number from 1 to 10000, inclusive.”
↩︎ Manual, on-demand classification with SYSTEM$CLASSIFY“System tags are not set on any columns in the specified object.”
↩︎ Exam trap 3“System tags are not set on any columns in the specified object.”
↩︎ Prediction“When you use this argument, call the stored procedure with the role that has the OWNERSHIP privilege on the schema.”
↩︎ Checkpoint - 5.https://docs.snowflake.com/en/sql-reference/classes/classification_profile/commands/create-classification-profileOfficial docs
“Order the column_tag_map arrays from highest preference to lowest preference.”
↩︎ Wiring classification into governance policies“Executing a CREATE OR REPLACE command removes the classification profile from all databases and schemas, which turns off automatic classification.”
↩︎ Checkpoint