What you will be able to do
- Explain how DMFs, expectations and the DMF schedule combine into a data quality check
- Attach DMFs with expectations at creation time and run them on every DML change
- Write custom and multi-table DMFs for feature validation beyond null and duplicate counts
- Monitor freshness and volume with the FRESHNESS DMF and anomaly detection, and know what these sources do not cover
1.The parts of a data quality check
Snowflake data quality checks keep validating data on a schedule and report violations so you can act on them. A check has three parts, and the exam expects you to keep them apart.
A data metric function (DMF) measures something about the data, such as a NULL count or how recently the table changed. It returns a value but does not decide whether that value is good or bad. Snowflake provides system DMFs in SNOWFLAKE.CORE, and you can write custom ones. An expectation turns a DMF into a pass/fail check: the returned value is compared with the expectation, and failures are reported as expectation violations. The DMF schedule controls how often the DMFs run, and it defaults to once an hour.
DMFs can be set on regular, temporary and transient tables, dynamic tables, event tables, external tables, Iceberg tables, views and materialized views. They cannot be set on hybrid tables or streams, on shared tables or views, or on objects in a reader account. Each account can have up to 50,000 DMF associations.
Checkpoint 1 of 7· Match them up
Match each term to its role
Tap a term, then the definition that fits it.
A DMF only measures. The expectation turns the measurement into a pass/fail check, the schedule sets how often it runs, and anomaly detection judges values against learned history.
“An expectation is combined with a DMF to create a data quality check.”Source: docs.snowflake.com
Cost follows the schedule. Scheduled DMFs run on serverless compute billed under *Data Quality Monitoring*, and logging the results to the event table is billed separately as *Logging*. Creating a DMF costs nothing, and calling one ad hoc in a SELECT is not billed. To track spend, use DATA_QUALITY_MONITORING_USAGE_HISTORY.
Checkpoint 2 of 7· Check yourself
A team wants to validate rows the moment they land. Which target cannot have a DMF set on it?
Streams and hybrid tables are the documented exclusions. Dynamic, Iceberg and transient tables are all supported.
“You cannot set a DMF on a hybrid table or a stream object.”Source: docs.snowflake.com
Checkpoint 3 of 7· Exam question
Which privilege setup lets a user view data lineage for objects in Snowsight?
Correct answer: D — VIEW LINEAGE on the account, object-level privileges such as SELECT or REFERENCES on the objects, and USAGE on the containing database and schema.
- A. Incorrect: lineage is captured by the platform automatically and is not computed by user-owned tasks or a warehouse.
- B. Incorrect: that grant exposes shared usage views but does not confer VIEW LINEAGE, so the Lineage graph still withholds nodes.
- C. Incorrect: MONITOR does not control lineage, and ownership of every node is not required; read-level privileges on the objects suffice.
- D. Correct: lineage is gated by the account-level VIEW LINEAGE privilege and the user also needs access to the objects shown, including USAGE on their database and schema.
Sources1
2.Attaching and scheduling DMFs at ingestion
To associate a DMF with an existing table or view, use ALTER and name the columns passed as arguments. Some DMFs, such as ROW_COUNT, take no column, so you write ON (). To break results down by a dimension, add a WITHIN GROUP clause to the association.
ALTER TABLE t
ADD DATA METRIC FUNCTION SNOWFLAKE.CORE.NULL_COUNT
ON (c1);For an ingestion table, monitoring should start on day one. The WITH DATA METRIC FUNCTION clause attaches DMFs, and their expectations, inside the CREATE statement itself. For tables it comes after the column definitions. For views, materialized views and dynamic tables it must come before AS SELECT. All bindings attach atomically, so a single invalid binding fails the whole CREATE. CLONE and LIKE copy the bindings, and CREATE OR REPLACE rebuilds them from scratch. In an expectation, the left side must be the keyword VALUE.
CREATE OR REPLACE TABLE customers (
customer_id NUMBER,
email VARCHAR
)
WITH DATA METRIC FUNCTION (
SNOWFLAKE.CORE.NULL_COUNT
ON (email)
EXPECTATION no_null_email ( VALUE = 0 )
);The DATA_METRIC_SCHEDULE parameter sets when DMFs run. It can be a number of minutes, a cron expression, or TRIGGER_ON_CHANGES, which runs the DMFs whenever DML changes the table. That last option is the natural fit for validating data as it arrives, though reclustering does not trigger a run and only some table types support it. The schedule is set per object, not per DMF. To pause one DMF, MODIFY it with SUSPEND. To pause all of them, set the schedule to an empty string. Changes to the schedule of existing DMFs take up to 10 minutes to apply. When you evaluate results in DATA_QUALITY_MONITORING_RESULTS, use measurement_time, because rows can be inserted between the scheduled time and the actual evaluation.
Checkpoint 4 of 7· Fill the gap
Complete the statement so the DMFs on this table run whenever a DML operation such as an INSERT changes it.
ALTER TABLE hr.tables.empl_info SET DATA_METRIC_SCHEDULE = ' ? ';TRIGGER_ON_CHANGES ties the run to DML changes. The other real options are time-based, and ON_INSERT is not a valid value.
Source: docs.snowflake.comSources2
3.Feature validation beyond basic checks
Null and duplicate counts catch broken loads, but feature validation often needs domain rules. The first step up is ACCEPTED_VALUES. It takes a column and a lambda expression and counts the records that do not satisfy the lambda.
ALTER TABLE t1
ADD DATA METRIC FUNCTION SNOWFLAKE.CORE.ACCEPTED_VALUES ON (age, age -> age = 5);When no system DMF expresses the rule, write one with CREATE DATA METRIC FUNCTION. A custom DMF takes a TABLE argument and returns a number. A custom DMF can also take more than one table argument, which lets it check consistency across datasets: referential integrity, matching, or conditional comparisons. When you attach it, the table you attach it to becomes the first argument, and any other table must be given by its fully qualified name. The documented example, governance.dmfs.referential_check, counts rows whose key has no match in a reference table.
ALTER TABLE salesorders ADD DATA METRIC FUNCTION governance.dmfs.referential_check ON (sp_id, TABLE (my_db.sch1.salespeople(sp_id)));Custom DMFs work like other functions. You can call one manually to test it before attaching it, make it secure with ALTER FUNCTION ... SET SECURE, and tag it. You cannot drop one while it is still associated with a table or view. Use DATA_METRIC_FUNCTION_REFERENCES to find those associations first.
Checkpoint 5 of 7· Check yourself
After the referential_check DMF above runs, it returns 12. What does that mean?
salesorders is the table the DMF is attached to, so it is the first argument. The function counts its rows whose sp_id has no match in the second table.
“A value greater than 0 indicates that there are sp_id values in salesorders”Source: docs.snowflake.com
Sources3
4.Freshness, volume anomalies and the limits of drift monitoring
The FRESHNESS system DMF returns the number of seconds since a table was last modified. Pass a timestamp column (DATE, TIMESTAMP_LTZ or TIMESTAMP_TZ) and it measures from the maximum value in that column, which is what you want when event time matters more than load time. Leave the column out and it measures from the last DML. A column is required when you attach it to a view or external table. In both forms, the scheduled run time is used for the comparison, not the time the run actually happened.
A fixed expectation such as "fresher than one hour" breaks down when update patterns vary. Anomaly detection handles this by training on historical DMF values and flagging any value above or below a predicted range. It is available for just two system DMFs: ROW_COUNT for volume and FRESHNESS for update frequency. For DMFs that run frequently, it needs at least two weeks of history so it can learn weekly seasonality. Anomaly checks also run on their own cadence, independent of the DMF schedule.
ALTER TABLE t1
ADD DATA METRIC FUNCTION SNOWFLAKE.CORE.ROW_COUNT ON ()
ANOMALY_DETECTION = TRUE;Checkpoint 6 of 7· Check yourself
Which DMF associations can use ANOMALY_DETECTION = TRUE?
Anomaly detection is documented only for the ROW_COUNT (volume) and FRESHNESS (update frequency) system DMFs.
“Currently, Snowflake can automatically detect anomalies in the volume and freshness of your data.”Source: docs.snowflake.com
Checkpoint 7 of 7· Exam question
A fraud team receives card-swipe events from an application as individual rows and needs them queryable in a feature table within seconds, without staging files in cloud storage. Which ingestion approach fits?
Correct answer: A — Snowpipe Streaming, writing rows with a client SDK straight into a PIPE-backed table so no intermediate files are staged.
- A. Correct: Snowpipe Streaming is built for row-oriented, low-latency ingestion directly from applications, with data queryable in seconds.
- B. Incorrect: a COPY task is batch and file-based, and the one-minute cadence cannot deliver the seconds-level freshness required.
- C. Incorrect: classic Snowpipe is file-based and adds staging and notification latency; per-event files would also be inefficient.
- D. Incorrect: per-event inserts keep a warehouse running and do not scale; Snowflake provides a purpose-built streaming path instead.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Each DMF on a table can run on its own schedule, for example NULL_COUNT every 5 minutes and ROW_COUNT daily.Why is that wrong?
DATA_METRIC_SCHEDULE is an object parameter, so every DMF on the table or view shares one schedule.
Covered in Attaching and scheduling DMFs at ingestion
2.Testing a DMF by calling it in a SELECT consumes Data Quality Monitoring credits.Why is that wrong?
Billing applies only when a scheduled DMF is computed on an object. Ad hoc calls are not billed.
Covered in The parts of a data quality check
3.Anomaly detection starts flagging problems as soon as it is enabled.Why is that wrong?
It first trains on historical DMF values, and frequently run DMFs need at least two weeks of data before it can detect anomalies.
Covered in Freshness, volume anomalies and the limits of drift monitoring
Practise it for real
Put quality checks on an ingestion table that run on every load and watch its volume for anomalies
1.Create the customers table with WITH DATA METRIC FUNCTION binding SNOWFLAKE.CORE.NULL_COUNT on email and EXPECTATION no_null_email ( VALUE = 0 ).
Why: Monitoring starts the moment the table exists, with no separate ALTER step.
You should see: The table is created with the DMF binding attached. An invalid binding would fail the whole CREATE.
2.Run ALTER TABLE customers SET DATA_METRIC_SCHEDULE = 'TRIGGER_ON_CHANGES';
Why: Ties validation to DML changes, so each load is checked, instead of waiting for the hourly default.
You should see: The statement succeeds, as long as the table type supports trigger-based scheduling.
3.Run SHOW PARAMETERS LIKE 'DATA_METRIC_SCHEDULE' IN TABLE customers;
Why: Confirms which schedule now applies to every DMF on the table.
You should see: The value column shows TRIGGER_ON_CHANGES at the TABLE level.
4.Run ALTER TABLE customers ADD DATA METRIC FUNCTION SNOWFLAKE.CORE.ROW_COUNT ON () ANOMALY_DETECTION = TRUE;
Why: Adds a learned volume check on top of the fixed null-count expectation.
You should see: The association is added. Anomalies are flagged only after enough history has been collected.
Stuck? Get a nudge
If results look inconsistent with recent inserts, filter DATA_QUALITY_MONITORING_RESULTS on measurement_time rather than the scheduled time.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“a DMF is a building block of a data quality check”
↩︎ The parts of a data quality check“By default, the DMF schedule runs a DMF once every hour.”
↩︎ The parts of a data quality check“You can only have 50,000 total associations of DMFs on objects per account.”
↩︎ The parts of a data quality check“You are not billed for unscheduled data metric function usage, such as calling a DMF with a SELECT statement.”
↩︎ Exam trap 2“An expectation is combined with a DMF to create a data quality check.”
↩︎ Checkpoint“You cannot set a DMF on a hybrid table or a stream object.”
↩︎ Checkpoint“Currently, Snowflake can automatically detect anomalies in the volume and freshness of your data.”
↩︎ Checkpoint - 2.
“Atomicity: All DMF bindings attach atomically.”
↩︎ Attaching and scheduling DMFs at ingestion“specify the measurement_time column in your query as the basis for the evaluation”
↩︎ Attaching and scheduling DMFs at ingestion“To break results down by a dimension (for example, null counts per region), include a WITHIN GROUP clause when you create the association.”
↩︎ Attaching and scheduling DMFs at ingestion“All data metric functions on a table or view follow the same schedule.”
↩︎ Exam trap 1 - 3.
“When you add the DMF to a table, that table is used as the first argument.”
↩︎ Feature validation beyond basic checks“You cannot drop a custom DMF from the system while it is still associated with a table or view.”
↩︎ Feature validation beyond basic checks“A value greater than 0 indicates that there are sp_id values in salesorders”
↩︎ Checkpoint - 4.
“Returns how much time in seconds has elapsed since a table was last modified.”
↩︎ Freshness, volume anomalies and the limits of drift monitoring“You must specify a column argument if you want to associate this function with a view or external table.”
↩︎ Freshness, volume anomalies and the limits of drift monitoring“comparing the current run of the function with the last time a DML command acted on the table.”
↩︎ Prediction - 5.
“Snowflake trains this algorithm with historical data, then automatically identifies return values that are above or below a predicted range.”
↩︎ Freshness, volume anomalies and the limits of drift monitoring“Snowflake requires at least two weeks of DMF data to start detecting anomalies.”
↩︎ Exam trap 3