What you will be able to do
- Explain what a Shapley value measures and why background data defines the baseline it is measured against
- Log a Model Registry version with explainability enabled, and add explainability to a model logged without it
- Call the explain function from Python and SQL and choose the right visualization for local or global feature impact
- Read feature importance scores from Snowflake ML Functions and know their limits
- Use a SHAP dependence plot to see how a feature's value shapes predictions, and read prediction-interval bounds from a forecast
Key concept
Shapley value (SHAP) — A per-feature number that says how far one feature's value pushed one specific prediction above or below the model's average prediction. The 'average' comes from background data, so every SHAP value is a deviation from that baseline.
1.Why interpret a model, and what a Shapley value says
A trained model learns relationships from data instead of having them written down in advance. That is what makes machine learning useful, and it also means the result can behave like a black box. When a model underperforms, it is hard to tell why. A black box can also hide implicit biases. Regulated industries such as finance and healthcare may need evidence that a model reaches the right answers for the right reasons. To deal with this, the Snowflake Model Registry has an explainability function built on Shapley values. For each feature, a Shapley value is that feature's marginal contribution to one prediction, averaged over all possible combinations of features. The method is computationally expensive, and Snowflake presents it as the basis for interpreting and debugging models.
| Feature | Value | Contribution vs. an average house |
|---|---|---|
| Size | 2000 | +$50,000 |
| Location | Beachside | +$75,000 |
| Bedrooms | 3 | +$50,000 |
| Pets | No | -$25,000 |
There are two things to notice in the table. First, contributions have a sign. Not allowing pets makes the house less desirable, so that feature pulls the price down by $25,000. Second, the reference point is not zero. It is the average prediction, and that average is calculated from background data, meaning a representative sample of the whole dataset. Background data is used throughout the rest of this lesson. It decides whether explainability can be switched on at all, and whether SHAP values from two runs can be compared.
Checkpoint 1 of 9· Check yourself
In Snowflake's explainability, what does the 'average' that each Shapley contribution is measured against come from?
The baseline comes from background data. It is not set per call or fixed at zero, which is why changing the background data changes the SHAP values.
“The average value is calculated using background data, a representative sample of the entire dataset.”Source: docs.snowflake.com
Sources1
2.Enabling explainability when you log a model
Explainability is a property of a model version in the registry, so it is set up when you log the model. The preview supports Python-native XGBoost, CatBoost, LightGBM and scikit-learn models, plus the XGBoost, LightGBM and scikit-learn classes from snowflake.ml.modeling. It is on by default for these models when they are logged with Snowpark ML 1.6.2 or later, and the implementation uses the SHAP library. You supply background data by passing up to 1,000 rows in sample_input_data. Some tree-based models encode background data in their structure during training and may not need it supplied explicitly. Most models do need it, and every model, tree-based ones included, gets more accurate explanations when you provide it.
mv = reg.log_model(
catboost_model,
model_name="diamond_catboost_explain_enabled",
version_name="explain_v0",
conda_dependencies=["snowflake-ml-python"],
sample_input_data = xs, # xs will be used as background data
)| Situation | What to do |
|---|---|
| Supported model logged with Snowpark ML 1.6.2 or later | Explainability is on by default; pass background data in sample_input_data for the best explanations |
| Passing both signatures and sample_input_data | Also set options={"enable_explainability": True} |
| You do not want explanations for this version | Set enable_explainability to False |
| Model type requires background data, none supplied | Explainability cannot be enabled |
| Version logged before Snowpark ML 1.6.2 | Load it and log a new version with sample_input_data |
| XGBoost UnicodeDecodeError on snowflake-ml-python before 1.7.0, and you cannot upgrade | Downgrade XGBoost to 2.0.3 and log with options={"relax_version": False} |
The Snowpark ML 1.6.2 row matters most on the exam. Model versions are immutable, so upgrading the library does not add an explain method to a version that already exists. You load the model's Python object with ModelVersion.load and log it again as a new version, passing the background data. The environment you load into must match the deployment environment exactly, including the Python version and every library version.
Checkpoint 2 of 9· Fill the gap
This code adds explainability to a model logged with an older Snowpark ML. Which method retrieves the model's Python object?
mv_old = reg.get_model("model_without_explain_enabled").default
model = mv_old. ? ()
mv_new = reg.log_model(
model,
model_name="model_with_explain_enabled",
version_name="explain_v0",
conda_dependencies=["snowflake-ml-python"],
sample_input_data = xs
)ModelVersion.load returns the model object, which you then log as a new, explainable version. The old version cannot be modified because model versions are immutable.
Source: docs.snowflake.comCheckpoint 3 of 9· Exam question
A data scientist logs an XGBoost churn classifier to the Snowflake Model Registry and inspects SHAP values for one customer. The base value is 0.18 (log-odds) and the per-feature SHAP values sum to +1.02. What does this tell the data scientist about that customer's prediction?
Correct answer: B — The model's raw output for this customer is about 1.20 log-odds, because the SHAP contributions add to the base value to reproduce the prediction.
- A. Incorrect: SHAP values sum to the difference between the prediction and the base value, which is non-zero for any row that differs from the average.
- B. Correct: SHAP is additive, so the base value plus all feature contributions equals the model output for that row (here in log-odds, before a sigmoid).
- C. Incorrect: the base value is added, not subtracted, and the sum is a log-odds shift relative to the base value rather than a probability gap.
- D. Incorrect: for a tree classifier the contributions are typically on the margin (log-odds) scale and the base value must be added; 1.02 cannot be a probability.
Sources1
3.Calling explain from Python and SQL
An explainable model version has a method named explain that returns a Shapley value for each feature. A Shapley value explains one prediction made from specific inputs, so explain does nothing on its own. You pass it the rows whose predictions you want explained. In Python you call it with ModelVersion.run and the function name.
reg = Registry(...)
mv = reg.get_model("diamond_catboost_explain_enabled").default
explanations = mv.run(input_data, function_name="explain")In SQL, the same method is exposed as a table function on a model alias. Because the explanation runs in SQL, it can be applied to every row of a table. That is useful when you need per-row reasons for a whole batch of predictions.
WITH MV_ALIAS AS MODEL DATABASE.SCHEMA.DIAMOND_CATBOOST_EXPLAIN_ENABLED VERSION EXPLAIN_V0
SELECT *
FROM DATABASE.SCHEMA.DIAMOND_DATA,
TABLE(MV_ALIAS!EXPLAIN(CUT, COLOR, CLARITY, CARAT, DEPTH, TABLE_PCT, X, Y, Z));To compute a Shapley value, the algorithm perturbs input features and replaces them with background data. The output is therefore a deviation from that background. If you compare SHAP values from two datasets, for example this month's scored rows against last month's, both must have been explained against the same background data. Otherwise the baselines differ and the numbers are not comparable. One practical issue also appears in the docs: with snowflake-ml-python older than 1.7.0, XGBoost models can fail with a UnicodeDecodeError. This is caused by an incompatibility between SHAP 0.42.1 and XGBoost 2.1.1. The fix is to upgrade. If you cannot, pin XGBoost to 2.0.3 and log with relax_version set to False.
Checkpoint 4 of 9· Check yourself
A team explains January's and June's scored rows and wants to compare the SHAP values for one feature across the two months. What has to be true for the comparison to be meaningful?
SHAP values are deviations from the background data, so comparisons are only valid when the background data is the same. Row counts and the calling interface do not change the baseline.
“it is important to use consistent background data when comparing Shapley values from multiple data sets.”Source: docs.snowflake.com
Checkpoint 5 of 9· Exam question
A team logs a scikit-learn model with `reg.log_model(...)` and wants accurate Shapley-based explanations later. Which logging step gives the Model Registry the data it needs to compute these explanations most accurately?
Correct answer: D — Pass a representative sample of the training features, up to 1,000 rows, as `sample_input_data` so it can serve as background data.
- A. Incorrect: explainability is built in for supported models in recent snowflake-ml-python versions; dependencies do not supply background data.
- B. Incorrect: the explainer needs representative feature values to build baselines; labels alone give it nothing to perturb.
- C. Incorrect: the background sample is capped at 1,000 rows, so the full table cannot be supplied and is not needed.
- D. Correct: Shapley values are computed relative to background data, and the registry accepts up to 1,000 sample rows via `sample_input_data`.
Sources1
4.Feature impact: one prediction, the whole model, and ML Functions scores
The output of explain is a table of SHAP values. Snowflake provides three visualization functions in snowflake.ml.monitoring.explain_visualize for interpreting it, and they answer different questions.
A force plot covers one row. Each feature appears as an arrow that pushes the prediction up from the base value, shown in red, or down, shown in blue. Arrow size shows magnitude. The base_value argument defaults to 0.0, but the docs say it should normally be set to the model's mean prediction. contribution_threshold, which defaults to 0.05, hides features whose share of total absolute SHAP is too small. If no feature meets the threshold, or the threshold is outside 0 to 1, the function raises a SnowflakeMLException.
A violin plot covers many rows. It shows the distribution of SHAP values for each feature and sorts features by their absolute mean SHAP value, which gives a global ranking of feature impact.
| Function | Scope | Key inputs | Answers |
|---|---|---|---|
| plot_force | One instance | shap_row, features_row, base_value | Which features pushed this prediction up or down, and by how much |
| plot_violin | Many rows, all features | shap_df, feature_df | Which features matter most overall, and how spread out their effects are |
| plot_influence_sensitivity | Many rows, one feature | shap_values, feature_values | How the influence changes as the feature's value changes |
Snowflake ML Functions report feature impact separately, through SQL methods on the model object. A classification model has SHOW_FEATURE_IMPORTANCE, and forecasting and anomaly-detection models have EXPLAIN_FEATURE_IMPORTANCE. Each returns a rank and a score between 0 and 1 for every feature. For classification, the score counts how often the model's trees used a feature to make a decision, normalized so the scores sum to 1. You cannot choose the technique. Features that are very similar to each other can split importance between them, so each looks less important than it is.
Both methods also return a feature type column, but it is not part of FORECAST output. In EXPLAIN_FEATURE_IMPORTANCE, FEATURE_TYPE separates user_provided features from those derived_from_timestamp or derived_from_endogenous, and aggregated_endogenous_features groups every feature built from the target variable. In SHOW_FEATURE_IMPORTANCE for classification, feature_type is currently always user_provided.
Checkpoint 6 of 9· Match them up
Match each tool to the interpretation question it answers
Tap a term, then the definition that fits it.
Force plots explain a single instance. Violin plots rank features by absolute mean SHAP across rows. Influence sensitivity plots one feature against its SHAP values. SHOW_FEATURE_IMPORTANCE reports split-count importance for ML Functions classification.
“The violin plots are sorted by the absolute mean SHAP value of each feature”Source: docs.snowflake.com
Checkpoint 7 of 9· Exam question
An analyst runs the following query against a model version alias `mv` that was logged with explainability enabled: ```sql SELECT * FROM src_table, TABLE(mv!EXPLAIN(age, tenure, plan_type)); ``` What does the result contain for each input row?
Correct answer: A — Per-feature SHAP value columns for the supplied features, describing each one's contribution to that row's prediction.
- A. Correct: the `EXPLAIN` method function returns Shapley contributions per feature for each row passed in.
- B. Incorrect: `EXPLAIN` returns attribution values, not prediction intervals.
- C. Incorrect: the explain method works at row level; any global ranking must be aggregated by the analyst from these per-row values.
- D. Incorrect: partial dependence is a separate technique that varies a feature over a grid; the explain function does not produce it.
5.Partial dependence: how a feature's value shapes predictions
The violin plot shows how much a feature matters. It does not show which values of the feature push predictions up and which push them down. plot_influence_sensitivity fills that gap. It draws a SHAP dependence scatter, with the feature's values on the x-axis and the matching SHAP values on the y-axis, one point per row. From the shape of the scatter you can read the trend, the strength and direction of the influence, and any clusters that point to interactions with other features. There is also an environment difference. In Snowflake Notebooks you can pass 2D arrays covering many features and pick one from an interactive dropdown. In a local notebook you must pass one feature's SHAP values and feature values.
Checkpoint 8 of 9· Check yourself
A data scientist calls plot_influence_sensitivity in a local Jupyter notebook and passes 2D arrays of SHAP values and feature values for all features. What happens?
The 2D-array input with a dropdown selector works only in Snowflake Notebooks. In a local notebook you pass one feature's SHAP values and feature values.
“The feature of providing a 2D array of SHAP values and feature values is only available in Snowflake Notebooks.”Source: docs.snowflake.com
Sources6
6.Confidence intervals: quantifying uncertainty in predictions
Snowflake's FORECAST method returns uncertainty by default. Every forecast row has a FORECAST value plus LOWER_BOUND and UPPER_BOUND columns, which mark the edges of a prediction interval. The width is set by the prediction_interval key in the optional CONFIG_OBJECT. It must be at least 0.0 and less than 1.0, and it defaults to 0.95. A wider setting, such as 0.99, produces wider bounds because the interval has to cover more future points. A wide interval tells whoever uses the forecast that the point value is less certain.
Checkpoint 9 of 9· Check yourself
A forecast is run with the default configuration. How should LOWER_BOUND and UPPER_BOUND be read?
The default prediction_interval is 0.95. It means 95% of future points are expected to fall inside the returned bounds.
“The default value of 0.95 means 95% of future points are expected to fall within the interval [lower_bound, upper_bound] from the forecast result.”Source: docs.snowflake.com
Sources7
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Upgrading snowflake-ml-python will add the explain method to a model version logged with an older release.Why is that wrong?
Model versions are immutable. You load the old model and log it as a new version, passing background data in sample_input_data.
Covered in Enabling explainability when you log a model
2.Background data is optional for every model, so explainability will always work without it.Why is that wrong?
Some model types need explicit background data, and without it explainability cannot be enabled. Even tree-based models get more accurate explanations when it is supplied.
Covered in Enabling explainability when you log a model
3.Feature importance scores from SHOW_FEATURE_IMPORTANCE are exact measures of each feature's effect.Why is that wrong?
The scores are normalized split counts and give an approximate ranking. Very similar features can split importance between them, and the values should be treated as estimates.
Covered in Feature impact: one prediction, the whole model, and ML Functions scores
4.A force plot's default base value is the model's average prediction.Why is that wrong?
base_value defaults to 0.0. You should normally set it to the model's mean prediction yourself.
Covered in Feature impact: one prediction, the whole model, and ML Functions scores
5.SHOW_FEATURE_IMPORTANCE scores are SHAP values, so they explain why one specific prediction was made.Why is that wrong?
They count how often the model's trees used each feature, normalized to sum to 1. That is a model-level approximate ranking, not a per-prediction contribution measured against background data.
Covered in Feature impact: one prediction, the whole model, and ML Functions scores
6.Setting prediction_interval to 1.0 gives the widest possible forecast bounds.Why is that wrong?
The value must be at least 0.0 and strictly less than 1.0, so 1.0 is not allowed. Values closer to 1.0, such as 0.99, give wider bounds.
Covered in Confidence intervals: quantifying uncertainty in predictions
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.snowflake.com/en/developer-guide/snowflake-ml/model-registry/model-explainabilityOfficial docs
“might require stronger evidence that the model is producing the correct results for the right reasons.”
↩︎ Why interpret a model, and what a Shapley value says“The average value is calculated using background data, a representative sample of the entire dataset.”
↩︎ Why interpret a model, and what a Shapley value says“Explainability is available by default for the above models logged using Snowpark ML 1.6.2 and later.”
↩︎ Enabling explainability when you log a model“You can provide up to 1,000 rows of background data when logging a model by passing it in the sample_input_data parameter”
↩︎ Enabling explainability when you log a model“you will need to set this flag in order to pass both signatures and background data”
↩︎ Enabling explainability when you log a model“downgrade the XGBoost version to 2.0.3 and log the model with the relax_version option set to False”
↩︎ Enabling explainability when you log a model“you must pass input data to explain to generate the predictions to be explained.”
↩︎ Calling explain from Python and SQL“The Shapley value is computed by systematically perturbing input features and replacing them with the background data.”
↩︎ Calling explain from Python and SQL“If you cannot upgrade snowflake-ml-python to version 1.7.0 or later, downgrade the XGBoost version to 2.0.3”
↩︎ Calling explain from Python and SQL“Shapley values measure the average marginal contribution of each feature to the model’s prediction.”
↩︎ Key concept“Since model versions are immutable, you must create a new model version to add explainability to an existing model.”
↩︎ Exam trap 1“If the model is a type that requires explicit background data to calculate Shapley values, explainability cannot be enabled without this data.”
↩︎ Exam trap 2“Together, these contributions explain why this particular house is priced $150,000 higher than an average home.”
↩︎ Prediction“it is important to use consistent background data when comparing Shapley values from multiple data sets.”
↩︎ Checkpoint - 2.
“A feature’s contribution is represented by an arrow that directs the model’s prediction higher or lower from the base value.”
↩︎ Feature impact: one prediction, the whole model, and ML Functions scores“This defaults to 0.0 but should typically be set to the model’s mean prediction value.”
↩︎ Exam trap 4 - 3.
“The SHOW_FEATURE_IMPORTANCE method counts the number of times the model’s trees used each feature to make a decision.”
↩︎ Feature impact: one prediction, the whole model, and ML Functions scores“You cannot choose the technique used to calculate feature importance.”
↩︎ Feature impact: one prediction, the whole model, and ML Functions scores“Using multiple features that are very similar to each other may result in reduced importance scores for those features.”
↩︎ Exam trap 3“The SHOW_FEATURE_IMPORTANCE method counts the number of times the model’s trees used each feature to make a decision.”
↩︎ Exam trap 5 - 4.https://docs.snowflake.com/en/sql-reference/classes/forecast/methods/explain_feature_importanceOfficial docs
“aggregated_endogenous_features represents all features derived as transformations of the target variable.”
↩︎ Feature impact: one prediction, the whole model, and ML Functions scores“The source of the feature.”
↩︎ Feature impact: one prediction, the whole model, and ML Functions scores - 5.https://docs.snowflake.com/en/sql-reference/classes/classification/methods/show_feature_importanceOfficial docs
“a value in [0, 1], with 0 being the lowest possible importance, and 1 the highest.”
↩︎ Feature impact: one prediction, the whole model, and ML Functions scores“Currently this is always user_provided, which denotes feature data provided by the user.”
↩︎ Feature impact: one prediction, the whole model, and ML Functions scores - 6.
“create a SHAP dependence scatter plot to visualize the relationship between feature values and their SHAP values”
↩︎ Partial dependence: how a feature's value shapes predictions“The function returns a chart that visualizes the feature values along the x-axis and their corresponding SHAP values along the y-axis.”
↩︎ Partial dependence: how a feature's value shapes predictions“If you are using a local notebook, you must pass a single feature’s SHAP values and feature values as arguments.”
↩︎ Partial dependence: how a feature's value shapes predictions“The feature of providing a 2D array of SHAP values and feature values is only available in Snowflake Notebooks.”
↩︎ Checkpoint - 7.
“A value greater than or equal to 0.0 and less than 1.0.”
↩︎ Confidence intervals: quantifying uncertainty in predictions“Lower boundary of prediction interval.”
↩︎ Confidence intervals: quantifying uncertainty in predictions“A value greater than or equal to 0.0 and less than 1.0.”
↩︎ Exam trap 6“The default value of 0.95 means 95% of future points are expected to fall within the interval [lower_bound, upper_bound] from the forecast result.”
↩︎ Checkpoint
Also cited
“The violin plots are sorted by the absolute mean SHAP value of each feature”
↩︎ Checkpoint