What you will be able to do
- Tell a supervised workload from an unsupervised one by checking whether the training data carries a label for each row
- Explain how Snowflake's Classification function trains on labelled data, predicts on new rows and evaluates itself on held-out rows
- Tell apart model parameters that are learned during training from hyperparameters that a data scientist sets
- Recognise that SNOWFLAKE.ML.ANOMALY_DETECTION can be trained either without labels or with them, and say which applies in a given case
- Say what the Snowflake documentation covers about reinforcement learning and what it does not
Key concept
Label (target) column — A label column holds the known answer for every training row, such as the class a row belongs to. Supervised learning needs one. Unsupervised learning has none and has to find structure in the data itself.
1.Does each training row already have the answer?
Before you choose an algorithm or a Snowflake feature, ask one question about your training data: does each row already include the answer you want the model to produce? The answer decides which of the two paradigms in the Snowflake documentation you are working in.
In supervised learning, each training row pairs some inputs (features) with a known outcome. Snowflake's Classification function makes this explicit. Its training data must include "a target column representing the labeled class of each data point" and at least one feature column. The model learns how the features relate to that label. It then applies what it learned to rows whose label is unknown. The documentation gives churn prediction, credit card fraud detection and spam detection as typical uses. The Auto ML skill in Snowflake describes the same kind of task as "predict a category".
In unsupervised learning, the data has no answer column. The model has to find structure without help, such as values outside the expected range or rows that fall naturally into groups. The Auto ML skill describes clustering as a way to "segment data into groups", for example customer segmentation. Snowflake's anomaly detection function can also be trained with no labels at all.
Reinforcement learning is the third paradigm in the exam guide. The documentation behind this lesson does not cover it. The last section explains what that means for your preparation.
Checkpoint 1 of 8· Check yourself
A team has a table of past loan applications. Each row has applicant features and a REPAID column set to TRUE or FALSE. They want to score new applications. Which paradigm fits, and why?
The paradigm depends on whether the training rows carry a label, not on whether the rows you score later have one. REPAID is the labelled class for each training row, so the task is supervised.
“a target column representing the labeled class of each data point”Source: docs.snowflake.com
2.Supervised learning in practice: training on labelled rows
Classification "uses machine learning algorithms to sort data into different classes using patterns detected in training data." It supports binary classification (two classes) and multi-class classification (more than two). Classification is one of two common supervised task types. The Auto ML skill lists the other as regression: "predict a number", such as revenue, price or demand. The difference lies in the target: classification predicts a category, regression predicts a number. In SQL, you train a classifier by creating a model object and passing it a reference to the training data. You name the label column with the TARGET_COLNAME argument. The training data also needs at least one feature column. Fill in the argument that tells the model which column is the label:
Checkpoint 2 of 8· Fill the gap
This statement trains a binary classifier on a view with the columns user_interest_score, user_rating and label. Which argument names the label column?
CREATE OR REPLACE SNOWFLAKE.ML.CLASSIFICATION model_binary(
INPUT_DATA => SYSTEM$REFERENCE('view', 'binary_classification_view'),
? => 'label'
);Classification takes its label through TARGET_COLNAME. LABEL_COLNAME is an argument of anomaly detection, where it marks known anomalies, and TIMESTAMP_COLNAME belongs to the time-series functions.
Source: docs.snowflake.comCreating the object also fits the model: "The model is fitted to the provided training data." After that you call the model's PREDICT method on new rows. This step, applying a trained model to unlabelled data, is called inference. Inference only works if the new rows look like the training rows. "Inference data must have the same feature names and types as training data." Within that rule, the function tolerates some variation. A categorical feature can contain a value the model never saw during training. Columns that were not in the training data are ignored.
The label column has rules of its own. It needs at least two classes, because a model cannot learn to tell classes apart if every row has the same one. It also cannot have a distinct value on every row, and it may hold at most 255 distinct classes. Numeric features are treated as continuous. String and Boolean features are treated as categorical. To make the model treat a numeric code as a category, cast it to a string.
3.Labels make evaluation possible, and parameters are learned
Because a supervised model learns from known answers, you can measure how often it gets them right. The useful test is whether it gets them right on rows it has not seen. Snowflake's Classification function does this automatically. During evaluation, "an additional model is trained on the original data but with some data points withheld." It then runs inference on the withheld rows and compares its predicted classes with the actual ones. You read the results through the model's evaluation methods:
CALL <model_name>!SHOW_EVALUATION_METRICS();
CALL <model_name>!SHOW_GLOBAL_EVALUATION_METRICS();
CALL <model_name>!SHOW_THRESHOLD_METRICS();
CALL <model_name>!SHOW_CONFUSION_MATRIX();Checkpoint 3 of 8· Check yourself
How does the Classification function estimate how well its model predicts?
Rows the model trained on would make its accuracy look better than it is. The function holds rows back and scores predictions only on those.
“The withheld data points are then used for inference, and the predicted classes are compared to the actual classes.”Source: docs.snowflake.com
Checkpoint 4 of 8· Exam question
A telecom provider keeps three years of customer records in Snowflake. A CHURNED column was filled in 90 days after each account closed or renewed. The team wants a model that scores today's active customers by cancellation risk. Which learning paradigm fits MOST directly?
Correct answer: C — Supervised classification, since every historical row carries a known churn label the model can learn to predict from account features.
- A. Clustering finds similar groups but never learns a mapping to a known outcome, so its output is a segment id and not a calibrated churn probability.
- B. Reinforcement learning needs an agent that acts in an environment and learns from feedback to its own choices. A static table of past outcomes has no such interaction loop.
- C. A labeled binary outcome plus input features is the textbook supervised classification setup, and the fitted model can then score current accounts.
- D. Semi-supervised learning targets cases where labels are scarce. Here every historical row is labeled, so no pseudo-labeling step is needed.
A model has two kinds of settings. Parameters are the values the model learns from the data during training. Hyperparameters are values a person chooses before training starts. The ML Functions hide both kinds from you. For Classification, "You cannot choose or modify the classification algorithm", and "Model parameters cannot be manually specified or adjusted." If you write your own training code on Snowflake's Container Runtime, you control them. The PyTorch loop below is from the Snowflake training guide. The learning rate (lr=1e-3) and EPOCHS = 5 are set before the loop runs. Inside the loop, criterion(logits, yb) compares the model's output with the labels yb, and opt.step() updates the learned weights. This is supervised learning in its most basic form. Snowflake ML can also "Accelerate hyperparameter tuning" with distributed HPO (hyperparameter optimisation).
for epoch in range(1, EPOCHS + 1):
model.train()
for xb, yb in train_loader:
xb, yb = xb.to(DEVICE), yb.to(DEVICE)
logits = model(xb)
loss = criterion(logits, yb)
opt.zero_grad()
loss.backward()
opt.step()
acc = evaluate(val_loader)
print(f"epoch {epoch} val_acc={acc:.3f}")Checkpoint 5 of 8· Exam question
A retailer has no labeled segments for 2 million loyalty-card customers. Marketing wants natural groups based on purchase frequency, basket size and recency, with all three features on very different numeric scales. Which approach is MOST appropriate?
Correct answer: A — Standardize the three features, then fit `snowflake.ml.modeling.cluster.KMeans` and compare several values of k by inertia.
- A. KMeans is distance based, so features on different scales must be standardized first or the largest-scale column dominates. It needs no labels and k can be compared.
- B. Logistic regression is supervised and would only learn to separate the invented spend label. It cannot reveal a set of natural groups that were never labeled.
- C. Predicting an existing tier reproduces a business rule instead of discovering structure. The result would simply echo the labels that were supplied as ground truth.
- D. Skipping scaling lets the widest-range feature dominate Euclidean distance, so the groups reflect one column and not the combined purchasing behavior.
4.Unsupervised learning: finding outliers and groups without labels
"Anomaly detection is the process of identifying outliers in data." Outliers are data points that fall outside the expected range. The function works on time series and needs two columns: a timestamp, and a target column holding "some quantity of interest at each timestamp". The word *target* is easy to misread here. In Classification, the target is the label to predict. In anomaly detection, it is the metric being watched, such as daily sales. What makes anomaly detection supervised or unsupervised is a separate argument, LABEL_COLNAME. In the documentation's basic example, the label argument is an empty string: "Because this example uses unsupervised training, you do not need to use the label column." If you do know which past rows were anomalies, a Boolean label column tells the model so, and training becomes supervised.
| Setup | What the label or target argument holds | Paradigm |
|---|---|---|
| SNOWFLAKE.ML.CLASSIFICATION with TARGET_COLNAME | The labeled class of each training row | Supervised |
| SNOWFLAKE.ML.ANOMALY_DETECTION with LABEL_COLNAME => '' | No labels. TARGET_COLNAME is just the metric being watched | Unsupervised |
| SNOWFLAKE.ML.ANOMALY_DETECTION with a Boolean label column | Marks which historical rows are known anomalies | Supervised (labeled) |
Clustering is the other common unsupervised workload. The Snowflake Auto ML skill lists it as a task type for segmenting data into groups, for example customer segmentation or grouping anomalies. Unlike classification, clustering is given no classes in advance: it discovers the groups itself, so there is no label column to name.
Because no label column exists, the result cannot be scored against correct answers the way a classifier's withheld rows are. Snowflake's Feature Store guide fits a KMeans model and lists techniques for judging clusters: Within-Cluster Sum of Squares, Silhouette Score and Gap Statistics. The guide itself only plots the clusters and reviews them visually. The same guide scales the input variables with min-max scaling before fitting, so that columns on large numeric ranges do not dominate the grouping.
Clustering is not the only unsupervised technique. Snowflake's documentation on GPU-accelerated libraries also describes dimensionality reduction, which shrinks high-dimensional data (such as text embeddings) to fewer dimensions. It uses that step together with clustering to extract topics from product reviews, and it lists UMAP and HDBSCAN among the libraries. None of this needs a label column either.
Checkpoint 6 of 8· Exam question
A data scientist must predict the number of minutes each food delivery will take, using a labeled Snowflake table with distance, courier load and weather columns. Which `snowflake.ml.modeling` estimator choice is MOST appropriate for this target?
Correct answer: D — `xgboost.XGBRegressor` fitted on the features with the duration column as label, since the target is a continuous number.
- A. A classifier treats durations as unrelated categories, discards their ordering and cannot predict values missing from training, so it suits a continuous target poorly.
- B. KMeans is unsupervised and ignores the duration label during fitting. Its clusters are not optimized to predict the target, so it is not a proper regression approach.
- C. PCA is an unsupervised projection that maximizes variance, not agreement with the label. Its component is a rescaled mix of features and has no meaning as minutes.
- D. A continuous numeric label calls for supervised regression, and XGBRegressor learns that mapping and returns a predicted number of minutes for each delivery.
Checkpoint 7 of 8· Match them up
Match each argument setting with what it tells the model
Tap a term, then the definition that fits it.
Both functions have a 'target', but only Classification's target is a label. In anomaly detection, LABEL_COLNAME controls supervision: an empty string means unsupervised, and a Boolean column supplies known anomalies.
“The purpose of the label column is to tell the model which rows are known anomalies.”Source: docs.snowflake.com
5.Reinforcement learning: not covered by this documentation
The documentation does show one contrast. Every model in this lesson learns from a fixed dataset you hand it when you create the model. Once trained, it does not keep learning as it is used. "Models are immutable and cannot be updated in place." To include new outcomes, you drop the model and train a new one, often with CREATE OR REPLACE. So if a question describes a model that keeps changing its behaviour based on feedback from new results, that is not what the ML Functions shown here do.
Checkpoint 8 of 8· Check yourself
A trained SNOWFLAKE.ML.CLASSIFICATION model starts to fall behind newly observed outcomes. According to the documentation, how do you bring it up to date?
ML Functions models cannot be changed after training, and their parameters cannot be set by hand. The only way to update one is to train a replacement.
“To update a model, drop the existing model and train a new one.”Source: docs.snowflake.com
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Anomaly detection is always unsupervised, because outliers are never labelled in advance.Why is that wrong?
SNOWFLAKE.ML.ANOMALY_DETECTION can train without labels (LABEL_COLNAME => '') or with a Boolean column that marks known anomalies. With that column, training is supervised.
Covered in Unsupervised learning: finding outliers and groups without labels
2.Any function with a TARGET_COLNAME argument is doing supervised learning.Why is that wrong?
In anomaly detection, the target column is just the metric being tracked over time. Whether training is supervised depends on the separate label column.
Covered in Unsupervised learning: finding outliers and groups without labels
3.Classification evaluation metrics are measured on the rows the model was trained on.Why is that wrong?
Evaluation trains an additional model with some rows withheld, then compares its predictions on those withheld rows with their actual classes.
Covered in Labels make evaluation possible, and parameters are learned
4.Any column can be the label, even one with a single value or a unique value on every row.Why is that wrong?
The label column must have more than one distinct value and fewer distinct values than there are rows. It may also hold at most 255 classes.
Covered in Supervised learning in practice: training on labelled rows
5.Predicting a number, such as price or demand, is classification because it is supervised.Why is that wrong?
Both classification and regression are supervised, but classification predicts a category and regression predicts a number.
Covered in Supervised learning in practice: training on labelled rows
6.Clustering is just classification under another name, so it needs a label column.Why is that wrong?
Classification sorts rows into classes using a labeled target column. Clustering segments data into groups without labels, so its groups are judged by measures such as silhouette score instead of correct answers.
Covered in Unsupervised learning: finding outliers and groups without labels
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“a target column representing the labeled class of each data point”
↩︎ Does each training row already have the answer?“Classification uses machine learning algorithms to sort data into different classes using patterns detected in training data.”
↩︎ Supervised learning in practice: training on labelled rows“Inference data must have the same feature names and types as training data.”
↩︎ Supervised learning in practice: training on labelled rows“Suitable training datasets for use with classification include a target column representing the labeled class of each data point and at least one feature column.”
↩︎ Supervised learning in practice: training on labelled rows“an additional model is trained on the original data but with some data points withheld”
↩︎ Labels make evaluation possible, and parameters are learned“Model parameters cannot be manually specified or adjusted.”
↩︎ Labels make evaluation possible, and parameters are learned“Models are immutable and cannot be updated in place.”
↩︎ Reinforcement learning: not covered by this documentation“a target column representing the labeled class of each data point”
↩︎ Key concept“an additional model is trained on the original data but with some data points withheld”
↩︎ Exam trap 3“The cardinality of the label (target) column must be greater than one and less than the number of rows in the dataset.”
↩︎ Exam trap 4“The withheld data points are then used for inference, and the predicted classes are compared to the actual classes.”
↩︎ Checkpoint“To update a model, drop the existing model and train a new one.”
↩︎ Checkpoint - 2.
“Clustering: segment data into groups (customer segmentation, anomaly grouping)”
↩︎ Does each training row already have the answer?“Classification: predict a category (churn yes/no, fraud detection, lead scoring)”
↩︎ Does each training row already have the answer?“Regression: predict a number (revenue forecast, price estimation, demand prediction)”
↩︎ Supervised learning in practice: training on labelled rows“Clustering: segment data into groups (customer segmentation, anomaly grouping)”
↩︎ Unsupervised learning: finding outliers and groups without labels“Regression: predict a number (revenue forecast, price estimation, demand prediction)”
↩︎ Exam trap 5“Clustering: segment data into groups (customer segmentation, anomaly grouping)”
↩︎ Exam trap 6 - 3.
“Accelerate hyperparameter tuning with Snowflake ML’s distributed HPO, optimized for data stored in Snowflake.”
↩︎ Labels make evaluation possible, and parameters are learned - 4.
“Anomaly detection is the process of identifying outliers in data.”
↩︎ Unsupervised learning: finding outliers and groups without labels“Because this example uses unsupervised training, you do not need to use the label column.”
↩︎ Unsupervised learning: finding outliers and groups without labels“Optionally, individual historical rows can be labeled as anomalous or non-anomalous by using a separate Boolean column.”
↩︎ Exam trap 1“A target column representing some quantity of interest at each timestamp.”
↩︎ Exam trap 2“Optionally, individual historical rows can be labeled as anomalous or non-anomalous by using a separate Boolean column.”
↩︎ Prediction“The purpose of the label column is to tell the model which rows are known anomalies.”
↩︎ Checkpoint - 5.https://www.snowflake.com/en/developers/guides/advanced-guide-to-snowflake-feature-storeSecondary source
“Within-Cluster Sum of Squares”
↩︎ Unsupervised learning: finding outliers and groups without labels“to scale our input variables using min-max scaling”
↩︎ Unsupervised learning: finding outliers and groups without labels - 6.
“Applying dimensionality reduction at scale”
↩︎ Unsupervised learning: finding outliers and groups without labels