What you will be able to do
- Work out the ML problem type (regression, binary or multi-class classification, time-series forecasting, clustering) from the shape of the data and the target
- Match each structured-data problem type to the Snowflake tool that handles it: scikit-learn in Snowpark, SNOWFLAKE.ML.CLASSIFICATION or SNOWFLAKE.ML.FORECAST
- Find the configuration mistakes the exam tests: single series vs multi-series forecasts, numeric codes treated as continuous, and future feature values
- Tell structured-data problems apart from unstructured image problems, and see where the provided documentation does not cover a problem type
Key concept
The target decides the problem type — Before choosing a tool, ask what the model has to output. A labelled class means classification. A continuous number means regression. A numeric value at future timestamps means forecasting. No label at all, only groups to discover, means clustering.
1.Framing the problem: what does the model have to output?
Exam questions in this subdomain usually describe a business situation and ask which kind of model fits. The fastest way to answer is to look at the target, the column the model has to produce. Supervised problems learn from rows where the answer is already known. Snowflake's classification documentation states this directly: a suitable training set must include a target column with the labelled class of each data point, plus at least one feature column. Change the target and the problem type changes with it, even when the features stay the same.
Snowflake gives you two ways to build these models. The ML Functions are packaged SQL tools: Snowflake picks the model type for each function, so you don't need to be an ML expert. They split into time-series functions (Forecasting, Anomaly Detection) and functions that work without time-series data (Classification, Top Insights). For problem types that have no ML Function, such as plain linear regression, you train open-source models (scikit-learn, XGBoost, LightGBM, PyTorch) yourself inside Snowflake.
| Problem type | What the target looks like | Tool shown in the sources |
|---|---|---|
| Linear regression | A continuous number, e.g. REVENUE | scikit-learn LinearRegression in a Snowpark stored procedure |
| Binary classification | Exactly two classes, e.g. a TRUE/FALSE label | SNOWFLAKE.ML.CLASSIFICATION |
| Multi-class classification | More than two classes, e.g. not_interested / add_to_wishlist / purchase | SNOWFLAKE.ML.CLASSIFICATION |
| Time-series forecasting | A numeric metric over a timestamp column, predicted forward | SNOWFLAKE.ML.FORECAST |
| Clustering | No label; groups are discovered | Snowpark-ML Kmeans (guide example) |
Checkpoint 1 of 9· Check yourself
A team wants to predict whether each customer will cancel their subscription next month. Their history has one row per customer with a TRUE/FALSE column for whether they cancelled. Which problem type is this?
The target is a labelled column with two values, which makes this binary classification. Snowflake lists customer churn prediction as a typical classification use case.
“Common use cases of classification include customer churn prediction, credit card fraud detection, and spam detection.”Source: docs.snowflake.com
2.Linear regression: predicting a continuous number
When the target is a quantity, such as revenue, cost or units, you have a regression problem. Snowflake has no dedicated ML Function for plain regression, so the documentation trains one with scikit-learn. In the Snowpark Python example, a stored procedure running on a Snowpark-optimized warehouse trains a linear regression model on a table called MARKETING_BUDGETS_FEATURES. The target is REVENUE. The features are spend on SEARCH_ENGINE, SOCIAL_MEDIA, VIDEO and EMAIL.
Notice how the code separates the target from the features and keeps a test split back:
# Load features
df = session.table('MARKETING_BUDGETS_FEATURES').to_pandas()
X = df.drop('REVENUE', axis = 1)
y = df['REVENUE']
# Split dataset into training and test
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state = 42)The procedure returns an R² score for both the training data and the test data, which is a typical way to judge a regression model. Linear regression is only one choice: the Snowflake ML training docs point out that XGBoost and LightGBM can also build classification, regression and ranking models. Either way, a numeric target is what makes this regression.
Checkpoint 2 of 9· Fill the gap
The procedure predicts REVENUE, a continuous value. Which estimator completes the pipeline?
# Create pipeline and train
pipeline = Pipeline(steps=[('preprocessor', preprocessor),('classifier', ? (n_jobs=-1))])
model = GridSearchCV(pipeline, param_grid={}, n_jobs=-1, cv=10)
model.fit(X_train, y_train)The doc says this stored procedure trains a linear regression model on REVENUE. StandardScaler and PolynomialFeatures appear earlier in the same code, but as preprocessing steps, not estimators. The pipeline step happens to be named 'classifier', but that name doesn't change the problem type.
Source: docs.snowflake.com3.Binary and multi-class classification
Classification sorts rows into classes it learned from labelled examples. The difference between binary and multi-class is just how many classes the target has. Binary means two classes (purchase / no purchase). Multi-class means more than two (not_interested / add_to_wishlist / purchase). SNOWFLAKE.ML.CLASSIFICATION handles both. It runs on a gradient boosting machine, and it trains binary and multi-class problems with different loss functions: an area-under-the-curve loss for binary and a logistic loss for multi-class. You cannot choose the algorithm or set its parameters.
The documentation's example table has both kinds of target side by side: a Boolean label column for the binary case and a string class column for the multi-class case. Training the binary model only requires naming the target:
CREATE OR REPLACE SNOWFLAKE.ML.CLASSIFICATION model_binary(
INPUT_DATA => SYSTEM$REFERENCE('view', 'binary_classification_view'),
TARGET_COLNAME => 'label'
);PREDICT returns an object containing the probability of each class and the predicted class, which is the one with the highest probability. The function also enforces limits that define what counts as a classification problem. The label must have more than one distinct value and fewer distinct values than there are rows, and the target can have at most 255 distinct classes. How a feature is treated depends on its data type, and this is where problems get set up wrong:
| Data type | Treated as | Consequence |
|---|---|---|
| Numeric | Continuous | Cast numeric codes to strings if they should be treated as categories |
| String | Categorical | High-cardinality values are fine; full free text such as sentences or paragraphs is not supported |
| Boolean | Categorical | Can also be used as a binary label |
| TIMESTAMP_NTZ | Source of derived features | Epoch, day, week and month features are created and appear as derived_features |
Checkpoint 3 of 9· Match them up
Match each classification detail to what the Snowflake documentation says about it
Tap a term, then the definition that fits it.
The function uses a different loss for each classification type, caps the number of classes, and treats numbers as continuous unless you cast them to strings.
“For multi-class classification, the model is trained using a logistic loss function.”Source: docs.snowflake.com
Sources1
4.Time-series forecasting: a numeric target at future timestamps
Forecasting is still supervised, because the model learns from past values of the target. What sets it apart is that time is the organising variable and the output is a value at timestamps that haven't happened yet. The minimum input is a table or view with one timestamp column and one numeric column. The docs ask for timestamps at a fixed interval without too many gaps. They also note that training can cope with real-world data that has missing, duplicate or misaligned time steps. The model works out the interval itself from the training data.
-- Train your model
CREATE SNOWFLAKE.ML.FORECAST my_model(
INPUT_DATA => TABLE(my_view),
TIMESTAMP_COLNAME => 'my_timestamps',
TARGET_COLNAME => 'my_metric'
);
-- Generate forecasts using your model
SELECT * FROM TABLE(my_model!FORECAST(FORECASTING_PERIODS => 7));For a multi-series forecast, the docs build one column that combines store_id and item into store_item and pass it as the series column. A single CREATE statement then trains a model for every series. To get forecasts for just one series, pass SERIES_VALUE to the FORECAST method. That is cheaper than generating every series and filtering the results.
Forecasts can also use features such as temperature, humidity or holidays. Any column you don't name as the timestamp or target is assumed to be a feature. There is a catch: once the model is trained on features, you have to supply *future* values of those features when you forecast. The timestamps of the forecast then come from that future-features input, not from FORECASTING_PERIODS. Every forecast comes with LOWER_BOUND and UPPER_BOUND, a prediction interval whose width you can set with prediction_interval. SHOW_EVALUATION_METRICS reports accuracy measured on out-of-sample data.
Checkpoint 4 of 9· Fill the gap
This model must produce a separate sales forecast for each store/item combination. Which parameter completes it?
CREATE SNOWFLAKE.ML.FORECAST model2(
INPUT_DATA => TABLE(v3),
? => 'store_item',
TIMESTAMP_COLNAME => 'date',
TARGET_COLNAME => 'sales'
);SERIES_COLNAME tells the model which column identifies each series when it trains. SERIES_VALUE is a FORECAST-method argument that picks one series at prediction time.
Source: docs.snowflake.comCheckpoint 5 of 9· Exam question
A logistics analyst at a courier company wants to estimate the number of minutes each parcel will take to reach its recipient, using distance, parcel weight, weather score and hour of day stored in a Snowflake table with historical delivery times. Which machine learning problem type does this represent?
Correct answer: A — Supervised regression, because the target is a continuous number learned from labelled historical deliveries
- A. Correct. A labelled, continuous target such as delivery minutes is the defining trait of supervised regression, which linear regression addresses.
- B. Incorrect. Binary classification predicts one of two categories; the stated goal is the number of minutes, which is continuous rather than a late/on-time flag.
- C. Incorrect. Clustering needs no label and only groups similar rows; here the historical delivery time is a known label that the model must predict.
- D. Incorrect. Bucketing would discard the numeric precision the analyst asked for; predicting minutes directly is a regression task, not a category assignment.
Checkpoint 6 of 9· Exam question
A support organisation has 40,000 historical tickets in a Snowflake table, each tagged with exactly one of six queues: Billing, Access, Outage, Feature Request, Security and Other. The team wants a model that routes new tickets automatically. Which problem type fits best?
Correct answer: C — Multi-class classification, because each ticket receives exactly one label out of six mutually exclusive classes
- A. Incorrect. A single yes/no question for one queue ignores the other five queues, so it cannot route a ticket to its correct destination.
- B. Incorrect. Existing queue tags provide supervision, and discovered clusters would not map reliably onto the six business queues the team defined.
- C. Correct. One label from more than two exclusive categories, learned from tagged examples, is the definition of multi-class classification.
- D. Incorrect. The target is a category, not a number; scoring queues is how classifiers work internally rather than a separate regression framing.
5.Unstructured data: image classification and segmentation
Everything so far has used structured, tabular data. The exam guide also lists supervised problems on unstructured data: image classification, where one label is assigned to a whole image, and segmentation. The structured-data tools don't carry over. SNOWFLAKE.ML.CLASSIFICATION accepts numeric, Boolean, string and TIMESTAMP_NTZ inputs, and it does not support even full free text. A picture of a product cannot go through it.
The sources show two other routes. The first is the Cortex AI Functions. Their multimodal capabilities cover documents, images, audio and video, and the documented use cases include classifying images. The second is training your own deep learning model on Snowflake ML's GPU container runtime with PyTorch or TensorFlow, which you would do when you need a custom vision model.
Checkpoint 7 of 9· Check yourself
A retailer has product photos in a stage and wants each photo labelled with a product category. Which statement fits the sources?
The input is an image, so the tabular classification function does not apply; a file path string carries no visual content. Cortex AI multimodal functions explicitly list classifying images as a use case.
“Content understanding: Summarize, classify, and describe documents, images, audio, and video.”Source: docs.snowflake.com
6.Unsupervised clustering, and the gap on association models
Take away the target column and the question is no longer "predict this label" but "what natural groups exist in this data?" That is clustering, which the exam guide puts under unsupervised learning. None of the ML Functions covers it: they are Forecasting, Anomaly Detection, Classification and Top Insights. You train clustering models in Snowflake ML with open-source frameworks, and the agentic ML documentation lists clustering next to classification, regression and forecasting as model types you can train.
The Snowflake Feature Store guide walks through a concrete example. A Snowpark-ML Kmeans model is fitted on a training dataset. Before that, the input variables are scaled with min-max scaling, and the scaling is applied at model time because it captures global state (the minimum and maximum of each column) from the training sample. There is no label to compare predictions against, so clustering is evaluated differently from classification. The guide lists Within-Cluster Sum of Squares, Silhouette Score, Gap Statistics and Cross-Validation as validation techniques, and in its notebook it simply plots the clusters and inspects them visually.
Checkpoint 8 of 9· Check yourself
After fitting a Kmeans model to unlabelled customer data, which of these is listed as a way to validate the clusters?
A confusion matrix and AUC loss need known labels, so they belong to classification. Prediction bounds belong to forecasting. The clustering guide lists Silhouette Score with Within-Cluster Sum of Squares and Gap Statistics.
“Silhouette Score”Source: www.snowflake.com
Checkpoint 9 of 9· Exam question
A bank stores two years of card transactions in Snowflake, each labelled fraudulent or legitimate, with only 0.3% fraud. A data scientist plans to use the Snowflake ML Classification function to flag new transactions. How should the problem be framed?
Correct answer: D — As binary classification on a heavily imbalanced target, judged by precision and recall for the fraud class
- A. Incorrect. Merchant category is a feature here; the target has only two values, fraud or legitimate, so the problem is binary rather than multi-class.
- B. Incorrect. The labelled outcome is categorical; a probability score is just the output of a classifier, not evidence that the target is a continuous variable.
- C. Incorrect. Labels exist for every transaction, so supervised learning applies; rarity calls for imbalance handling, not for abandoning the labels.
- D. Correct. Two labelled outcomes make this binary classification, and with 0.3% positives, class-level precision and recall reveal more than overall accuracy.
The exam guide also lists association models under GenAI. None of the documentation provided for this lesson describes association models: there is no association-rule or market-basket feature, and no definition of the term. The ML Functions overview names only Forecasting, Anomaly Detection, Classification and Top Insights. Top Insights finds dimensions and values that affect a metric in surprising ways, but the documentation does not call it an association model, so don't treat the two as the same thing. For this bullet, study from the current Snowflake documentation rather than this lesson.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If the table contains several stores, CREATE SNOWFLAKE.ML.FORECAST will forecast each store separately automatically.Why is that wrong?
Separate series have to be declared with SERIES_COLNAME, often a combined column such as store_item. Without it, you don't get one forecast per store.
Covered in Time-series forecasting: a numeric target at future timestamps
2.A forecast model trained with extra columns such as temperature can still be called with only FORECASTING_PERIODS.Why is that wrong?
Extra columns become features, and a model trained on features needs their future values passed in as INPUT_DATA at forecast time.
Covered in Time-series forecasting: a numeric target at future timestamps
3.Numeric category codes (for example region 1, 2, 3) are treated as categories automatically by SNOWFLAKE.ML.CLASSIFICATION.Why is that wrong?
Numeric features are always treated as continuous. To have them treated as categories, cast them to strings.
Covered in Binary and multi-class classification
4.Unstructured inputs such as free-text reviews can go straight into the classification ML Function as string features.Why is that wrong?
String features are treated as categorical values. Full free text is not supported, and images need multimodal or deep learning tooling.
Covered in Unstructured data: image classification and segmentation
Practise it for real
Train single-series and multi-series forecasting models on the documentation's sample sales data and compare the output
1.Create the sales_data table from the forecasting docs, then create view v1 with SELECT date, sales FROM sales_data WHERE store_id=1 AND item='jacket'.
Why: A single series needs only one timestamp column and one numeric target.
You should see: v1 returns daily rows from 2020-01-01, with sales rising by 1 each day.
2.Run CREATE SNOWFLAKE.ML.FORECAST model1 with INPUT_DATA => TABLE(v1), TIMESTAMP_COLNAME => 'date', TARGET_COLNAME => 'sales', then CALL model1!FORECAST(FORECASTING_PERIODS => 3).
Why: This shows forecasting as supervised learning on past values of the target.
You should see: Three rows for 2020-01-13 to 2020-01-15 with SERIES NULL and forecasts 14, 15, 16.
3.Create view v3 that combines [store_id, item] AS store_item, train model2 with SERIES_COLNAME => 'store_item', and CALL model2!FORECAST(FORECASTING_PERIODS => 2).
Why: Declaring the series column is what gives you one forecast per store/item.
You should see: Two forecast rows each for [ 1, "jacket" ] and [ 2, "umbrella" ].
Stuck? Get a nudge
If every SERIES value comes back NULL, check whether you trained without SERIES_COLNAME.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“include a target column representing the labeled class of each data point”
↩︎ Framing the problem: what does the model have to output?“Common use cases of classification include customer churn prediction, credit card fraud detection, and spam detection.”
↩︎ Framing the problem: what does the model have to output?“Binary classification (two classes) and multi-class classification (more than two classes) are supported.”
↩︎ Binary and multi-class classification“For binary classification, the model is trained using an area-under-the-curve loss function.”
↩︎ Binary and multi-class classification“Your target column must contain no more than 255 distinct classes.”
↩︎ Binary and multi-class classification“The cardinality of the label (target) column must be greater than one and less than the number of rows in the dataset.”
↩︎ Binary and multi-class classification“The prediction object includes predicted probabilities for each class and the predicted class based on the maximum predicted probability.”
↩︎ Binary and multi-class classification“It does not support full free text, like sentences or paragraphs.”
↩︎ Unstructured data: image classification and segmentation“include a target column representing the labeled class of each data point”
↩︎ Key concept“Numeric features are treated as continuous. To treat numeric features as categorical, cast them to strings.”
↩︎ Exam trap 3“It does not support full free text, like sentences or paragraphs.”
↩︎ Exam trap 4“For multi-class classification, the model is trained using a logistic loss function.”
↩︎ Checkpoint - 2.
“Snowflake provides an appropriate type of model for each feature, so you don't have to be a machine learning expert to take advantage of them.”
↩︎ Framing the problem: what does the model have to output?“Top Insights helps you find dimensions and values that affect the metric in surprising ways.”
↩︎ Unsupervised clustering, and the gap on association models - 3.https://docs.snowflake.com/en/developer-guide/snowpark/python/python-snowpark-training-mlOfficial docs
“The example then creates a stored procedure that trains a linear regression model.”
↩︎ Linear regression: predicting a continuous number“Return model R2 score on train and test data”
↩︎ Linear regression: predicting a continuous number - 4.
“you can use the XGBoost and LightGBM libraries to develop powerful classification, regression, and ranking models.”
↩︎ Linear regression: predicting a continuous number“You can use a GPU-powered container runtime image to train deep learning models with PyTorch, TensorFlow, and other frameworks.”
↩︎ Unstructured data: image classification and segmentation - 5.
“Forecasting uses a machine learning algorithm that predicts future numeric data from historical time series data.”
↩︎ Time-series forecasting: a numeric target at future timestamps“However, model training can handle real-world data that has missing, duplicate, or misaligned time steps.”
↩︎ Time-series forecasting: a numeric target at future timestamps“Note that the model has inferred the interval between timestamps from the training data.”
↩︎ Time-series forecasting: a numeric target at future timestamps“additional columns in the input data are assumed to be features for use in training.”
↩︎ Time-series forecasting: a numeric target at future timestamps“To create a forecasting model for multiple series at once, use the series_colname parameter.”
↩︎ Exam trap 1“To generate forecasts with this model, you must provide future values for the features to the model”
↩︎ Exam trap 2“To create a forecasting model for multiple series at once, use the series_colname parameter.”
↩︎ Prediction - 6.
“method provides evaluation metrics on out-of-sample data.”
↩︎ Time-series forecasting: a numeric target at future timestamps - 7.
“Content understanding: Summarize, classify, and describe documents, images, audio, and video.”
↩︎ Unstructured data: image classification and segmentation - 8.
“Use CoCo to train classification, regression, forecasting, and clustering models using scikit-learn, XGBoost, LightGBM, PyTorch, or AutoGluon.”
↩︎ Unsupervised clustering, and the gap on association models - 9.https://www.snowflake.com/en/developers/guides/advanced-guide-to-snowflake-feature-storeSecondary source
“We use the training Dataset we created in the previous step to fit a Snowpark-ML Kmeans model.”
↩︎ Unsupervised clustering, and the gap on association models“includes some feature pre-processing to scale our input variables using min-max scaling”
↩︎ Unsupervised clustering, and the gap on association models“In the Notebook we have simply plotted the clusters to review visually.”
↩︎ Unsupervised clustering, and the gap on association models“Silhouette Score”
↩︎ Checkpoint