CertSafari
    Snowflake SnowPro Advanced: Data Scientist (DSA-C03)· Lessons

    Domain 1 · Lesson 2/16

    Machine Learning Problem Types: Regression, Classification, Forecasting, Clustering

    Identify machine learning problem types.

    16 min read
    4.25% of exam
    9 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Work out the ML problem type (regression, binary or multi-class classification, time-series forecasting, clustering) from the shape of the data and the target
    • Match each structured-data problem type to the Snowflake tool that handles it: scikit-learn in Snowpark, SNOWFLAKE.ML.CLASSIFICATION or SNOWFLAKE.ML.FORECAST
    • Find the configuration mistakes the exam tests: single series vs multi-series forecasts, numeric codes treated as continuous, and future feature values
    • Tell structured-data problems apart from unstructured image problems, and see where the provided documentation does not cover a problem type

    Key concept

    The target decides the problem type — Before choosing a tool, ask what the model has to output. A labelled class means classification. A continuous number means regression. A numeric value at future timestamps means forecasting. No label at all, only groups to discover, means clustering.

    1.Framing the problem: what does the model have to output?

    Exam questions in this subdomain usually describe a business situation and ask which kind of model fits. The fastest way to answer is to look at the target, the column the model has to produce. Supervised problems learn from rows where the answer is already known. Snowflake's classification documentation states this directly: a suitable training set must include a target column with the labelled class of each data point, plus at least one feature column. Change the target and the problem type changes with it, even when the features stay the same.

    Snowflake gives you two ways to build these models. The ML Functions are packaged SQL tools: Snowflake picks the model type for each function, so you don't need to be an ML expert. They split into time-series functions (Forecasting, Anomaly Detection) and functions that work without time-series data (Classification, Top Insights). For problem types that have no ML Function, such as plain linear regression, you train open-source models (scikit-learn, XGBoost, LightGBM, PyTorch) yourself inside Snowflake.

    Reading the problem type from the target, and the Snowflake tool the sources show for each
    Problem typeWhat the target looks likeTool shown in the sources
    Linear regressionA continuous number, e.g. REVENUEscikit-learn LinearRegression in a Snowpark stored procedure
    Binary classificationExactly two classes, e.g. a TRUE/FALSE labelSNOWFLAKE.ML.CLASSIFICATION
    Multi-class classificationMore than two classes, e.g. not_interested / add_to_wishlist / purchaseSNOWFLAKE.ML.CLASSIFICATION
    Time-series forecastingA numeric metric over a timestamp column, predicted forwardSNOWFLAKE.ML.FORECAST
    ClusteringNo label; groups are discoveredSnowpark-ML Kmeans (guide example)

    Checkpoint 1 of 9· Check yourself

    A team wants to predict whether each customer will cancel their subscription next month. Their history has one row per customer with a TRUE/FALSE column for whether they cancelled. Which problem type is this?

    Sources12

    2.Linear regression: predicting a continuous number

    When the target is a quantity, such as revenue, cost or units, you have a regression problem. Snowflake has no dedicated ML Function for plain regression, so the documentation trains one with scikit-learn. In the Snowpark Python example, a stored procedure running on a Snowpark-optimized warehouse trains a linear regression model on a table called MARKETING_BUDGETS_FEATURES. The target is REVENUE. The features are spend on SEARCH_ENGINE, SOCIAL_MEDIA, VIDEO and EMAIL.

    Notice how the code separates the target from the features and keeps a test split back:

    Separating the continuous REVENUE target from its features and holding back 20% as a test setpython
    # Load features
     df = session.table('MARKETING_BUDGETS_FEATURES').to_pandas()
     X = df.drop('REVENUE', axis = 1)
     y = df['REVENUE']
    
     # Split dataset into training and test
     X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state = 42)

    The procedure returns an R² score for both the training data and the test data, which is a typical way to judge a regression model. Linear regression is only one choice: the Snowflake ML training docs point out that XGBoost and LightGBM can also build classification, regression and ranking models. Either way, a numeric target is what makes this regression.

    Checkpoint 2 of 9· Fill the gap

    The procedure predicts REVENUE, a continuous value. Which estimator completes the pipeline?

    # Create pipeline and train
     pipeline = Pipeline(steps=[('preprocessor', preprocessor),('classifier',  ? (n_jobs=-1))])
     model = GridSearchCV(pipeline, param_grid={}, n_jobs=-1, cv=10)
     model.fit(X_train, y_train)

    Sources34

    3.Binary and multi-class classification

    Classification sorts rows into classes it learned from labelled examples. The difference between binary and multi-class is just how many classes the target has. Binary means two classes (purchase / no purchase). Multi-class means more than two (not_interested / add_to_wishlist / purchase). SNOWFLAKE.ML.CLASSIFICATION handles both. It runs on a gradient boosting machine, and it trains binary and multi-class problems with different loss functions: an area-under-the-curve loss for binary and a logistic loss for multi-class. You cannot choose the algorithm or set its parameters.

    The documentation's example table has both kinds of target side by side: a Boolean label column for the binary case and a string class column for the multi-class case. Training the binary model only requires naming the target:

    Training a binary classifier; the target column is the only thing you namesql
    CREATE OR REPLACE SNOWFLAKE.ML.CLASSIFICATION model_binary(
        INPUT_DATA => SYSTEM$REFERENCE('view', 'binary_classification_view'),
        TARGET_COLNAME => 'label'
    );

    PREDICT returns an object containing the probability of each class and the predicted class, which is the one with the highest probability. The function also enforces limits that define what counts as a classification problem. The label must have more than one distinct value and fewer distinct values than there are rows, and the target can have at most 255 distinct classes. How a feature is treated depends on its data type, and this is where problems get set up wrong:

    How SNOWFLAKE.ML.CLASSIFICATION treats each feature data type
    Data typeTreated asConsequence
    NumericContinuousCast numeric codes to strings if they should be treated as categories
    StringCategoricalHigh-cardinality values are fine; full free text such as sentences or paragraphs is not supported
    BooleanCategoricalCan also be used as a binary label
    TIMESTAMP_NTZSource of derived featuresEpoch, day, week and month features are created and appear as derived_features

    Checkpoint 3 of 9· Match them up

    Match each classification detail to what the Snowflake documentation says about it

    Tap a term, then the definition that fits it.

    Sources1

    4.Time-series forecasting: a numeric target at future timestamps

    Forecasting is still supervised, because the model learns from past values of the target. What sets it apart is that time is the organising variable and the output is a value at timestamps that haven't happened yet. The minimum input is a table or view with one timestamp column and one numeric column. The docs ask for timestamps at a fixed interval without too many gaps. They also note that training can cope with real-world data that has missing, duplicate or misaligned time steps. The model works out the interval itself from the training data.

    The minimal forecasting workflow: train on one timestamp and one target column, then forecast 7 periods aheadsql
    -- Train your model
    CREATE SNOWFLAKE.ML.FORECAST my_model(
      INPUT_DATA => TABLE(my_view),
      TIMESTAMP_COLNAME => 'my_timestamps',
      TARGET_COLNAME => 'my_metric'
    );
    
    -- Generate forecasts using your model
    SELECT * FROM TABLE(my_model!FORECAST(FORECASTING_PERIODS => 7));

    For a multi-series forecast, the docs build one column that combines store_id and item into store_item and pass it as the series column. A single CREATE statement then trains a model for every series. To get forecasts for just one series, pass SERIES_VALUE to the FORECAST method. That is cheaper than generating every series and filtering the results.

    Forecasts can also use features such as temperature, humidity or holidays. Any column you don't name as the timestamp or target is assumed to be a feature. There is a catch: once the model is trained on features, you have to supply *future* values of those features when you forecast. The timestamps of the forecast then come from that future-features input, not from FORECASTING_PERIODS. Every forecast comes with LOWER_BOUND and UPPER_BOUND, a prediction interval whose width you can set with prediction_interval. SHOW_EVALUATION_METRICS reports accuracy measured on out-of-sample data.

    Checkpoint 4 of 9· Fill the gap

    This model must produce a separate sales forecast for each store/item combination. Which parameter completes it?

    CREATE SNOWFLAKE.ML.FORECAST model2(
      INPUT_DATA => TABLE(v3),
       ?  => 'store_item',
      TIMESTAMP_COLNAME => 'date',
      TARGET_COLNAME => 'sales'
    );

    Checkpoint 5 of 9· Exam question

    A logistics analyst at a courier company wants to estimate the number of minutes each parcel will take to reach its recipient, using distance, parcel weight, weather score and hour of day stored in a Snowflake table with historical delivery times. Which machine learning problem type does this represent?

    Checkpoint 6 of 9· Exam question

    A support organisation has 40,000 historical tickets in a Snowflake table, each tagged with exactly one of six queues: Billing, Access, Outage, Feature Request, Security and Other. The team wants a model that routes new tickets automatically. Which problem type fits best?

    Sources56

    5.Unstructured data: image classification and segmentation

    Everything so far has used structured, tabular data. The exam guide also lists supervised problems on unstructured data: image classification, where one label is assigned to a whole image, and segmentation. The structured-data tools don't carry over. SNOWFLAKE.ML.CLASSIFICATION accepts numeric, Boolean, string and TIMESTAMP_NTZ inputs, and it does not support even full free text. A picture of a product cannot go through it.

    The sources show two other routes. The first is the Cortex AI Functions. Their multimodal capabilities cover documents, images, audio and video, and the documented use cases include classifying images. The second is training your own deep learning model on Snowflake ML's GPU container runtime with PyTorch or TensorFlow, which you would do when you need a custom vision model.

    Checkpoint 7 of 9· Check yourself

    A retailer has product photos in a stage and wants each photo labelled with a product category. Which statement fits the sources?

    Sources741

    6.Unsupervised clustering, and the gap on association models

    Take away the target column and the question is no longer "predict this label" but "what natural groups exist in this data?" That is clustering, which the exam guide puts under unsupervised learning. None of the ML Functions covers it: they are Forecasting, Anomaly Detection, Classification and Top Insights. You train clustering models in Snowflake ML with open-source frameworks, and the agentic ML documentation lists clustering next to classification, regression and forecasting as model types you can train.

    The Snowflake Feature Store guide walks through a concrete example. A Snowpark-ML Kmeans model is fitted on a training dataset. Before that, the input variables are scaled with min-max scaling, and the scaling is applied at model time because it captures global state (the minimum and maximum of each column) from the training sample. There is no label to compare predictions against, so clustering is evaluated differently from classification. The guide lists Within-Cluster Sum of Squares, Silhouette Score, Gap Statistics and Cross-Validation as validation techniques, and in its notebook it simply plots the clusters and inspects them visually.

    Checkpoint 8 of 9· Check yourself

    After fitting a Kmeans model to unlabelled customer data, which of these is listed as a way to validate the clusters?

    Checkpoint 9 of 9· Exam question

    A bank stores two years of card transactions in Snowflake, each labelled fraudulent or legitimate, with only 0.3% fraud. A data scientist plans to use the Snowflake ML Classification function to flag new transactions. How should the problem be framed?

    The exam guide also lists association models under GenAI. None of the documentation provided for this lesson describes association models: there is no association-rule or market-basket feature, and no definition of the term. The ML Functions overview names only Forecasting, Anomaly Detection, Classification and Top Insights. Top Insights finds dimensions and values that affect a metric in surprising ways, but the documentation does not call it an association model, so don't treat the two as the same thing. For this bullet, study from the current Snowflake documentation rather than this lesson.

    Sources892

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.If the table contains several stores, CREATE SNOWFLAKE.ML.FORECAST will forecast each store separately automatically.Why is that wrong?

      Separate series have to be declared with SERIES_COLNAME, often a combined column such as store_item. Without it, you don't get one forecast per store.

      Covered in Time-series forecasting: a numeric target at future timestamps

    2. 2.A forecast model trained with extra columns such as temperature can still be called with only FORECASTING_PERIODS.Why is that wrong?

      Extra columns become features, and a model trained on features needs their future values passed in as INPUT_DATA at forecast time.

      Covered in Time-series forecasting: a numeric target at future timestamps

    3. 3.Numeric category codes (for example region 1, 2, 3) are treated as categories automatically by SNOWFLAKE.ML.CLASSIFICATION.Why is that wrong?

      Numeric features are always treated as continuous. To have them treated as categories, cast them to strings.

      Covered in Binary and multi-class classification

    4. 4.Unstructured inputs such as free-text reviews can go straight into the classification ML Function as string features.Why is that wrong?

      String features are treated as categorical values. Full free text is not supported, and images need multimodal or deep learning tooling.

      Covered in Unstructured data: image classification and segmentation

    Practise it for real

    Train single-series and multi-series forecasting models on the documentation's sample sales data and compare the output

    1. 1.Create the sales_data table from the forecasting docs, then create view v1 with SELECT date, sales FROM sales_data WHERE store_id=1 AND item='jacket'.

      Why: A single series needs only one timestamp column and one numeric target.

      You should see: v1 returns daily rows from 2020-01-01, with sales rising by 1 each day.

    2. 2.Run CREATE SNOWFLAKE.ML.FORECAST model1 with INPUT_DATA => TABLE(v1), TIMESTAMP_COLNAME => 'date', TARGET_COLNAME => 'sales', then CALL model1!FORECAST(FORECASTING_PERIODS => 3).

      Why: This shows forecasting as supervised learning on past values of the target.

      You should see: Three rows for 2020-01-13 to 2020-01-15 with SERIES NULL and forecasts 14, 15, 16.

    3. 3.Create view v3 that combines [store_id, item] AS store_item, train model2 with SERIES_COLNAME => 'store_item', and CALL model2!FORECAST(FORECASTING_PERIODS => 2).

      Why: Declaring the series column is what gives you one forecast per store/item.

      You should see: Two forecast rows each for [ 1, "jacket" ] and [ 2, "umbrella" ].

    Stuck? Get a nudge

    If every SERIES value comes back NULL, check whether you trained without SERIES_COLNAME.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “include a target column representing the labeled class of each data point”
      ↩︎ Framing the problem: what does the model have to output?
      “Common use cases of classification include customer churn prediction, credit card fraud detection, and spam detection.”
      ↩︎ Framing the problem: what does the model have to output?
      “Binary classification (two classes) and multi-class classification (more than two classes) are supported.”
      ↩︎ Binary and multi-class classification
      “For binary classification, the model is trained using an area-under-the-curve loss function.”
      ↩︎ Binary and multi-class classification
      “Your target column must contain no more than 255 distinct classes.”
      ↩︎ Binary and multi-class classification
      “The cardinality of the label (target) column must be greater than one and less than the number of rows in the dataset.”
      ↩︎ Binary and multi-class classification
      “The prediction object includes predicted probabilities for each class and the predicted class based on the maximum predicted probability.”
      ↩︎ Binary and multi-class classification
      “It does not support full free text, like sentences or paragraphs.”
      ↩︎ Unstructured data: image classification and segmentation
      “include a target column representing the labeled class of each data point”
      ↩︎ Key concept
      “Numeric features are treated as continuous. To treat numeric features as categorical, cast them to strings.”
      ↩︎ Exam trap 3
      “It does not support full free text, like sentences or paragraphs.”
      ↩︎ Exam trap 4
      “For multi-class classification, the model is trained using a logistic loss function.”
      ↩︎ Checkpoint
    2. 2.
      “Snowflake provides an appropriate type of model for each feature, so you don't have to be a machine learning expert to take advantage of them.”
      ↩︎ Framing the problem: what does the model have to output?
      “Top Insights helps you find dimensions and values that affect the metric in surprising ways.”
      ↩︎ Unsupervised clustering, and the gap on association models
    3. 3.
      “The example then creates a stored procedure that trains a linear regression model.”
      ↩︎ Linear regression: predicting a continuous number
      “Return model R2 score on train and test data”
      ↩︎ Linear regression: predicting a continuous number
    4. 4.
      “you can use the XGBoost and LightGBM libraries to develop powerful classification, regression, and ranking models.”
      ↩︎ Linear regression: predicting a continuous number
      “You can use a GPU-powered container runtime image to train deep learning models with PyTorch, TensorFlow, and other frameworks.”
      ↩︎ Unstructured data: image classification and segmentation
    5. 5.
      “Forecasting uses a machine learning algorithm that predicts future numeric data from historical time series data.”
      ↩︎ Time-series forecasting: a numeric target at future timestamps
      “However, model training can handle real-world data that has missing, duplicate, or misaligned time steps.”
      ↩︎ Time-series forecasting: a numeric target at future timestamps
      “Note that the model has inferred the interval between timestamps from the training data.”
      ↩︎ Time-series forecasting: a numeric target at future timestamps
      “additional columns in the input data are assumed to be features for use in training.”
      ↩︎ Time-series forecasting: a numeric target at future timestamps
      “To create a forecasting model for multiple series at once, use the series_colname parameter.”
      ↩︎ Exam trap 1
      “To generate forecasts with this model, you must provide future values for the features to the model”
      ↩︎ Exam trap 2
      “To create a forecasting model for multiple series at once, use the series_colname parameter.”
      ↩︎ Prediction
    6. 7.
      “Content understanding: Summarize, classify, and describe documents, images, audio, and video.”
      ↩︎ Unstructured data: image classification and segmentation
    7. 8.
      “Use CoCo to train classification, regression, forecasting, and clustering models using scikit-learn, XGBoost, LightGBM, PyTorch, or AutoGluon.”
      ↩︎ Unsupervised clustering, and the gap on association models
    8. 9.
      “We use the training Dataset we created in the previous step to fit a Snowpark-ML Kmeans model.”
      ↩︎ Unsupervised clustering, and the gap on association models
      “includes some feature pre-processing to scale our input variables using min-max scaling”
      ↩︎ Unsupervised clustering, and the gap on association models
      “In the Notebook we have simply plotted the clusters to review visually.”
      ↩︎ Unsupervised clustering, and the gap on association models
      “Silhouette Score”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 16 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.