CertSafari
    Snowflake SnowPro Advanced: Data Scientist (DSA-C03)· Lessons

    Domain 1 · Lesson 3/16

    ML Lifecycle in Snowflake: Data Collection, Exploration, Feature Engineering and Training

    Summarize the machine learning lifecycle.

    11 min read
    4.25% of exam
    6 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Name the stages of the machine learning lifecycle and the Snowflake ML capability that supports each one
    • Explain why an immutable dataset snapshot matters for reproducible training
    • Identify data quality issues that exploration must surface before modeling begins
    • Explain how a feature store prevents training-serving skew
    • Describe where models are trained in Snowflake and how experiments compare training runs

    Key concept

    End-to-end ML lifecycle — Building a model is a chain of stages: collect and prepare data, explore it, engineer features, train and compare models, deploy, monitor, and version. Each stage feeds the next, so a mistake made early, such as leakage or features computed inconsistently, carries through to production.

    1.The lifecycle as a chain of stages

    Every exam question about the ML lifecycle assumes you know the order of the stages and what each one produces. Snowflake's ML overview lists them as capabilities on one platform. You prepare data, create and use features with the Snowflake Feature Store, and train models in notebooks on Container Runtime. Experiments evaluate the trained models against set metrics. ML Jobs operationalize the pipelines, the Snowflake Model Registry deploys models for inference at scale, and ML Observability and Explainability monitor the models in production. Running underneath all of these is ML Lineage, which traces artifacts from source data to features, datasets and models, for reproducibility, compliance and debugging.

    The stages form a loop rather than a straight line. Monitoring finds drift, drift triggers retraining, and retraining produces a new model version. This page covers the first half of the loop, from raw data to a trained model.

    Lifecycle stage and the Snowflake ML capability that supports it
    Lifecycle stageSnowflake ML capability
    Feature engineeringSnowflake Feature Store
    TrainingSnowflake Notebooks on Container Runtime, Snowflake ML Jobs
    Comparing trained modelsExperiments
    DeploymentSnowflake Model Registry, Model Serving to Snowpark Container Services
    Monitoring and explanationML Observability, ML Explainability
    Tracing artifacts end to endML Lineage

    Checkpoint 1 of 7· Match them up

    Match each Snowflake ML capability to the lifecycle job it does

    Tap a term, then the definition that fits it.

    Sources1

    2.Data collection: getting data in and freezing it

    Collection means getting the data a model needs into a place where it can be governed and queried. Snowflake Notebooks let you explore data that is already in Snowflake, or upload new data from local files, from external cloud storage, or from Snowflake Marketplace datasets. When the data has to keep arriving, for example new transactions every hour, the Feature Store supports automated, incremental refresh from both batch and streaming sources. You define a feature pipeline once and it keeps updating as new data lands.

    Collection also has to be reproducible. If the training table changes after a model is trained, nobody can recreate the exact rows that model learned from. Snowflake Datasets solve this by giving you an immutable, file-based snapshot of the data, with a name and an optional version. A model trained on dataset version v1 can always be traced back to v1.

    These sources describe the Snowflake mechanics of collection. They do not cover general collection methodology such as sampling design or how labels are gathered, so this lesson does not teach those topics.

    Checkpoint 2 of 7· Check yourself

    A team wants to be able to retrain last quarter's model on exactly the rows it originally used, even though the source tables have changed since then. What should they have created at training time?

    Sources21

    3.Visualization and exploration: find problems before you model

    Exploratory data analysis (EDA) comes before any modeling. Its purpose is to find patterns and problems while they are still cheap to fix. Snowflake Notebooks give you a cell-based environment for Python, SQL and Markdown, where you can perform exploratory data analysis and develop models in the same interface. For charts, Altair and Streamlit are imported by default. You can also install matplotlib, plotly or seaborn from the notebook's Packages menu. If you work in SQL worksheets instead, Snowsight can turn query results into bar charts, line charts, scatterplots, heat grids and scorecards, which let you quickly spot patterns and outliers.

    Plotting the notebook's toy penguin dataset with plotlypython
    import plotly.express as px
    px.bar(df, x='measurement', y='value', color='species')

    Exploration is a gate, not just a set of charts. Snowflake's Auto ML skill in CoCo treats it that way. Its four-step workflow puts mandatory quality gates (leakage detection, baseline scoring and EDA) ahead of any modeling. Leakage is the finding that matters most. If a column carries information that would not be available at prediction time, the model will look excellent in testing and fail in production. A baseline score gives you a number that any real model has to beat.

    Checkpoint 3 of 7· Put it in order

    Put the Auto ML skill's four workflow steps in order

    1. 1.Build: feature engineering, model training and evaluation
    2. 2.Understand and configure: explore the data and propose task type, metric and trial plan
    3. 3.Quality gates: leakage detection, baseline scoring, EDA
    4. 4.Deliver: update the experiment manifest and present next steps

    Checkpoint 4 of 7· Exam question

    A binary churn classifier is evaluated on 1,000 customers and produces this confusion matrix: TP = 80, FP = 20, FN = 40, TN = 860. Which interpretation is correct?

    Sources23

    4.Feature engineering: define once, use in training and serving

    Feature engineering turns raw columns into the inputs a model learns from. The lifecycle risk here is training-serving skew: the features a model sees in production are computed differently from the features it was trained on. The Snowflake Feature Store is built to prevent this. It is an integrated place to define, manage, store and discover features. It supports both batch and low-latency online retrieval, and it keeps training and inference consistent because both read the same feature definitions.

    Scale is the other concern. When the data is too large for one node, Snowflake's distributed APIs spread feature engineering and training across multiple nodes. The preprocessing functions live in snowflake.ml.modeling.preprocessing. In a well-run pipeline, preprocessing is registered once and reused everywhere, never re-implemented by hand for production.

    Checkpoint 5 of 7· Check yourself

    Which step best removes training-serving skew caused by preprocessing that was re-implemented separately for production?

    Checkpoint 6 of 7· Exam question

    A fraud model scores card transactions where only 0.4% are fraudulent. A candidate model that labels every transaction as legitimate reports 99.6% accuracy on the test set. Which evaluation approach is MOST appropriate for judging whether the model is useful?

    Sources4

    5.Training models and comparing them with experiments

    You train a model in either a Snowflake Notebook or a Snowflake ML Job. Both run on Container Runtime, which comes with packages such as scikit-learn, numpy and scipy preinstalled. XGBoost and LightGBM are also available, and GPU images support PyTorch and TensorFlow. ML Jobs let teams that work in an external IDE send their code to the same runtime. To load training data, the DataConnector reads a Snowflake table into a pandas DataFrame, a PyTorch dataset or a TensorFlow dataset, and it uses the compute pool's distributed processing to load the data faster.

    Loading a table through DataConnector and training an XGBoost classifierpython
    from snowflake.ml.data.data_connector import DataConnector
    from snowflake.snowpark.context import get_active_session
    import xgboost as xgb
    
    session = get_active_session()
    
    # Specify training table location
    table_name = "TRAINING_TABLE"
    
    # Load table into DataConnector
    data_connector = DataConnector.from_dataframe(session.table(table_name))
    
    pandas_df = data_connector.to_pandas()
    label_column_name = 'TARGET'
    X, y = pandas_df.drop(label_column_name, axis=1), pandas_df[label_column_name]
    
    clf = xgb.Classifier()
    clf.fit(X, y)

    Training is rarely a single run. You tune hyperparameters, either with Snowflake ML's distributed HPO or with open-source libraries such as hyperopt or optuna. Each candidate run produces a result, and experiments keep those results organized. A run can be logged to an experiment while it trains on Snowflake, or you can upload metadata and artifacts from training done earlier. You then compare all the runs in Snowsight and pick the model that goes forward to deployment.

    Checkpoint 7 of 7· Check yourself

    A data scientist has run twenty training configurations and needs to compare them against the same metrics before choosing one for production. Which Snowflake ML capability is designed for this?

    Sources4

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Re-implementing training-time preprocessing in production SQL is fine as long as the logic is equivalent.Why is that wrong?

      Two separate implementations of the same logic drift apart and cause training-serving skew. Define the features once in a Feature Store so training and inference read the same definitions.

      Covered in Feature engineering: define once, use in training and serving

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “ML Lineage is a capability to trace end-to-end lineage of ML artifacts from source data to features, datasets, and models.”
      ↩︎ The lifecycle as a chain of stages
      “supports automated, incremental refresh from batch and streaming data sources”
      ↩︎ Data collection: getting data in and freezing it
      “Snowflake ML is an integrated set of capabilities for end-to-end machine learning in a single platform on top of your governed data.”
      ↩︎ Key concept
      “ensures consistency between training and inference to reduce training-serving skew”
      ↩︎ Exam trap 1
      “Deploy your model for inference at scale with the Snowflake Model Registry”
      ↩︎ Checkpoint
      “ensures consistency between training and inference to reduce training-serving skew”
      ↩︎ Prediction
      “It supports both batch and low-latency online feature retrieval, and ensures consistency between training and inference to reduce training-serving skew.”
      ↩︎ Checkpoint
      “Use experiments to record the results of your model training, and evaluate a collection of models in an organized way.”
      ↩︎ Checkpoint
    2. 2.
      “upload new data to Snowflake from local files, external cloud storage, or datasets from the Snowflake Marketplace”
      ↩︎ Data collection: getting data in and freezing it
      “perform exploratory data analysis, develop machine learning models”
      ↩︎ Visualization and exploration: find problems before you model
    3. 3.
      “Charts let you quickly identify and understand patterns and outliers in data.”
      ↩︎ Visualization and exploration: find problems before you model
    4. 4.
      “Leverage distributed preprocessing functions in snowflake.ml.modeling.preprocessing.”
      ↩︎ Feature engineering: define once, use in training and serving
      “You can train a model within either a Snowflake Notebook or a Snowflake ML Job.”
      ↩︎ Training models and comparing them with experiments

    Also cited

    Continue to page 2 of 2

    Evaluating, Explaining, Deploying, Monitoring and Versioning ML Models in Snowflake

    Spotted a mistake, or was something unclear? Tell us.