CertSafari
    Snowflake SnowPro Advanced: MLOps Engineer (MLA-B01)· Lessons

    Domain 2 · Lesson 7/17

    Open-Source Packages and DataConnector for ML Training in Snowflake Notebooks

    Utilize Snowflake Workspaces.

    9 min read
    8% of exam
    5 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Pick the right way to add open-source packages to a notebook's Container Runtime
    • Load Snowflake tables or Datasets with DataConnector into pandas, PyTorch or TensorFlow
    • Pass a DataConnector directly to Snowflake's distributed XGBoost and PyTorch training APIs

    1.Open-source packages in the Container Runtime

    Snowflake Notebooks run on Snowflake Container Runtime, which comes with about 100 data science and ML packages preinstalled, including scikit-learn, numpy and scipy, plus first-party packages such as snowflake-snowpark-python. You choose Python 3.10 to 3.12 when you create the service. The point is to use familiar open-source frameworks without moving data out of Snowflake. One detail matters for exam questions: the CPU and GPU runtimes ship different packages. For deep learning with PyTorch, TensorFlow and similar frameworks, use a GPU-powered runtime image, which draws GPU compute on demand from your compute pools.

    When the default runtime doesn't have what you need, there are four ways to add packages. They trade ease of use against governance and control.

    Ways to add packages beyond the default runtime
    OptionWhen to use itEffort
    Uploaded .whl or .py files (Workspaces or Stages)Extra or private packages; use a Shared workspace or stage for team-wide filesEasy: upload, then !pip install and import
    Artifact Repositories (Preview)Extra packages with package policy enforcement; Snowflake PyPI repo open to PUBLIC by defaultEasy; non-Snowflake repos need admin setup
    External Access Integrations (EAIs)Reach external repos or network resources such as internal PyPI, Artifactory or GitHubHarder: account admin must provision
    Custom Images (Preview)Full control: system-level dependencies, security tooling, standardized imagesMost effort; startup may be slower

    You run pip install in a Python cell or in the terminal. To pin versions, use !pip install -r requirements.txt, but check compatibility first: a version that conflicts with the preinstalled packages can break the environment. If you select an artifact repository, packages are not installed through EAIs. Installing from external stages is not supported, though you can pull a module from an internal stage with the Snowpark session:

    Import a module from a stage into the notebook containerpython
    from snowflake.snowpark.context import get_active_session
    import sys
    
    session = get_active_session()
    session.file.get("@db.schema.stage_name/math_tools.py", "/tmp")
    
    sys.path.append("/tmp")
    import math_tools
    
    math_tools.add_one(3)

    Checkpoint 1 of 5· Check yourself

    A platform team wants data scientists to install extra Python packages that aren't in the default runtime, but only from a curated source with package policy enforcement. Which option fits best?

    Sources123

    2.DataConnector: faster ingestion into open-source objects

    Reading a large table straight into an object like a pandas dataframe can be slow. DataConnector, in the snowflake.ml.data namespace, speeds this up by parallelizing the reads across the compute pool's nodes. You build it from one of two inputs. A Snowpark DataFrame gives direct access to tables and is best during development. A Snowflake Dataset is a versioned schema-level object and is best for production. (from_sources accepts a list of either.) You then convert the connector into whatever your open-source library expects.

    DataConnector inputs and outputs
    MethodInput or outputTypical consumer
    DataConnector.from_dataframeSnowpark DataFrame (development)Any of the outputs below
    DataConnector.from_datasetSnowflake Dataset (production, versioned)Any of the outputs below
    to_pandaspandas dataframescikit-learn, XGBoost and other pandas-compatible libraries
    to_torch_datasetPyTorch datasetPyTorch DataLoader
    to_tf_datasetTensorFlow dataset, streamedTensorFlow models
    Load a table through DataConnector into pandas before training with open-source XGBoostpython
    from snowflake.ml.data.data_connector import DataConnector
    from snowflake.snowpark.context import get_active_session
    import xgboost as xgb
    
    session = get_active_session()
    
    # Specify training table location
    table_name = "TRAINING_TABLE"
    
    # Load table into DataConnector
    data_connector = DataConnector.from_dataframe(session.table(table_name))
    
    # Convert to pandas dataframe
    pandas_df = data_connector.to_pandas()

    From there it's ordinary open-source code: split features and label, then call fit. DataConnector speeds up both the load and the pandas conversion, so the gain is in ingestion, not in the training library. For deep learning, to_torch_dataset(batch_size=32) feeds a standard PyTorch DataLoader. to_tf_dataset streams batches into TensorFlow.

    Checkpoint 2 of 5· Exam question

    An administrator wants to know which compute pools exist in every Snowflake account for notebook workloads without any setup. Which pair is correct?

    Checkpoint 3 of 5· Fill the gap

    Which DataConnector method produces a streamed dataset for TensorFlow?

    # Convert to TensorFlow dataset
    tf_ds = data_connector. ? (
        batch_size=4,
        shuffle=True,
        drop_last_batch=True
    )

    Sources42

    3.Skipping conversion: DataConnector into distributed trainers

    Converting to pandas still puts all the data in one process's memory, and training on large datasets can exceed a single node's resources. For the best performance, pass the DataConnector straight to Snowflake's distributed training APIs in snowflake.ml.modeling.distributors. These provide distributed versions of XGBoost, LightGBM and PyTorch, with APIs close to the standard ones. A scaling config controls resources, for example whether XGBoost uses GPUs.

    Train Snowflake's distributed XGBEstimator directly from a DataConnectorpython
    # Create DataConnector from a Snowpark dataframe
    snowflake_df = session.table("TRAINING_TABLE")
    data_connector = DataConnector.from_dataframe(snowflake_df)
    
    # Create Snowflake XGBoost estimator
    snowflake_est = XGBEstimator(
        n_estimators=1,
        objective="reg:squarederror",
        scaling_config=XGBScalingConfig(use_gpu=False),
    )
    
    # Train using the data connector
    # When using a data connector, input_cols and label_col must be provided
    fit_booster = snowflake_est.fit(
        data_connector,
        input_cols=NUMERICAL_COLS,
        label_col=LABEL_COL
    )

    For PyTorch, ShardedDataConnector splits the data so that each worker process gets its own shard, which it reads with context.get_dataset_map()["train"].get_shard(). The trainer's ScalingConfig sets the number of nodes, the workers per node, and the CPUs and GPUs per worker. This is where you set GPU allocation explicitly instead of accepting the use-all-GPUs default.

    Scaling configuration for the Snowflake PyTorch trainerpython
    # Create PyTorch trainer with scaling configuration
    pytorch_trainer = PyTorchTrainer(
        train_func=train_func,
        scaling_config=ScalingConfig(
            num_nodes=1,
            num_workers_per_node=4,
            resource_requirements_per_worker=WorkerResourceConfig(num_cpus=1, num_gpus=0),
        ),
    )

    Checkpoint 4 of 5· Exam question

    Three data scientists share a custom compute pool with `MAX_NODES = 1` for their Workspaces notebooks. The second and third users cannot start their notebooks while the first is running. What is the cause and the fix?

    Checkpoint 5 of 5· Check yourself

    You pass a DataConnector, not a pandas dataframe, to Snowflake's XGBEstimator.fit. What must you also supply?

    Sources21

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.pip install -r requirements.txt is always safe because it just adds packages on top of the runtime.Why is that wrong?

      Pinned versions that conflict with the preinstalled packages can break the Python environment, so check compatibility before installing.

      Covered in Open-source packages in the Container Runtime

    2. 2.DataConnector only helps by producing a pandas dataframe, so large-model training always ends in single-process pandas.Why is that wrong?

      You can pass a DataConnector straight to the distributed XGBoost, LightGBM and PyTorch APIs, and ShardedDataConnector spreads shards across nodes for distributed PyTorch training.

      Covered in Skipping conversion: DataConnector into distributed trainers

    Practise it for real

    Train an open-source scikit-learn model in a Workspaces notebook, loading the data through DataConnector

    1. 1.In a notebook connected to a notebook service, run session = get_active_session() and create data_connector = DataConnector.from_dataframe(session.table(table_name)), using your own training table.

      Why: from_dataframe wraps a Snowpark DataFrame, the input recommended during development.

      You should see: A DataConnector object is created without errors.

    2. 2.Call pandas_df = data_connector.to_pandas().

      Why: DataConnector parallelizes the read and the pandas conversion.

      You should see: A pandas dataframe containing the table's rows.

    3. 3.Split into X, y = pandas_df.drop(label_column_name, axis=1), pandas_df[label_column_name] and fit LogisticRegression(max_iter=1000).

      Why: scikit-learn ships with the Container Runtime, so no install is needed.

      You should see: The model fits and returns a fitted estimator.

    4. 4.Suspend the service, reconnect, and check whether pandas_df still exists.

      Why: Suspension clears in-memory state and variables.

      You should see: A NameError, because variables don't survive a suspend.

    Stuck? Get a nudge

    If the label column isn't called TARGET, change label_column_name to match your table.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “The CPU container runtime has different packages than the GPU container runtime.”
      ↩︎ Open-source packages in the Container Runtime
      “users can employ familiar and innovative open source frameworks inside Snowflake Notebooks, without moving data out of Snowflake.”
      ↩︎ Open-source packages in the Container Runtime
      “The APIs of the distributed classes are similar to those of the standard versions.”
      ↩︎ Skipping conversion: DataConnector into distributed trainers
    2. 2.
      “You can use a GPU-powered container runtime image to train deep learning models with PyTorch, TensorFlow, and other frameworks.”
      ↩︎ Open-source packages in the Container Runtime
      “The DataConnector accelerates data loading and pandas dataframe conversion.”
      ↩︎ DataConnector: faster ingestion into open-source objects
      “Training ML models on large datasets can exceed the resources of a single node.”
      ↩︎ Skipping conversion: DataConnector into distributed trainers
    3. 4.
      “The DataConnector accelerates data loading by parallelizing the reads across multiple compute nodes.”
      ↩︎ DataConnector: faster ingestion into open-source objects
      “Snowflake Datasets: Versioned schema-level objects. Best used for production workflows.”
      ↩︎ DataConnector: faster ingestion into open-source objects
      “You can use the ShardedDataConnector to shard your data across multiple nodes for distributed training with the Snowflake PyTorch distributor.”
      ↩︎ Exam trap 2
      “When using a data connector, input_cols and label_col must be provided”
      ↩︎ Checkpoint

    Also cited

    Ready to test yourself?

    Practise the 29 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.