What you will be able to do
- Pick the right way to add open-source packages to a notebook's Container Runtime
- Load Snowflake tables or Datasets with DataConnector into pandas, PyTorch or TensorFlow
- Pass a DataConnector directly to Snowflake's distributed XGBoost and PyTorch training APIs
1.Open-source packages in the Container Runtime
Snowflake Notebooks run on Snowflake Container Runtime, which comes with about 100 data science and ML packages preinstalled, including scikit-learn, numpy and scipy, plus first-party packages such as snowflake-snowpark-python. You choose Python 3.10 to 3.12 when you create the service. The point is to use familiar open-source frameworks without moving data out of Snowflake. One detail matters for exam questions: the CPU and GPU runtimes ship different packages. For deep learning with PyTorch, TensorFlow and similar frameworks, use a GPU-powered runtime image, which draws GPU compute on demand from your compute pools.
When the default runtime doesn't have what you need, there are four ways to add packages. They trade ease of use against governance and control.
| Option | When to use it | Effort |
|---|---|---|
| Uploaded .whl or .py files (Workspaces or Stages) | Extra or private packages; use a Shared workspace or stage for team-wide files | Easy: upload, then !pip install and import |
| Artifact Repositories (Preview) | Extra packages with package policy enforcement; Snowflake PyPI repo open to PUBLIC by default | Easy; non-Snowflake repos need admin setup |
| External Access Integrations (EAIs) | Reach external repos or network resources such as internal PyPI, Artifactory or GitHub | Harder: account admin must provision |
| Custom Images (Preview) | Full control: system-level dependencies, security tooling, standardized images | Most effort; startup may be slower |
You run pip install in a Python cell or in the terminal. To pin versions, use !pip install -r requirements.txt, but check compatibility first: a version that conflicts with the preinstalled packages can break the environment. If you select an artifact repository, packages are not installed through EAIs. Installing from external stages is not supported, though you can pull a module from an internal stage with the Snowpark session:
from snowflake.snowpark.context import get_active_session
import sys
session = get_active_session()
session.file.get("@db.schema.stage_name/math_tools.py", "/tmp")
sys.path.append("/tmp")
import math_tools
math_tools.add_one(3)Checkpoint 1 of 5· Check yourself
A platform team wants data scientists to install extra Python packages that aren't in the default runtime, but only from a curated source with package policy enforcement. Which option fits best?
Artifact Repositories are the option that supports package policies. Installing from external stages is not supported, and an EAI to public PyPI enforces no policy.
“Use when you need additional Python packages not included in the default runtime or from private/curated repositories, with package policy enforcement.”Source: docs.snowflake.com
2.DataConnector: faster ingestion into open-source objects
Reading a large table straight into an object like a pandas dataframe can be slow. DataConnector, in the snowflake.ml.data namespace, speeds this up by parallelizing the reads across the compute pool's nodes. You build it from one of two inputs. A Snowpark DataFrame gives direct access to tables and is best during development. A Snowflake Dataset is a versioned schema-level object and is best for production. (from_sources accepts a list of either.) You then convert the connector into whatever your open-source library expects.
| Method | Input or output | Typical consumer |
|---|---|---|
| DataConnector.from_dataframe | Snowpark DataFrame (development) | Any of the outputs below |
| DataConnector.from_dataset | Snowflake Dataset (production, versioned) | Any of the outputs below |
| to_pandas | pandas dataframe | scikit-learn, XGBoost and other pandas-compatible libraries |
| to_torch_dataset | PyTorch dataset | PyTorch DataLoader |
| to_tf_dataset | TensorFlow dataset, streamed | TensorFlow models |
from snowflake.ml.data.data_connector import DataConnector
from snowflake.snowpark.context import get_active_session
import xgboost as xgb
session = get_active_session()
# Specify training table location
table_name = "TRAINING_TABLE"
# Load table into DataConnector
data_connector = DataConnector.from_dataframe(session.table(table_name))
# Convert to pandas dataframe
pandas_df = data_connector.to_pandas()From there it's ordinary open-source code: split features and label, then call fit. DataConnector speeds up both the load and the pandas conversion, so the gain is in ingestion, not in the training library. For deep learning, to_torch_dataset(batch_size=32) feeds a standard PyTorch DataLoader. to_tf_dataset streams batches into TensorFlow.
Checkpoint 2 of 5· Exam question
An administrator wants to know which compute pools exist in every Snowflake account for notebook workloads without any setup. Which pair is correct?
Correct answer: D — SYSTEM_COMPUTE_POOL_CPU and SYSTEM_COMPUTE_POOL_GPU, provisioned automatically for notebooks on Container Runtime.
- A. Incorrect. These names do not exist; notebooks rely on the two SYSTEM_COMPUTE_POOL_* pools.
- B. Incorrect. Notebook Container Runtime uses compute pools, not warehouses, and no such warehouses are created.
- C. Incorrect. The system pools already exist; creating a custom pool is optional and uses names you choose.
- D. Correct. Snowflake provisions one CPU and one GPU system compute pool per account for notebook use.
Checkpoint 3 of 5· Fill the gap
Which DataConnector method produces a streamed dataset for TensorFlow?
# Convert to TensorFlow dataset
tf_ds = data_connector. ? (
batch_size=4,
shuffle=True,
drop_last_batch=True
)to_tf_dataset returns a TensorFlow dataset loaded in a streaming fashion. to_torch_dataset targets PyTorch, and from_dataset builds a connector rather than converting one.
Source: docs.snowflake.com3.Skipping conversion: DataConnector into distributed trainers
Converting to pandas still puts all the data in one process's memory, and training on large datasets can exceed a single node's resources. For the best performance, pass the DataConnector straight to Snowflake's distributed training APIs in snowflake.ml.modeling.distributors. These provide distributed versions of XGBoost, LightGBM and PyTorch, with APIs close to the standard ones. A scaling config controls resources, for example whether XGBoost uses GPUs.
# Create DataConnector from a Snowpark dataframe
snowflake_df = session.table("TRAINING_TABLE")
data_connector = DataConnector.from_dataframe(snowflake_df)
# Create Snowflake XGBoost estimator
snowflake_est = XGBEstimator(
n_estimators=1,
objective="reg:squarederror",
scaling_config=XGBScalingConfig(use_gpu=False),
)
# Train using the data connector
# When using a data connector, input_cols and label_col must be provided
fit_booster = snowflake_est.fit(
data_connector,
input_cols=NUMERICAL_COLS,
label_col=LABEL_COL
)For PyTorch, ShardedDataConnector splits the data so that each worker process gets its own shard, which it reads with context.get_dataset_map()["train"].get_shard(). The trainer's ScalingConfig sets the number of nodes, the workers per node, and the CPUs and GPUs per worker. This is where you set GPU allocation explicitly instead of accepting the use-all-GPUs default.
# Create PyTorch trainer with scaling configuration
pytorch_trainer = PyTorchTrainer(
train_func=train_func,
scaling_config=ScalingConfig(
num_nodes=1,
num_workers_per_node=4,
resource_requirements_per_worker=WorkerResourceConfig(num_cpus=1, num_gpus=0),
),
)Checkpoint 4 of 5· Exam question
Three data scientists share a custom compute pool with `MAX_NODES = 1` for their Workspaces notebooks. The second and third users cannot start their notebooks while the first is running. What is the cause and the fix?
Correct answer: B — One notebook runs per user on each node, so MAX_NODES must exceed one to give concurrent users separate nodes.
- A. Incorrect. There is no account-wide limit of one notebook; the constraint is node capacity within the pool.
- B. Correct. Notebook services are limited per node, and the pool needs additional nodes to host other users' notebooks at once.
- C. Incorrect. Missing privileges would block access to objects, not prevent a pool from scheduling a second user's service.
- D. Incorrect. Notebooks on Container Runtime do not use warehouse clusters; capacity comes from compute pool nodes.
Checkpoint 5 of 5· Check yourself
You pass a DataConnector, not a pandas dataframe, to Snowflake's XGBEstimator.fit. What must you also supply?
Fitting the distributed estimator from a DataConnector requires explicit input and label columns. Converting to pandas first is exactly the step this path avoids.
“When using a data connector, input_cols and label_col must be provided”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.pip install -r requirements.txt is always safe because it just adds packages on top of the runtime.Why is that wrong?
Pinned versions that conflict with the preinstalled packages can break the Python environment, so check compatibility before installing.
2.DataConnector only helps by producing a pandas dataframe, so large-model training always ends in single-process pandas.Why is that wrong?
You can pass a DataConnector straight to the distributed XGBoost, LightGBM and PyTorch APIs, and ShardedDataConnector spreads shards across nodes for distributed PyTorch training.
Covered in Skipping conversion: DataConnector into distributed trainers
Practise it for real
Train an open-source scikit-learn model in a Workspaces notebook, loading the data through DataConnector
1.In a notebook connected to a notebook service, run
session = get_active_session()and createdata_connector = DataConnector.from_dataframe(session.table(table_name)), using your own training table.Why: from_dataframe wraps a Snowpark DataFrame, the input recommended during development.
You should see: A DataConnector object is created without errors.
2.Call
pandas_df = data_connector.to_pandas().Why: DataConnector parallelizes the read and the pandas conversion.
You should see: A pandas dataframe containing the table's rows.
3.Split into
X, y = pandas_df.drop(label_column_name, axis=1), pandas_df[label_column_name]and fitLogisticRegression(max_iter=1000).Why: scikit-learn ships with the Container Runtime, so no install is needed.
You should see: The model fits and returns a fitted estimator.
4.Suspend the service, reconnect, and check whether
pandas_dfstill exists.Why: Suspension clears in-memory state and variables.
You should see: A NameError, because variables don't survive a suspend.
Stuck? Get a nudge
If the label column isn't called TARGET, change label_column_name to match your table.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The CPU container runtime has different packages than the GPU container runtime.”
↩︎ Open-source packages in the Container Runtime“users can employ familiar and innovative open source frameworks inside Snowflake Notebooks, without moving data out of Snowflake.”
↩︎ Open-source packages in the Container Runtime“The APIs of the distributed classes are similar to those of the standard versions.”
↩︎ Skipping conversion: DataConnector into distributed trainers - 2.
“You can use a GPU-powered container runtime image to train deep learning models with PyTorch, TensorFlow, and other frameworks.”
↩︎ Open-source packages in the Container Runtime“The DataConnector accelerates data loading and pandas dataframe conversion.”
↩︎ DataConnector: faster ingestion into open-source objects“Training ML models on large datasets can exceed the resources of a single node.”
↩︎ Skipping conversion: DataConnector into distributed trainers - 3.https://docs.snowflake.com/en/user-guide/ui-snowsight/notebooks-in-workspaces/notebooks-in-workspaces-limitationsOfficial docs
“Installing packages from external stages is not supported.”
↩︎ Open-source packages in the Container Runtime - 4.
“The DataConnector accelerates data loading by parallelizing the reads across multiple compute nodes.”
↩︎ DataConnector: faster ingestion into open-source objects“Snowflake Datasets: Versioned schema-level objects. Best used for production workflows.”
↩︎ DataConnector: faster ingestion into open-source objects“You can use the ShardedDataConnector to shard your data across multiple nodes for distributed training with the Snowflake PyTorch distributor.”
↩︎ Exam trap 2“When using a data connector, input_cols and label_col must be provided”
↩︎ Checkpoint
Also cited
“If the package version specified in requirements.txt conflicts with supported versions of the pre-installed packages, the Python environment may break.”
↩︎ Exam trap 1“Use when you need additional Python packages not included in the default runtime or from private/curated repositories, with package policy enforcement.”
↩︎ Checkpoint