What you will be able to do
- Explain how a notebook service maps to compute-pool nodes and how distributed training runs inside Notebooks in Workspaces
- Identify the default GPU allocation for a single notebook service and how to get dedicated or larger compute
- Predict what is lost when a notebook service suspends, idles out or is restarted for maintenance
- Choose between a notebook and a Python file in Workspaces, and know where the provided documentation stops
- Build and evaluate a model with open-source packages in a notebook, loading data with DataConnector and installing extra packages safely
Key concept
Notebook service — The Snowflake-managed container service that hosts your notebook kernel. It belongs to one user and runs on exactly one node of the compute pool you pick. Every notebook or Python file attached to it shares that node's CPUs, memory and GPUs.
1.Where notebook code actually runs
Notebooks in Workspaces replaces the Legacy Notebooks experience. It gives you a Jupyter-compatible editor with file management and a terminal, and your code runs in a pre-built container environment built for AI/ML work. For ML, Snowflake points to two strengths. First, you can run the whole lifecycle, from reading source data to model inference, in a single notebook that inherits the platform's existing governance. Second, the architecture supports scalable model development: you get distributed data loading and training on CPU or GPU compute pools without configuring distributed infrastructure yourself.
Running a notebook or a .py file creates a notebook service. This is a Snowflake-managed service that hosts the kernel. When you create one, you choose the Python version, the Snowflake Container Runtime version, the compute pool, the idle timeout and any external access integrations. The service belongs to a single user and runs on exactly one node of the chosen compute pool. You can share a service across Workspaces, but it is always scoped to one user.
Inside that container, the Snowflake ML APIs spread the work across whatever resources are available. When one node is not enough, Container Runtime 2.3 and above supports multi-node clusters, so frameworks such as PyTorch, XGBoost and LightGBM can train across several nodes.
Checkpoint 1 of 6· Check yourself
A data scientist connects a notebook in Workspaces to a new notebook service on a GPU compute pool. What does that service occupy?
A notebook service is a per-user service that takes up one node on the pool you choose. Scaling past that node goes through the distributed APIs and multi-node clusters, not through the service itself.
“Each notebook service is scoped to a single user and occupies one node on the selected compute pool.”Source: docs.snowflake.com
2.Default GPU allocation for a notebook service
Two facts together give you the default GPU allocation for one notebook instance. First, the service sits on one node, so the most GPUs it can see is whatever that node's instance has. Second, Snowflake ML's distributed processing framework uses all GPUs on a multi-GPU node by default. You do not have to add code to use the node's other GPUs. You can change the number of GPUs, and the memory, for each task through the provided APIs.
The second constraint is sharing. The node's resources belong to the service, not to each notebook. Every notebook and Python file attached to the same service shares them. If a training run needs the GPUs to itself, create a separate notebook service for it and don't attach anything else. There is also a hardware limit: notebook services run only on x86-based compute pools, and ARM-based pools are not yet supported.
| Constraint | Default behaviour | How to change it |
|---|---|---|
| Nodes per service | One node on the selected compute pool | Use multi-node clusters (Container Runtime 2.3+) with distributed APIs |
| GPUs on that node | Snowflake ML distributed framework uses all of them | Configure GPU and memory per task through the APIs |
| Sharing | All attached notebooks and Python files share the node | Create a separate notebook service for dedicated compute |
| Scope | Single user; can be shared across Workspaces | Each user creates their own services |
Checkpoint 2 of 6· Exam question
A data science lead asks what compute foundation Snowflake Notebooks in Workspaces use so the team can plan GPU access and package flexibility. Which statement is accurate?
Correct answer: C — Notebooks run on Container Runtime, which is powered by Snowpark Container Services and runs on compute pools of CPU or GPU nodes.
- A. Incorrect. Python code runs on a container service in Snowflake, not in the browser, so heavy training does not depend on the client machine.
- B. Incorrect. Virtual warehouses have no GPU option; GPU workloads need a GPU compute pool, and resizing a warehouse only adds CPU and memory.
- C. Correct. Container Runtime is built on Snowpark Container Services, so notebooks execute on compute pool nodes rather than on a virtual warehouse.
- D. Incorrect. The stored procedure sandbox is the warehouse model; Container Runtime notebooks can also install packages from PyPI via external access integrations.
Checkpoint 3 of 6· Check yourself
Three notebooks are attached to one GPU notebook service. One of them runs a long deep-learning job that should not compete for the node's GPUs. What is the documented fix?
Attached notebooks share the node's resources. A workload that needs dedicated compute gets its own service. ARM-based pools are not supported for notebook services.
“If a notebook or Python file requires dedicated compute resources, create a separate notebook service”Source: docs.snowflake.com
3.Suspension, idle timeout and maintenance windows
Long training sessions depend on the service's lifecycle. You can suspend a service manually, or it suspends itself when it reaches its idle timeout. Suspending wipes the session. It disconnects every attached notebook, clears in-memory state, and removes installed packages and variables. Files that your code or the terminal created in the Workspace file system and in /tmp are deleted too. To resume, connect a notebook to the service or run a notebook that was connected to it before.
Maintenance adds a time limit. Notebook services are a type of Snowpark Container Services and need periodic maintenance, which takes about five minutes and suspends and restarts the service. Once a service enters RUNNING, it is guaranteed not to be disrupted by maintenance for seven calendar days (168 hours). After that, it may be suspended for mandatory maintenance. The limitations page also warns that services may be restarted over the weekend, after which you must rerun notebooks and reinstall packages. Plan multi-day work around these windows, and save anything you need outside the session.
| Event | What happens |
|---|---|
| Manual Suspend or idle timeout reached | Notebooks disconnect; in-memory state, packages, variables and created files (including /tmp) are removed |
| Service enters RUNNING | No maintenance disruption for seven calendar days (168 hours) |
| Seven days after creation | Service may be suspended for mandatory maintenance (about five minutes) |
| Connect or run a previously connected notebook | Suspended service resumes |
Administrators can list notebook services, monitor consumption per compute pool, and apply budgets to specific pools. Snowflake recommends a separate compute pool for each role if you want to see consumption by role.
SHOW SERVICES OF TYPE NOTEBOOK;Checkpoint 4 of 6· Match them up
Match each lifecycle event to its effect on a notebook service
Tap a term, then the definition that fits it.
Suspension, for any reason, discards session state. The maintenance guarantee lasts seven days from RUNNING, and connecting a notebook resumes a suspended service.
“Suspending a service disconnects all notebooks connected to it, clears in-memory states, and removes all packages and variables.”Source: docs.snowflake.com
4.Notebook, Python file, or something else?
Workspaces runs two kinds of executable file on the same infrastructure. Python files (.py) use the same Snowflake-managed notebook services as notebooks. When you run either one for the first time, you configure the Python version, runtime version, compute pool, idle timeout and integrations in the same way. A single running service can serve both. The difference is how code runs. A notebook runs cell by cell, which suits interactive exploration. A Python file's Run executes the whole file in one pass, as an IDE or terminal would, with no per-cell control. Snowflake recommends Python files for script-style work in the same governed environment, such as utilities, training scripts, or ETL jobs that you also deploy to production.
| Aspect | Notebook (.ipynb) | Python file (.py) |
|---|---|---|
| Compute | Snowflake-managed notebook service | Same notebook services as notebooks |
| Execution | Interactive, cell by cell | Run executes the entire file in one pass; no per-cell Run all |
| Best fit | Exploration and end-to-end interactive ML work | Script-style workflows: utilities, training scripts, ETL jobs deployed to production |
Notebooks can also go to production. You can use the native scheduler or call notebooks from orchestration scripts. If you need to run ML work asynchronously from any development environment instead, the documentation points to Snowflake ML Jobs, which use the same Container Runtime. Note one SQL quirk: queries in notebook SQL cells don't appear in Query History until you shut down the kernel.
Checkpoint 5 of 6· Check yourself
An engineer has a training script that should run top to bottom, the way it runs from a terminal, and will later be deployed to production. Everything should stay in the governed Workspaces environment. What should they use?
Python files run on the same notebook services as notebooks, and Run executes the whole file in one pass. Snowflake names training scripts and production-bound jobs as the use case.
“Use Python files when you want a script-style workflow in the same governed Workspaces environment as notebooks”Source: docs.snowflake.com
5.Building and evaluating models with open-source packages
You do not have to give up familiar open-source tools to train inside Snowflake. The Container Runtime that backs Notebooks in Workspaces includes approximately 100 popular data science and machine learning packages, plus first-party packages such as snowflake-snowpark-python. The pre-installed set includes scikit-learn, numpy and scipy, and you can also use XGBoost and LightGBM. CPU and GPU runtimes ship different package sets, and the full list is in the Container Runtime release notes. Your data stays in Snowflake while you work with these libraries.
The usual pattern has three steps. First, load the data with the DataConnector, which parallelizes reads across compute nodes, and convert it to a pandas DataFrame. Second, build the model with the open-source library: split the features and label, then fit a scikit-learn or XGBoost estimator. Third, evaluate it. The documented example splits the data with scikit-learn's train_test_split, fits an XGBClassifier on the training portion and calls predict on the held-out portion. When the data or model outgrows one node, the same workflow can move to Snowflake's distributed versions of XGBoost, LightGBM and PyTorch, which have APIs similar to the standard ones. You can also pass a DataConnector straight to those distributed estimators instead of converting to pandas first.
data_connector = DataConnector.from_dataframe(session.table(table_name))If a package is not in the runtime, install it. After an administrator configures External Access Integrations (EAIs), you can run pip install in a Python cell or in the notebook terminal to pull from sources such as PyPI. Other options are an artifact repository, .whl or .py files uploaded to the workspace, or a custom image. You can pin versions in a requirements.txt and install it with !pip install -r requirements.txt. Remember that a suspended service removes installed packages, so reinstall them after a resume.
Checkpoint 6 of 6· Check yourself
A notebook in Workspaces needs a PyPI package that is not in the default Container Runtime. Which approach does the documentation describe?
Once EAIs are configured, you can install directly from external sources such as PyPI with pip. Installed packages do not survive a suspension.
“After configuring External Access Integrations (EAIs) for secure repository access, you can install packages directly from external sources such as PyPI.”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Each notebook attached to a GPU service gets its own GPU, or a notebook service can use GPUs across the whole compute pool.Why is that wrong?
A service occupies one node, and every notebook attached to it shares that node's resources. For dedicated compute, create a separate service. To go beyond one node, use multi-node clusters with the distributed APIs.
2.A notebook session left running stays intact indefinitely, so in-memory models and pip-installed packages will still be there more than a week later.Why is that wrong?
The no-maintenance guarantee lasts seven days from RUNNING. After that, or after a weekend restart or an idle timeout, the service suspends, and packages and variables are gone.
3.Any package version can be installed freely with pip in a notebook, including one that conflicts with the runtime's pre-installed packages.Why is that wrong?
A requirements.txt version that conflicts with the supported pre-installed versions can break the Python environment, so validate compatibility before installing.
Covered in Building and evaluating models with open-source packages
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.snowflake.com/en/user-guide/ui-snowsight/notebooks-in-workspaces/notebooks-in-workspaces-overviewOfficial docs
“Users can access distributed data loading and training across designated CPU or GPU compute pools.”
↩︎ Where notebook code actually runs“Use the native scheduler or incorporate notebooks into orchestration scripts for production pipelines.”
↩︎ Notebook, Python file, or something else? - 2.https://docs.snowflake.com/en/user-guide/ui-snowsight/notebooks-in-workspaces/notebooks-in-workspaces-compute-setupOfficial docs
“This notebook is optimized for Snowflake Container Runtime 2.3 or above, which introduces support for multi-node clusters.”
↩︎ Where notebook code actually runs“Any files created from code or the terminal in the Workspace file system and the /tmp directory are also removed.”
↩︎ Suspension, idle timeout and maintenance windows“Each notebook service is scoped to a single user and occupies one node on the selected compute pool.”
↩︎ Key concept“All notebooks and Python files connected to the same service share the compute resources on that node.”
↩︎ Exam trap 1“it is guaranteed not to be disrupted for seven calendar days (168 hours) due to service maintenance.”
↩︎ Exam trap 2“Each notebook service is scoped to a single user and occupies one node on the selected compute pool.”
↩︎ Checkpoint“If a notebook or Python file requires dedicated compute resources, create a separate notebook service”
↩︎ Checkpoint“Suspending a service disconnects all notebooks connected to it, clears in-memory states, and removes all packages and variables.”
↩︎ Checkpoint - 3.
“When using the Snowflake ML APIs, the Container Runtime distributes the processing across available resources.”
↩︎ Where notebook code actually runs“The number of resources, including GPU and memory allocation for each task, can be easily configured through the provided APIs.”
↩︎ Default GPU allocation for a notebook service“users can employ familiar and innovative open source frameworks inside Snowflake Notebooks, without moving data out of Snowflake”
↩︎ Building and evaluating models with open-source packages“By default, this framework uses all GPUs on multi-GPU nodes, offering significant performance improvements compared to open-source packages and reduces overall runtime.”
↩︎ Prediction - 4.https://docs.snowflake.com/en/user-guide/ui-snowsight/notebooks-in-workspaces/notebooks-in-workspaces-limitationsOfficial docs
“Notebook services can only run on x86-based compute pools. ARM-based compute pools are not yet supported.”
↩︎ Default GPU allocation for a notebook service“After a restart, you must rerun notebooks and reinstall any packages to restore variables and packages.”
↩︎ Suspension, idle timeout and maintenance windows - 5.
“Python files use the same Snowflake-managed notebook services as notebooks in Workspaces.”
↩︎ Notebook, Python file, or something else?“Run executes the entire file in one pass”
↩︎ Notebook, Python file, or something else?“Use Python files when you want a script-style workflow in the same governed Workspaces environment as notebooks”
↩︎ Checkpoint - 6.
“Snowflake Container Runtime includes approximately 100 packages and libraries that support a wide range of analytics, data engineering, and machine learning development tasks inside Snowflake.”
↩︎ Building and evaluating models with open-source packages“If the package version specified in requirements.txt conflicts with supported versions of the pre-installed packages, the Python environment may break.”
↩︎ Exam trap 3“After configuring External Access Integrations (EAIs) for secure repository access, you can install packages directly from external sources such as PyPI.”
↩︎ Checkpoint - 7.
“The DataConnector accelerates data loading and pandas dataframe conversion.”
↩︎ Building and evaluating models with open-source packages