CertSafari
    Snowflake SnowPro Advanced: MLOps Engineer (MLA-B01)· Lessons

    Domain 2 · Lesson 7/17

    Notebooks in Workspaces: Notebook Services, GPU Allocation and When to Use Them

    Utilize Snowflake Workspaces.

    12 min read
    8% of exam
    7 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Explain how a notebook service maps to compute-pool nodes and how distributed training runs inside Notebooks in Workspaces
    • Identify the default GPU allocation for a single notebook service and how to get dedicated or larger compute
    • Predict what is lost when a notebook service suspends, idles out or is restarted for maintenance
    • Choose between a notebook and a Python file in Workspaces, and know where the provided documentation stops
    • Build and evaluate a model with open-source packages in a notebook, loading data with DataConnector and installing extra packages safely

    Key concept

    Notebook service — The Snowflake-managed container service that hosts your notebook kernel. It belongs to one user and runs on exactly one node of the compute pool you pick. Every notebook or Python file attached to it shares that node's CPUs, memory and GPUs.

    1.Where notebook code actually runs

    Notebooks in Workspaces replaces the Legacy Notebooks experience. It gives you a Jupyter-compatible editor with file management and a terminal, and your code runs in a pre-built container environment built for AI/ML work. For ML, Snowflake points to two strengths. First, you can run the whole lifecycle, from reading source data to model inference, in a single notebook that inherits the platform's existing governance. Second, the architecture supports scalable model development: you get distributed data loading and training on CPU or GPU compute pools without configuring distributed infrastructure yourself.

    Running a notebook or a .py file creates a notebook service. This is a Snowflake-managed service that hosts the kernel. When you create one, you choose the Python version, the Snowflake Container Runtime version, the compute pool, the idle timeout and any external access integrations. The service belongs to a single user and runs on exactly one node of the chosen compute pool. You can share a service across Workspaces, but it is always scoped to one user.

    Inside that container, the Snowflake ML APIs spread the work across whatever resources are available. When one node is not enough, Container Runtime 2.3 and above supports multi-node clusters, so frameworks such as PyTorch, XGBoost and LightGBM can train across several nodes.

    Checkpoint 1 of 6· Check yourself

    A data scientist connects a notebook in Workspaces to a new notebook service on a GPU compute pool. What does that service occupy?

    Sources123

    2.Default GPU allocation for a notebook service

    Two facts together give you the default GPU allocation for one notebook instance. First, the service sits on one node, so the most GPUs it can see is whatever that node's instance has. Second, Snowflake ML's distributed processing framework uses all GPUs on a multi-GPU node by default. You do not have to add code to use the node's other GPUs. You can change the number of GPUs, and the memory, for each task through the provided APIs.

    The second constraint is sharing. The node's resources belong to the service, not to each notebook. Every notebook and Python file attached to the same service shares them. If a training run needs the GPUs to itself, create a separate notebook service for it and don't attach anything else. There is also a hardware limit: notebook services run only on x86-based compute pools, and ARM-based pools are not yet supported.

    What one notebook service gets, and how to change it
    ConstraintDefault behaviourHow to change it
    Nodes per serviceOne node on the selected compute poolUse multi-node clusters (Container Runtime 2.3+) with distributed APIs
    GPUs on that nodeSnowflake ML distributed framework uses all of themConfigure GPU and memory per task through the APIs
    SharingAll attached notebooks and Python files share the nodeCreate a separate notebook service for dedicated compute
    ScopeSingle user; can be shared across WorkspacesEach user creates their own services

    Checkpoint 2 of 6· Exam question

    A data science lead asks what compute foundation Snowflake Notebooks in Workspaces use so the team can plan GPU access and package flexibility. Which statement is accurate?

    Checkpoint 3 of 6· Check yourself

    Three notebooks are attached to one GPU notebook service. One of them runs a long deep-learning job that should not compete for the node's GPUs. What is the documented fix?

    Sources34

    3.Suspension, idle timeout and maintenance windows

    Long training sessions depend on the service's lifecycle. You can suspend a service manually, or it suspends itself when it reaches its idle timeout. Suspending wipes the session. It disconnects every attached notebook, clears in-memory state, and removes installed packages and variables. Files that your code or the terminal created in the Workspace file system and in /tmp are deleted too. To resume, connect a notebook to the service or run a notebook that was connected to it before.

    Maintenance adds a time limit. Notebook services are a type of Snowpark Container Services and need periodic maintenance, which takes about five minutes and suspends and restarts the service. Once a service enters RUNNING, it is guaranteed not to be disrupted by maintenance for seven calendar days (168 hours). After that, it may be suspended for mandatory maintenance. The limitations page also warns that services may be restarted over the weekend, after which you must rerun notebooks and reinstall packages. Plan multi-day work around these windows, and save anything you need outside the session.

    Notebook service lifecycle events and their effect
    EventWhat happens
    Manual Suspend or idle timeout reachedNotebooks disconnect; in-memory state, packages, variables and created files (including /tmp) are removed
    Service enters RUNNINGNo maintenance disruption for seven calendar days (168 hours)
    Seven days after creationService may be suspended for mandatory maintenance (about five minutes)
    Connect or run a previously connected notebookSuspended service resumes

    Administrators can list notebook services, monitor consumption per compute pool, and apply budgets to specific pools. Snowflake recommends a separate compute pool for each role if you want to see consumption by role.

    List the notebook services in the accountsql
    SHOW SERVICES OF TYPE NOTEBOOK;

    Checkpoint 4 of 6· Match them up

    Match each lifecycle event to its effect on a notebook service

    Tap a term, then the definition that fits it.

    Sources24

    4.Notebook, Python file, or something else?

    Workspaces runs two kinds of executable file on the same infrastructure. Python files (.py) use the same Snowflake-managed notebook services as notebooks. When you run either one for the first time, you configure the Python version, runtime version, compute pool, idle timeout and integrations in the same way. A single running service can serve both. The difference is how code runs. A notebook runs cell by cell, which suits interactive exploration. A Python file's Run executes the whole file in one pass, as an IDE or terminal would, with no per-cell control. Snowflake recommends Python files for script-style work in the same governed environment, such as utilities, training scripts, or ETL jobs that you also deploy to production.

    Notebook vs Python file in Workspaces
    AspectNotebook (.ipynb)Python file (.py)
    ComputeSnowflake-managed notebook serviceSame notebook services as notebooks
    ExecutionInteractive, cell by cellRun executes the entire file in one pass; no per-cell Run all
    Best fitExploration and end-to-end interactive ML workScript-style workflows: utilities, training scripts, ETL jobs deployed to production

    Notebooks can also go to production. You can use the native scheduler or call notebooks from orchestration scripts. If you need to run ML work asynchronously from any development environment instead, the documentation points to Snowflake ML Jobs, which use the same Container Runtime. Note one SQL quirk: queries in notebook SQL cells don't appear in Query History until you shut down the kernel.

    Checkpoint 5 of 6· Check yourself

    An engineer has a training script that should run top to bottom, the way it runs from a terminal, and will later be deployed to production. Everything should stay in the governed Workspaces environment. What should they use?

    Sources51

    5.Building and evaluating models with open-source packages

    You do not have to give up familiar open-source tools to train inside Snowflake. The Container Runtime that backs Notebooks in Workspaces includes approximately 100 popular data science and machine learning packages, plus first-party packages such as snowflake-snowpark-python. The pre-installed set includes scikit-learn, numpy and scipy, and you can also use XGBoost and LightGBM. CPU and GPU runtimes ship different package sets, and the full list is in the Container Runtime release notes. Your data stays in Snowflake while you work with these libraries.

    The usual pattern has three steps. First, load the data with the DataConnector, which parallelizes reads across compute nodes, and convert it to a pandas DataFrame. Second, build the model with the open-source library: split the features and label, then fit a scikit-learn or XGBoost estimator. Third, evaluate it. The documented example splits the data with scikit-learn's train_test_split, fits an XGBClassifier on the training portion and calls predict on the held-out portion. When the data or model outgrows one node, the same workflow can move to Snowflake's distributed versions of XGBoost, LightGBM and PyTorch, which have APIs similar to the standard ones. You can also pass a DataConnector straight to those distributed estimators instead of converting to pandas first.

    Load a Snowflake table through a DataConnector (then call to_pandas() and train with any open-source library)python
    data_connector = DataConnector.from_dataframe(session.table(table_name))

    If a package is not in the runtime, install it. After an administrator configures External Access Integrations (EAIs), you can run pip install in a Python cell or in the notebook terminal to pull from sources such as PyPI. Other options are an artifact repository, .whl or .py files uploaded to the workspace, or a custom image. You can pin versions in a requirements.txt and install it with !pip install -r requirements.txt. Remember that a suspended service removes installed packages, so reinstall them after a resume.

    Checkpoint 6 of 6· Check yourself

    A notebook in Workspaces needs a PyPI package that is not in the default Container Runtime. Which approach does the documentation describe?

    Sources673

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Each notebook attached to a GPU service gets its own GPU, or a notebook service can use GPUs across the whole compute pool.Why is that wrong?

      A service occupies one node, and every notebook attached to it shares that node's resources. For dedicated compute, create a separate service. To go beyond one node, use multi-node clusters with the distributed APIs.

      Covered in Default GPU allocation for a notebook service

    2. 2.A notebook session left running stays intact indefinitely, so in-memory models and pip-installed packages will still be there more than a week later.Why is that wrong?

      The no-maintenance guarantee lasts seven days from RUNNING. After that, or after a weekend restart or an idle timeout, the service suspends, and packages and variables are gone.

      Covered in Suspension, idle timeout and maintenance windows

    3. 3.Any package version can be installed freely with pip in a notebook, including one that conflicts with the runtime's pre-installed packages.Why is that wrong?

      A requirements.txt version that conflicts with the supported pre-installed versions can break the Python environment, so validate compatibility before installing.

      Covered in Building and evaluating models with open-source packages

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Users can access distributed data loading and training across designated CPU or GPU compute pools.”
      ↩︎ Where notebook code actually runs
      “Use the native scheduler or incorporate notebooks into orchestration scripts for production pipelines.”
      ↩︎ Notebook, Python file, or something else?
    2. 2.
      “This notebook is optimized for Snowflake Container Runtime 2.3 or above, which introduces support for multi-node clusters.”
      ↩︎ Where notebook code actually runs
      “Any files created from code or the terminal in the Workspace file system and the /tmp directory are also removed.”
      ↩︎ Suspension, idle timeout and maintenance windows
      “Each notebook service is scoped to a single user and occupies one node on the selected compute pool.”
      ↩︎ Key concept
      “All notebooks and Python files connected to the same service share the compute resources on that node.”
      ↩︎ Exam trap 1
      “it is guaranteed not to be disrupted for seven calendar days (168 hours) due to service maintenance.”
      ↩︎ Exam trap 2
      “Each notebook service is scoped to a single user and occupies one node on the selected compute pool.”
      ↩︎ Checkpoint
      “If a notebook or Python file requires dedicated compute resources, create a separate notebook service”
      ↩︎ Checkpoint
      “Suspending a service disconnects all notebooks connected to it, clears in-memory states, and removes all packages and variables.”
      ↩︎ Checkpoint
    3. 3.
      “When using the Snowflake ML APIs, the Container Runtime distributes the processing across available resources.”
      ↩︎ Where notebook code actually runs
      “The number of resources, including GPU and memory allocation for each task, can be easily configured through the provided APIs.”
      ↩︎ Default GPU allocation for a notebook service
      “users can employ familiar and innovative open source frameworks inside Snowflake Notebooks, without moving data out of Snowflake”
      ↩︎ Building and evaluating models with open-source packages
      “By default, this framework uses all GPUs on multi-GPU nodes, offering significant performance improvements compared to open-source packages and reduces overall runtime.”
      ↩︎ Prediction
    4. 4.
      “Notebook services can only run on x86-based compute pools. ARM-based compute pools are not yet supported.”
      ↩︎ Default GPU allocation for a notebook service
      “After a restart, you must rerun notebooks and reinstall any packages to restore variables and packages.”
      ↩︎ Suspension, idle timeout and maintenance windows
    5. 5.
      “Python files use the same Snowflake-managed notebook services as notebooks in Workspaces.”
      ↩︎ Notebook, Python file, or something else?
      “Run executes the entire file in one pass”
      ↩︎ Notebook, Python file, or something else?
      “Use Python files when you want a script-style workflow in the same governed Workspaces environment as notebooks”
      ↩︎ Checkpoint
    6. 6.
      “Snowflake Container Runtime includes approximately 100 packages and libraries that support a wide range of analytics, data engineering, and machine learning development tasks inside Snowflake.”
      ↩︎ Building and evaluating models with open-source packages
      “If the package version specified in requirements.txt conflicts with supported versions of the pre-installed packages, the Python environment may break.”
      ↩︎ Exam trap 3
      “After configuring External Access Integrations (EAIs) for secure repository access, you can install packages directly from external sources such as PyPI.”
      ↩︎ Checkpoint
    7. 7.
      “The DataConnector accelerates data loading and pandas dataframe conversion.”
      ↩︎ Building and evaluating models with open-source packages

    Continue to page 2 of 2

    Open-Source Packages and DataConnector for ML Training in Snowflake Notebooks

    Spotted a mistake, or was something unclear? Tell us.