CertSafari
    Snowflake SnowPro Core Certification (COF-C03)· Lessons

    Domain 1 · Lesson 6/19

    Snowpark, Notebooks, Streamlit and Snowflake ML

    Explain AI/ML and application development features

    9 min read
    5.17% of exam
    4 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Explain how Snowpark DataFrames and UDFs push processing to Snowflake, and why DataFrame evaluation is lazy
    • Describe what Snowflake Notebooks offer for exploration, ML development and scheduled pipelines
    • Describe how Streamlit in Snowflake hosts, governs and runs data apps
    • Name the Snowflake ML components and match each to its stage of the ML lifecycle

    Key concept

    Bring the code to the data — Snowflake's developer and ML features run your Python, Java or Scala code, your apps and your models inside Snowflake, next to the governed data. You don't export the data to an outside cluster to work on it.

    1.Snowpark: DataFrames that run where the data lives

    Snowpark is the base layer for the other features on this page. It is a client library for writing data-processing code in a general-purpose language instead of building SQL strings by hand. Snowflake ships Snowpark libraries for three languages: Java, Python and Scala. You can write Snowpark code in local tools such as Jupyter, VS Code or IntelliJ, but the work itself runs in Snowflake. The docs compare it with the Snowflake Connector for Spark: Snowpark pushes down every operation, including UDFs, and needs no separate compute cluster. Snowflake handles the scale and compute management.

    The main abstraction is the DataFrame. You call methods such as select instead of writing 'select column_name' as a string, so you get code completion and type checking. A DataFrame describes a query. It does not hold results.

    Building a DataFrame does not run anything. collect() sends the query and returns Rows.python
    >>> # Create a DataFrame with the "id" and "name" columns from the "sample_product_data" table.
    >>> # This does not execute the query.
    >>> df = session.table("sample_product_data").select(col("id"), col("name"))
    
    >>> # Send the query to the server for execution and
    >>> # return a list of Rows containing the results.
    >>> results = df.collect()

    Lazy evaluation lets Snowpark combine many transformations into a single operation. That cuts the data transferred between the client and Snowflake and improves performance. The same idea applies to custom logic. When you define a user-defined function (UDF) inline, for example with a Python lambda, Snowpark pushes the code to the server. When you call it, it runs where the data is, and Snowflake can parallelise it across the data.

    Checkpoint 1 of 6· Fill the gap

    Which Snowpark function turns this Python lambda into a server-side function named my_udf?

    >>> from snowflake.snowpark.types import IntegerType
    >>> add_one =  ? (lambda x: x+1, return_type=IntegerType(), input_types=[IntegerType()], name="my_udf", replace=True)

    Checkpoint 2 of 6· Check yourself

    A team moving off the Snowflake Connector for Spark asks what Snowpark changes about compute. Which statement is correct?

    Checkpoint 3 of 6· Exam question

    A data scientist wants to explore a sales table with Python, SQL, and Markdown cells side by side, run the code against a Snowflake virtual warehouse, and keep the exploration artifact as a schema-level object that teammates can open in Snowsight. Which capability should they use?

    Sources1

    2.Snowflake Notebooks: a cell-based workspace in Snowsight

    Snowpark gives you an API. Snowflake Notebooks give you a place to use it interactively. A notebook is a development interface in Snowsight made of cells. Each cell holds Python, SQL or Markdown, and you run and compare results one cell at a time. One notebook can cover exploratory data analysis, ML model development and data engineering work.

    Abilities the documentation lists:

    - Work with data already in Snowflake, or upload it from local files, external cloud storage or the Snowflake Marketplace. - Visualise results with embedded Streamlit visualisations and libraries such as Altair, Matplotlib or seaborn. - Sync with a Git repository for version control. - Annotate results in Markdown cells. - Run the notebook on a schedule to automate a pipeline. - Use role-based access control, so other users with the same role can view and collaborate.

    Notebooks are also where ML training starts. Notebooks on Container Runtime provide a Jupyter-like environment for training and fine-tuning large models without managing infrastructure.

    A naming note: the original experience has been renamed Legacy Notebooks. Notebooks in Workspaces is the next generation and adds Jupyter compatibility.

    Checkpoint 4 of 6· Check yourself

    A data engineer has prototyped a transformation in a Snowflake Notebook and now wants it to run automatically every night. What does the documentation say Notebooks support for this?

    Sources2

    3.Streamlit in Snowflake: data apps without moving data

    A notebook is a workspace for developers. Streamlit is how you share results with people who don't write code. Streamlit is an open-source Python library for building custom web apps for machine learning and data science. Streamlit in Snowflake lets you build, deploy and share those apps inside Snowflake. The app reads Snowflake data without the data or the app code moving to an external system.

    Three points are worth remembering:

    - Managed infrastructure. Snowflake manages the app's compute and storage. You choose whether it runs on a warehouse or a container runtime. - Governance through an object. The source code and environment configuration are stored in a Snowflake object, and RBAC controls who can access the app. - Works with the rest of the platform. Streamlit in Snowflake integrates with Snowpark, UDFs, stored procedures and the Snowflake Native App Framework. In Snowsight, a side-by-side editor and app preview let you edit and test quickly.

    Checkpoint 5 of 6· Check yourself

    How is access to a Streamlit in Snowflake app controlled?

    Sources3

    4.Snowflake ML: the end-to-end model lifecycle

    Snowflake ML connects the tools above into a full machine-learning workflow. The docs call it an integrated set of capabilities for end-to-end ML on top of your governed data, with distributed feature engineering, training and inference on CPU or GPU compute.

    Training runs in Container Runtime. This is a pre-built environment with packages such as PyTorch, XGBoost and Scikit-learn already installed, and you can add others from PyPI or Hugging Face. You reach it from Snowflake Notebooks. Teams that prefer an external IDE can use Snowflake ML Jobs to send functions, files or modules to Container Runtime and to automate pipelines. The platform is also open in both directions: models developed in Snowflake can be deployed outside it, and externally trained models can be brought in for inference.

    Snowflake ML components and their place in the lifecycle
    ComponentWhat it does
    Snowflake Feature StoreDefines, manages and stores features. Refreshes incrementally from batch and streaming sources, and keeps training and inference consistent
    Container RuntimePre-built environment for data loading, distributed training and hyperparameter tuning
    ExperimentsRecord training results and compare models to choose one for production
    Snowflake ML JobsDevelop and automate ML pipelines, including code sent from an external IDE
    Snowflake Model RegistryLog and manage models trained anywhere, and run inference at scale. Model Serving deploys to Snowpark Container Services
    ML Observability / ML ExplainabilityMonitor performance and drift with alerts, and compute Shapley values
    ML LineageTrace ML artifacts from source data to features, datasets and models

    For exam questions, place each component at its lifecycle stage: features → training → experiments → registry and inference → monitoring, with lineage covering every stage. If a question describes training-serving skew, the answer is the Feature Store. If it mentions Shapley values, the answer is ML Explainability.

    Checkpoint 6 of 6· Match them up

    Match each Snowflake ML component to the problem it solves

    Tap a term, then the definition that fits it.

    Sources4

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Assigning a Snowpark DataFrame runs the query and pulls the rows to the client.Why is that wrong?

      Snowpark DataFrames are lazy. Constructing one only defines the query, and an action such as collect() sends the SQL to Snowflake to run there.

      Covered in Snowpark: DataFrames that run where the data lives

    2. 2.The Snowflake Model Registry only accepts models trained inside Snowflake.Why is that wrong?

      The Registry logs and manages models whether or not they were trained in Snowflake, and externally trained models can be brought in for inference.

      Covered in Snowflake ML: the end-to-end model lifecycle

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Snowflake currently provides Snowpark libraries for three languages: Java, Python, and Scala.”
      ↩︎ Snowpark: DataFrames that run where the data lives
      “When you call the UDF in your client code, your custom code is executed on the server (where the data is).”
      ↩︎ Snowpark: DataFrames that run where the data lives
      “you can build applications that process data in Snowflake without moving data to the system where your application code runs”
      ↩︎ Key concept
      “Snowpark operations are executed lazily on the server”
      ↩︎ Exam trap 1
      “Snowpark operations are executed lazily on the server”
      ↩︎ Prediction
      “No requirement for a separate cluster outside of Snowflake for computations.”
      ↩︎ Checkpoint
    2. 2.
      “an interactive, cell-based programming environment for Python, SQL, and Markdown”
      ↩︎ Snowflake Notebooks: a cell-based workspace in Snowsight
      “Interactively visualize your data using embedded Streamlit visualizations and other libraries like Altair, Matplotlib, or seaborn.”
      ↩︎ Snowflake Notebooks: a cell-based workspace in Snowsight
      “Notebooks in Workspaces is the next generation of Snowflake Notebooks, providing Jupyter compatibility and full integration with Workspaces.”
      ↩︎ Snowflake Notebooks: a cell-based workspace in Snowsight
      “Run your notebook on a schedule to automate pipelines.”
      ↩︎ Checkpoint
    3. 3.
      “Snowflake manages the underlying compute and storage for your Streamlit app.”
      ↩︎ Streamlit in Snowflake: data apps without moving data
      “Streamlit in Snowflake works seamlessly with Snowpark, user-defined functions (UDFs), stored procedures, and Snowflake Native App Framework.”
      ↩︎ Streamlit in Snowflake: data apps without moving data
      “stores your source code and environment configuration within a Snowflake object that uses Role-based Access Control (RBAC)”
      ↩︎ Checkpoint
    4. 4.
      “an integrated set of capabilities for end-to-end machine learning in a single platform on top of your governed data”
      ↩︎ Snowflake ML: the end-to-end model lifecycle
      “You can use Model Serving to deploy the models to Snowpark Container Services for inference.”
      ↩︎ Snowflake ML: the end-to-end model lifecycle
      “use the ML Explainability function to compute Shapley values for models in the Snowflake Model Registry”
      ↩︎ Snowflake ML: the end-to-end model lifecycle
      “externally trained models can easily be brought into Snowflake for inference”
      ↩︎ Exam trap 2
      “ensures consistency between training and inference to reduce training-serving skew”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Snowflake Cortex: AI SQL Functions, Cortex Search and Cortex Analyst

    Spotted a mistake, or was something unclear? Tell us.