What you will be able to do
- Name the three Snowpark languages and explain how Snowpark keeps compute inside Snowflake
- Create a Snowpark Python session from a connections.toml entry, a parameter dictionary, or browser-based SSO
- Explain why Snowpark DataFrames run lazily and what triggers execution
- Install snowflake-ml-python and connect it to Snowflake through a Snowpark Session, including picking the warehouse
Key concept
Pushdown: process the data where it lives — Snowpark builds SQL on the client and runs it inside Snowflake. The data stays in Snowflake, and only the results you ask for come back to your machine, so the client can be far smaller than the data.
1.Snowpark: one API in three languages, with compute inside Snowflake
Snowpark is the Snowflake library for querying and processing data at scale. The idea that matters for data science is where the work runs. Your code builds the operations, Snowflake executes them, and the data never has to be copied to the machine that runs your script. This is why a laptop with modest memory can prepare training data from a very large table. The laptop only describes the work. The warehouse does it.
Snowflake publishes Snowpark libraries for three languages: Java, Python and Scala. Each one has its own developer guide and API reference. On an account where some people write Python and others write Java, both groups use the same Snowpark model. They simply use different client libraries.
| Aspect | Snowpark |
|---|---|
| Languages | Java, Python, Scala |
| Where you write code | Local tools such as Jupyter, VS Code, or IntelliJ |
| Where transformations run | Pushed down to Snowflake, including Snowflake UDFs |
| Extra compute cluster | None; all computations are done within Snowflake |
You can also write Snowpark Python in a Python worksheet in Snowsight instead of a local environment. The docs point to Snowsight worksheets when you want to write a stored procedure that automates tasks inside Snowflake. They point to a local environment when you are building a client application.
Checkpoint 1 of 5· Check yourself
A platform team supports Python data scientists and Java developers on one Snowflake account. Which statement about Snowpark is accurate?
Snowflake ships Snowpark libraries for Java, Python and Scala, and all of them run their computation inside Snowflake.
“Snowflake currently provides Snowpark libraries for three languages: Java, Python, and Scala.”Source: docs.snowflake.com
2.Opening a Snowpark Python session
Every Snowpark Python program starts with a Session, imported with from snowflake.snowpark import Session. Snowpark has no separate login system of its own. It authenticates with the same mechanisms and the same parameters (account, user, role, warehouse, database, schema) as the connect function of the Snowflake Connector for Python.
There are two common ways to supply those parameters.
1. **A named connection in connections.toml.** You add a block such as [myconnection] with account, user, password, warehouse, database and schema. Then you refer to it by name: session = Session.builder.config("connection_name", "myconnection").create(). Keeping credentials in this file keeps them out of your notebooks.
2. A Python dictionary that you pass to the builder, as shown below.
connection_parameters = {
"account": "<your snowflake account>",
"user": "<your snowflake user>",
"password": "<your snowflake password>",
"role": "<your snowflake role>", # optional
"warehouse": "<your snowflake warehouse>", # optional
"database": "<your snowflake database>", # optional
"schema": "<your snowflake schema>", # optional
}
new_session = Session.builder.configs(connection_parameters).create()Two details come up often. First, account takes your account identifier without the snowflakecomputing.com suffix. Second, if your organisation uses single sign-on, add "authenticator": "externalbrowser" to the dictionary and drop the password. The docs also list MFA, key pair authentication, proxies and OAuth as other ways to connect. When you are finished, new_session.close() closes the session and cancels any queries that are still running.
Checkpoint 2 of 5· Put it in order
Put the steps for creating a Snowpark session from a dictionary in order
- 1.Pass the dict to Session.builder.configs to get a builder with those parameters
- 2.Call the builder's create method to establish the session
- 3.Create a Python dict with the names and values of the connection parameters
The docs give exactly this sequence: build the dict, give it to the builder through configs, then create.
“Call the create method of the builder to establish the session.”Source: docs.snowflake.com
Sources3
3.Lazy DataFrames and UDFs: how the work stays on the server
The core Snowpark abstraction is the DataFrame. You use it to describe the data you want: which columns, which filters, which joins. Snowpark records these as language constructs such as select and col, rather than as SQL strings, so your editor can autocomplete and type-check them. Execution is lazy. Nothing runs until you call an action such as collect(). At that point Snowpark sends the combined SQL to Snowflake. Batching many operations into one query this way reduces how much data moves between your client and Snowflake.
>>> # Create a DataFrame with the "id" and "name" columns from the "sample_product_data" table.
>>> # This does not execute the query.
>>> df = session.table("sample_product_data").select(col("id"), col("name"))
>>> # Send the query to the server for execution and
>>> # return a list of Rows containing the results.
>>> results = df.collect()Sometimes the logic is custom, like a per-row Python function. In that case you wrap it as a user-defined function (UDF), for example with a Python lambda. Snowpark pushes the UDF code to Snowflake, and Snowflake runs it in parallel next to the data. You don't download the table to apply the function. Snowpark Python also supports UDTFs and stored procedures, and you can schedule stored procedures as tasks.
Checkpoint 3 of 5· Exam question
A data scientist working in Visual Studio Code on a laptop must engineer features from a 2-billion-row table without moving the data out of Snowflake. Which approach is MOST appropriate?
Correct answer: C — Build a Snowpark `Session` from connection parameters and chain DataFrame transformations on `session.table`, so the SQL runs on a warehouse.
- A. `to_pandas()` materialises the whole table on the client, so the features would be computed locally rather than pushed down to the warehouse.
- B. Pulling all 2 billion rows to the client moves the data out of Snowflake and will exhaust laptop memory long before features are computed.
- C. Snowpark DataFrame operations are compiled to SQL and executed on the warehouse, so only small results ever reach the laptop.
- D. Unloading and downloading copies the data out of Snowflake, which is the opposite of connecting the tool directly to the data.
Checkpoint 4 of 5· Check yourself
Which statement about a Python UDF created inline in a Snowpark app is accurate?
Snowpark sends the UDF code to the Snowflake engine, so the function runs on the server without moving data to the client.
“When you call the UDF in your client code, your custom code is executed on the server (where the data is).”Source: docs.snowflake.com
Sources1
4.Snowpark ML: connecting snowflake-ml-python through a session
The exam guide says "Snowpark ML". The current docs call it Snowflake ML and ship it as one package, snowflake-ml-python. This package holds the Python APIs for the Snowflake ML workflow. You can use it in a Python IDE on your own workstation, in Snowsight worksheets, or in Snowflake Notebooks. Notebooks offer both CPU and GPU runtimes.
Where you install it depends on where you work:
- In Snowsight worksheets or notebooks: choose the Anaconda package snowflake-ml-python from the Packages menu. Your organisation's package policy must allow it.
- On your own workstation: install it yourself. The docs prefer conda, using the Snowflake conda channel. If you use pip, use a virtualenv, and don't run pip inside a conda environment.
Snowpark Python is a dependency of the package and installs automatically with it. scikit-learn installs by default. Libraries such as lightgbm and xgboost are optional extras.
Snowflake ML connects through a Snowpark Session. The helper SnowflakeLoginOptions in snowflake.ml.utils.connection_params reads a named connection from your SnowSQL configuration file, or reads environment variables. It returns a dictionary that you pass to Session.builder.configs. If you already have a Snowflake Connector for Python connection, you can wrap it instead: Session.builder.configs({"connection": connection}).create().
from snowflake.snowpark import Session
from snowflake.ml.utils import connection_params
params = connection_params.SnowflakeLoginOptions("myaccount")
sp_session = Session.builder.configs(params).create()The session decides where the compute runs. Training and inference run in the warehouse the session specifies. If the configuration has no warehouse, or you want a different one, add a warehouse key to the params or switch the warehouse on the live session.
Checkpoint 5 of 5· Fill the gap
Which session method sets the warehouse that Snowflake ML operations will use?
sp_session. ? ("mlwarehouse")The docs say to call the session's use_warehouse method when no warehouse is set, or when you want a different one.
Source: docs.snowflake.comSources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Creating a Snowpark DataFrame with session.table(...).select(...) runs the query and loads the rows into client memory.Why is that wrong?
Building a DataFrame only describes the query. Nothing runs until an action such as collect() sends the SQL to Snowflake.
Covered in Lazy DataFrames and UDFs: how the work stays on the server
2.Snowpark, like Spark-based tools, needs its own compute cluster next to Snowflake.Why is that wrong?
Snowpark pushes all operations down to Snowflake. Snowflake handles scale and compute, and no outside cluster is needed.
Covered in Snowpark: one API in three languages, with compute inside Snowflake
3.After installing snowflake-ml-python, you still have to install the Snowpark Python package separately.Why is that wrong?
Snowpark Python is a dependency of snowflake-ml-python and installs with it. Some extra environment setup may still be needed.
Covered in Snowpark ML: connecting snowflake-ml-python through a session
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Support for authoring Snowpark code using local tools such as Jupyter, VS Code, or IntelliJ.”
↩︎ Snowpark: one API in three languages, with compute inside Snowflake“This reduces the amount of data transferred between your client and the Snowflake database.”
↩︎ Lazy DataFrames and UDFs: how the work stays on the server“Snowpark automatically pushes the custom code for UDFs to the Snowflake engine.”
↩︎ Lazy DataFrames and UDFs: how the work stays on the server“you can build applications that process data in Snowflake without moving data to the system where your application code runs”
↩︎ Key concept“Snowpark operations are executed lazily on the server”
↩︎ Exam trap 1“No requirement for a separate cluster outside of Snowflake for computations.”
↩︎ Exam trap 2“Snowflake currently provides Snowpark libraries for three languages: Java, Python, and Scala.”
↩︎ Checkpoint“The data isn’t retrieved when you construct the DataFrame object.”
↩︎ Prediction“When you call the UDF in your client code, your custom code is executed on the server (where the data is).”
↩︎ Checkpoint - 2.
“You can write Snowpark Python code in a local development environment or in a Python worksheet in Snowsight.”
↩︎ Snowpark: one API in three languages, with compute inside Snowflake - 3.
“To authenticate, you use the same mechanisms that the Snowflake Connector for Python supports.”
↩︎ Opening a Snowpark Python session“Note that the account identifier does not include the snowflakecomputing.com suffix.”
↩︎ Opening a Snowpark Python session“Set the authenticator option to externalbrowser.”
↩︎ Opening a Snowpark Python session“Call the create method of the builder to establish the session.”
↩︎ Checkpoint - 4.
“All Snowflake ML features are available in a single package, snowflake-ml-python.”
↩︎ Snowpark ML: connecting snowflake-ml-python through a session“These operations run in the warehouse specified by the session you use to connect.”
↩︎ Snowpark ML: connecting snowflake-ml-python through a session“Notebooks support both CPU and GPU runtime options.”
↩︎ Snowpark ML: connecting snowflake-ml-python through a session“Snowpark Python is a dependency of snowflake-ml-python and is installed automatically with it.”
↩︎ Exam trap 3