What you will be able to do
- Explain where Snowpark code runs and why Snowpark needs no separate compute cluster
- Describe the construct → transform → action lifecycle of a lazily evaluated DataFrame
- Choose the right Session method to build a DataFrame from a table, local values, a range, a staged file or a SQL query
Key concept
Lazily evaluated DataFrame — A Snowpark DataFrame is a description of a query, not a container of rows. Building it and transforming it sends nothing to Snowflake. SQL runs only when you call an action method.
1.Snowpark architecture: your code, Snowflake's compute
Snowpark is a client library for querying and processing data in Snowflake. Snowflake provides it in exactly three languages: Java, Python and Scala. Its main architectural idea is that the data never moves to your code. You write ordinary Java, Python or Scala on your laptop, in Jupyter, VS Code or IntelliJ. Snowpark turns that code into SQL, and the SQL runs inside Snowflake's engine. The documentation describes this as building applications "without moving data to the system where your application code runs".
The documentation compares Snowpark with the Snowflake Connector for Spark, and two benefits stand out for the exam. First, Snowpark offers pushdown for every operation, including Snowflake UDFs, so the heavy lifting happens where the data is stored. Second, Snowpark needs no separate cluster of its own. Snowflake does all the computation and also manages scaling and compute. So when a question asks where a Snowpark transformation executes, or what extra infrastructure you must provision, the answer is Snowflake itself and nothing extra.
| Aspect | How Snowpark handles it |
|---|---|
| Languages | Java, Python and Scala libraries |
| Where computation runs | Within Snowflake; no separate cluster outside Snowflake |
| Pushdown | Supported for all operations, including Snowflake UDFs |
| Authoring | Local tools such as Jupyter, VS Code or IntelliJ |
| Building queries | Language constructs such as a select method instead of a 'select column_name' string |
The last row matters in day-to-day work. Snowpark gives you programming-language constructs for building SQL statements. Instead of putting a SELECT inside a string, you call methods and pass column objects. You can still run a string of SQL if you want to, but the native constructs give you code completion and type checking. Here is the Python version of a three-column projection:
>>> # Import the col function from the functions module.
>>> from snowflake.snowpark.functions import col
>>> # Create a DataFrame that contains the id, name, and serial_number
>>> # columns in the "sample_product_data" table.
>>> df = session.table("sample_product_data").select(col("id"), col("name"), col("serial_number"))
>>> df.show()Checkpoint 1 of 5· Check yourself
A data engineer runs a Snowpark Python pipeline from a laptop against a 2 TB table. Where do the DataFrame transformations actually run?
Snowpark pushes operations down to Snowflake and needs no external cluster. Snowflake does the computation and manages scaling and compute.
“All of the computations are done within Snowflake. Scale and compute management are handled by Snowflake.”Source: docs.snowflake.com
Sources1
2.Lazy evaluation: construct, transform, then act
Since the work happens on the server, Snowpark has to decide when to send it there. That raises a question about the line of code below.
The DataFrame is Snowpark's core abstraction. The Python, Java and Scala guides all describe the same three-step lifecycle. First, construct a DataFrame and name its source, such as a table, a staged file or a SQL statement. Second, describe the transformations: which columns to select, how to filter rows, and how to sort and group the results. Third, call a method that performs an action, such as collect(). Only the third step contacts the database.
The documentation gives the reason for this design. Snowpark operations run lazily on the server, so a pipeline can delay execution as long as possible while "batching up many operations into a single operation. This reduces the amount of data transferred between your client and the Snowflake database." One way to picture it: a DataFrame is like a query waiting to be evaluated.
Checkpoint 2 of 5· Fill the gap
Which method sends this DataFrame's query to the server and returns a list of Rows?
>>> # Send the query to the server for execution and
>>> # return a list of Rows containing the results.
>>> results = df. ? ()collect() is an action. It evaluates the DataFrame and returns the results. select, filter and table only describe or build the query.
Source: docs.snowflake.comCheckpoint 3 of 5· Put it in order
Put the steps for retrieving data into a Snowpark DataFrame in order
- 1.Specify how the dataset should be transformed (columns, filters, sorting, grouping)
- 2.Invoke an action method, such as collect(), to execute the statement
- 3.Construct a DataFrame, specifying the source of the data
The source is defined first and the transformations are layered on top. The action comes last because it is what makes Snowflake run the SQL.
“you must invoke a method that performs an action (for example, the collect() method).”Source: docs.snowflake.com
Checkpoint 4 of 5· Exam question
A data engineer writes a Python script that builds a Snowpark DataFrame by chaining `.filter()`, `.select()`, and `.with_column()` calls against a table. When the engineer checks the Snowflake query history immediately after the chain, no query has run yet. A query only appears once the engineer calls `.collect()` at the end of the script. What explains this behavior?
Correct answer: A — Transformation methods such as `filter`, `select`, and `with_column` only extend the DataFrame's underlying query plan, and Snowflake compiles and runs the SQL only when an action method like `collect` executes.
- A. Snowpark DataFrames are lazily evaluated: transformation methods only describe how to build the resulting SQL statement and do not touch the database. The statement is compiled and pushed down to Snowflake only when an action method such as `collect`, `show`, or `count` is called.
- B. Snowpark does not batch DataFrame operations on a timer. There is no local buffering window that periodically synchronizes transformations with Snowflake; execution timing is governed entirely by lazy evaluation, not a clock.
- C. DataFrame transformation methods do not need a separate `session.sql()` call to be recognized. They are native DataFrame API calls that Snowpark translates directly into the query plan without any manual SQL registration step.
- D. Warehouse suspension would cause an action's query to wait or auto-resume the warehouse, not cause transformation calls to queue silently. The missing query history entries are explained by lazy evaluation, not warehouse state.
3.Constructing a DataFrame from a Session
Every DataFrame starts from a Session, your connection to Snowflake. In Python, one documented approach is to put the connection parameters (account, user, role, warehouse, database, schema) in a dict and call Session.builder.configs(connection_parameters).create(). After that, the Session's methods and properties build DataFrames, and each one reads from a different kind of source.
| Call | Data source | Example from the docs |
|---|---|---|
| session.table(...) | A table, view, or stream | session.table("sample_product_data") |
| session.create_dataframe(...) | Values you specify | session.create_dataframe([1, 2, 3, 4]).to_df("a") |
| session.range(...) | A range of values | session.range(1, 10, 2).to_df("a") gives 1, 3, 5, 7, 9 |
| session.read.<format>(...) | A file in a stage, via a DataFrameReader | session.read.json("@my_stage2/data1.json") |
| session.sql(...) | The results of a SQL query | session.sql("SELECT name from sample_product_data") |
Look closely at the read row. read is a property that returns a DataFrameReader, and you then call the method for the file's format, such as json or csv. You can also attach a schema first, built from StructType and StructField. You can likewise pass a StructType schema to create_dataframe so that local values get explicit column types.
session.sql can run SELECT statements against tables and staged files, but the documentation recommends something else. It says that "using the table method and read property offer better syntax highlighting, error highlighting, and intelligent code completion in development tools." The Java and Scala guides give the same advice. Two more details from those guides: in Java and Scala, the table method returns an Updatable object, which extends DataFrame with methods for updating and deleting table data. And in every language, words reserved by Snowflake are not valid as column names when you construct a DataFrame.
Checkpoint 5 of 5· Match them up
Match each Session call to the source it builds a DataFrame from
Tap a term, then the definition that fits it.
Each constructor matches one kind of source. table and read are preferred over sql for tables and staged files because they give better tooling support.
“To create a DataFrame from data in a table, view, or stream, call the table method”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Snowpark, like the Spark connector, needs its own compute cluster next to Snowflake for heavy transformations.Why is that wrong?
Snowpark pushes all operations down to Snowflake. No external cluster is involved, and Snowflake manages scale and compute.
Covered in Snowpark architecture: your code, Snowflake's compute
2.Calling session.table(...).select(...) runs a query and pulls the selected columns to the client.Why is that wrong?
Constructing and transforming a DataFrame only describes the query. Snowflake runs SQL only when an action such as collect() is called.
3.session.sql("SELECT ...") is the recommended way to load a table or staged file into a DataFrame.Why is that wrong?
sql works, but the docs recommend the table method and read property because they give better syntax highlighting, error highlighting and code completion.
Covered in Constructing a DataFrame from a Session
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Snowflake currently provides Snowpark libraries for three languages: Java, Python, and Scala.”
↩︎ Snowpark architecture: your code, Snowflake's compute“Support for pushdown for all operations, including Snowflake UDFs.”
↩︎ Snowpark architecture: your code, Snowflake's compute“The Snowpark API provides programming language constructs for building SQL statements.”
↩︎ Snowpark architecture: your code, Snowflake's compute“batching up many operations into a single operation. This reduces the amount of data transferred between your client and the Snowflake database.”
↩︎ Lazy evaluation: construct, transform, then act“No requirement for a separate cluster outside of Snowflake for computations.”
↩︎ Exam trap 1“All of the computations are done within Snowflake. Scale and compute management are handled by Snowflake.”
↩︎ Checkpoint“This does not execute the query.”
↩︎ Prediction - 2.
“In a sense, a DataFrame is like a query that needs to be evaluated in order to retrieve data.”
↩︎ Lazy evaluation: construct, transform, then act“use the read property to get a DataFrameReader object”
↩︎ Constructing a DataFrame from a Session“A DataFrame represents a relational dataset that is evaluated lazily: it only executes when a specific action is triggered.”
↩︎ Key concept“In order to retrieve the data into the DataFrame, you must invoke a method that performs an action”
↩︎ Exam trap 2“using the table method and read property offer better syntax highlighting, error highlighting, and intelligent code completion in development tools.”
↩︎ Exam trap 3“you must invoke a method that performs an action (for example, the collect() method).”
↩︎ Checkpoint“To create a DataFrame from data in a table, view, or stream, call the table method”
↩︎ Checkpoint - 3.
“new_session = Session.builder.configs(connection_parameters).create()”
↩︎ Constructing a DataFrame from a Session - 4.
“The table method returns an Updatable object.”
↩︎ Constructing a DataFrame from a Session“Words reserved by Snowflake are not valid as column names when constructing a DataFrame.”
↩︎ Constructing a DataFrame from a Session