What you will be able to do
- Name the categories of Spark actions and give an example of each
- Explain why collect is risky and why production pipelines should limit themselves to write actions
- Distinguish lazy query execution from eager or lazy schema analysis in Spark Classic and Spark Connect
1.What counts as an action, and where its result goes
Every Spark computation starts with an action, so you need to recognise actions on sight. Databricks describes actions as instructing Spark to compute a result from a series of transformations on one or more DataFrames, and groups them into four kinds: output to the console or editor (display, show), collecting data back as Row objects (take(n), first, head), writing to data sources (saveAsTable), and aggregations that trigger a computation (count).
Checkpoint 1 of 5· Match them up
Match each DataFrame method to the category of action it belongs to
Tap a term, then the definition that fits it.
These are the four categories in the Databricks list of actions. All four make Spark evaluate the transformations behind the DataFrame.
“Actions to collect data (returns Row objects), such as take(n), and first or head”Source: docs.databricks.com
Pay particular attention to the collecting actions. collect() returns every record of the DataFrame as a list of Row. That list is built on the driver, so its reference page warns it should only be used if the result is expected to be small, because all the data is loaded into the driver's memory. Look at the second call below. The filter is a transformation that runs on the cluster, and collect brings back only the rows that survived it.
df = spark.createDataFrame([(14, "Tom"), (23, "Alice"), (16, "Bob")], ["age", "name"])
df.collect()
# [Row(age=14, name='Tom'), Row(age=23, name='Alice'), Row(age=16, name='Bob')]
df.filter(df.age > 15).collect()
# [Row(age=23, name='Alice'), Row(age=16, name='Bob')]Checkpoint 2 of 5· Check yourself
A colleague calls collect() on a DataFrame built from a very large table, with no filter. What is the documented risk?
collect is an action that returns every row to the driver as a list, which is only safe when the result is small.
“should only be used if the resulting list is expected to be small, as all the data is loaded into the driver's memory”Source: docs.databricks.com
Checkpoint 3 of 5· Exam question
A developer wants to force Spark to materialize an intermediate DataFrame so its execution time can be measured separately from the rest of the pipeline, using the notebook cell below: ```python cleaned = raw_df.dropDuplicates().filter(col("status") == "active") # ADD A LINE HERE to force execution and print row count final = cleaned.join(other_df, "id") ``` Which line, inserted where indicated, forces Spark to execute the transformations that build `cleaned` at that point in the notebook, independent of the later `join`?
Correct answer: A — print(cleaned.count())
- A. This is correct: `count()` is an action, so it forces Spark to execute `dropDuplicates()` and `filter()` immediately and return a row count the developer can use to measure timing at that point.
- B. This is incorrect because `select("*")` is a transformation. It only reassigns `cleaned` to a new logical plan node and does not trigger any computation on the cluster.
- C. This is incorrect because `explain()` only prints the logical and physical query plan for inspection; it does not execute any of the plan's transformations or touch the underlying data.
- D. This is incorrect because `repartition(8)` is a transformation that adds a shuffle step to the plan. It does not force execution on its own and would not materialize `cleaned` at that point.
- E. This is incorrect because `printSchema()` reads only the DataFrame's known schema metadata, which Spark can resolve without scanning any data, so it does not trigger a job.
2.Each action recomputes: why production code should only write
Since the plan is only a set of instructions, every action can run it again from the source. The RDD guide's way out is persist() (or cache()), which keeps a dataset in memory after the first time it is computed. Storage levels are covered in their own lesson. What matters for the execution model is that, by default, two actions mean two computations.
This is why Databricks treats actions as optimization boundaries. Its optimizations consider all transformations triggered by a given action at once and find the best plan. Any extra action, such as a debugging count(), a preview display() or a manual cache, cuts that plan short. Databricks says: in production data pipelines, writing data is typically the only action that should be present. It also warns that manually caching data or returning preview results in production pipelines can increase cost and latency.
Checkpoint 4 of 5· Check yourself
A production pipeline reads three tables, joins and aggregates them, calls display() twice to sanity-check intermediate results, then writes to a target table. What does Databricks guidance recommend?
Every action other than the final write interrupts query optimization and adds cost. A production pipeline should keep only the write.
“In production data pipelines, writing data is typically the only action that should be present.”Source: docs.databricks.com
3.Lazy execution vs lazy analysis: Spark Classic and Spark Connect
"Lazy" covers two separate things: when Spark executes a query, and when it analyzes it, meaning it resolves column names and checks the plan is valid. Databricks compares the two Spark front ends. Spark Classic is the traditional in-process model. Spark Connect is a gRPC protocol in which a client sends plans to a remote Spark server, and it is used on serverless compute and Databricks Connect. On execution they agree: both follow the same lazy execution model for query execution.
| Operation | Spark Classic | Spark Connect |
|---|---|---|
| Transformations: df.filter(...), df.select(...), df.limit(...) | Lazy execution | Lazy execution |
| SQL queries: spark.sql("select ...") | Lazy execution | Lazy execution |
| Actions: df.collect(), df.show() | Eager execution | Eager execution |
| SQL commands: spark.sql("insert ..."), spark.sql("create ...") | Eager execution | Eager execution |
Analysis is where they differ. Spark Classic performs analysis eagerly during logical plan construction, so a mistake shows up straight away. Its example is spark.sql("select 1 as a, 2 as b").filter("c > 1"), which throws an error at once because column c doesn't exist. In Spark Connect, transformations are built on the client side and sent as unresolved plans to the server. The same line raises nothing. The error only appears on df.columns or df.show(), when the plan reaches the server for analysis. If your code needs to catch analysis errors early, Databricks suggests forcing eager analysis with df.columns, df.schema or df.collect().
Python UDFs are also deferred in Spark Connect. They are serialized only when an action runs, so a UDF picks up variable values from execution time:
from pyspark.sql.functions import udf
x = 123
@udf("INT")
def foo():
return x
df = spark.range(1).select(foo())
x = 456
df.show() # Prints 456Checkpoint 5 of 5· Check yourself
On serverless compute (Spark Connect), you run df = spark.sql("select 1 as a, 2 as b").filter("c > 1"). When is the missing-column error raised?
Spark Connect keeps the unresolved plan on the client. Anything that needs a resolved plan, whether schema access or an action, sends it to the server, and that is when the error appears.
“on df.columns or df.show() an error will be thrown because the unresolved plan is sent to the server for analysis”Source: docs.databricks.com
Sources5
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Anything passed to spark.sql() runs immediately, because SQL is executed when submitted.Why is that wrong?
A spark.sql select query is lazy, just like a DataFrame transformation. Only SQL commands such as insert or create run eagerly.
Covered in Lazy execution vs lazy analysis: Spark Classic and Spark Connect
2.A typo in a column name always raises an error on the line where the transformation is written.Why is that wrong?
That is true in Spark Classic, which analyzes eagerly. Spark Connect defers analysis until something like df.columns, df.schema or an action sends the plan to the server.
Covered in Lazy execution vs lazy analysis: Spark Classic and Spark Connect
3.Extra count() or display() calls in a production pipeline are harmless sanity checks.Why is that wrong?
Each extra action interrupts whole-plan optimization and can add cost and latency. Writing should typically be the only action.
Covered in Each action recomputes: why production code should only write
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/pysparkOfficial docs
“Actions to output data in the console or your editor, such as display or show”
↩︎ What counts as an action, and where its result goes“Actions to write to data sources, such as saveAsTable”
↩︎ What counts as an action, and where its result goes“Aggregations that trigger a computation, such as count”
↩︎ What counts as an action, and where its result goes“All other actions interrupt query optimization and can lead to bottlenecks.”
↩︎ Exam trap 3“Actions to collect data (returns Row objects), such as take(n), and first or head”
↩︎ Checkpoint“In production data pipelines, writing data is typically the only action that should be present.”
↩︎ Checkpoint - 2.
“should only be used if the resulting list is expected to be small, as all the data is loaded into the driver's memory”
↩︎ What counts as an action, and where its result goes - 3.https://spark.apache.org/docs/latest/rdd-programming-guide.htmlSecondary source
“which would cause lineLengths to be saved in memory after the first time it is computed”
↩︎ Each action recomputes: why production code should only write“By default, each transformed RDD may be recomputed each time you run an action on it.”
↩︎ Prediction - 4.https://docs.databricks.com/aws/en/spark/faqOfficial docs
“consider all transformations triggered by a given action at once and find the optimal plan”
↩︎ Each action recomputes: why production code should only write“Manually caching data or returning preview results in production pipelines can interrupt these optimizations and lead to increases in cost and latency.”
↩︎ Each action recomputes: why production code should only write - 5.
“Spark Classic performs analysis eagerly during logical plan construction.”
↩︎ Lazy execution vs lazy analysis: Spark Classic and Spark Connect“Transformations are constructed on the client side and sent as unresolved plans to the server.”
↩︎ Lazy execution vs lazy analysis: Spark Classic and Spark Connect“In Spark Connect, Python UDFs are lazy.”
↩︎ Lazy execution vs lazy analysis: Spark Classic and Spark Connect“Both Spark Classic and Spark Connect follow the same lazy execution model for query execution.”
↩︎ Exam trap 1“you can trigger eager analysis, for example with df.columns, df.schema, or df.collect()”
↩︎ Exam trap 2“on df.columns or df.show() an error will be thrown because the unresolved plan is sent to the server for analysis”
↩︎ Checkpoint