What you will be able to do
- Classify a Spark operation as a transformation or an action
- Trace when Spark actually computes a chain of operations
- Explain why lazy evaluation and immutable DataFrames let Spark optimize a whole query
Key concept
Lazy evaluation — Transformations only record logic in a plan. Spark computes nothing until an action asks for a result, and then it runs the whole recorded chain in one go.
1.Every operation is a transformation or an action
Spark's execution model splits everything you can do with data into two groups. Databricks says it plainly: in Apache Spark, all operations are defined as either transformations or actions. Once you know which group an operation belongs to, you can predict when Spark does any work.
A transformation adds processing logic to the plan. It takes a dataset and describes a new one. Databricks lists reading data, joins, aggregations and type casting as examples. In the original RDD API, map is the standard transformation: it passes each element through a function and returns a new RDD with the results.
An action triggers that processing logic to evaluate and output a result. The RDD programming guide describes actions as operations that return a value to the driver program after running a computation on the dataset. reduce is the standard example. It aggregates all elements with a function and returns the final result to the driver. Databricks' examples of actions include writes, displaying or previewing results, and counting rows.
| Aspect | Transformation | Action |
|---|---|---|
| What it does | Adds processing logic to the plan | Triggers that logic to evaluate and output a result |
| What you get back | A new dataset (RDD or DataFrame) | A value returned to the driver, or data written out |
| Examples in the sources | reading data, joins, aggregations, type casting, map, filter | writes, display/show, count, reduce, collect |
Checkpoint 1 of 6· Check yourself
Which of these operations is an action?
Counting rows has to evaluate the data and return a number, so it is an action. Joins, type casts and map only add logic to the plan.
“Examples include writes, displaying or previewing results, manual caching, or getting the count of rows.”Source: docs.databricks.com
Checkpoint 2 of 6· Exam question
Which statement correctly describes why Apache Spark defers execution of `.filter()`, `.select()`, and other DataFrame transformations until an action such as `.count()` or `.show()` is called?
Correct answer: A — Spark builds a logical plan of the chain of transformations and only launches computation once an action requests a concrete result, letting the optimizer rewrite the plan first.
- A. This is correct: lazy evaluation means transformations only extend a logical plan, and the Catalyst optimizer can rewrite and combine that whole plan before any physical execution happens. Actual work only starts when an action forces a concrete result to be produced.
- B. This is incorrect because Spark does not eagerly run transformations against a sample to validate them. No computation happens at all until an action is called, so there is no early sample-run followed by a full rerun.
- C. This is incorrect because Spark does not make literal in-memory copies of a DataFrame for each transformation. Transformations only add nodes to a logical plan; no data is materialized or duplicated until an action executes it.
- D. This is incorrect because transformations are not submitted as background jobs at all. Only actions cause the driver to submit a job to the cluster; transformations merely extend the plan an action will later execute.
- E. This is incorrect because transformations are not compiled into executor bytecode or cached automatically as they are written. Nothing runs on executors until an action triggers a job for the accumulated plan.
2.Watching laziness happen, line by line
lines = sc.textFile("data.txt")
lineLengths = lines.map(lambda s: len(s))
totalLength = lineLengths.reduce(lambda a, b: a + b)The guide goes through it one line at a time. Line 1 defines a base RDD, but nothing is loaded into memory: lines is merely a pointer to the file. Line 2 defines lineLengths as a map transformation, and it too is not computed yet, because of laziness. Line 3 calls reduce, an action. Only now does Spark break the computation into tasks that run on separate machines. Each machine runs its part of the map plus a local reduction, and returns only its answer to the driver.
DataFrames work the same way. Databricks says that because read is a transformation, Spark does not load data until you call an action such as display. Defining a DataFrame against a file or table costs almost nothing. The cost comes later, at the first action.
df_csv = (spark.read
.format("csv")
.option("header", True)
.option("inferSchema", True)
.load(volume_file_path)
)
display(df_csv)Checkpoint 3 of 6· Fill the gap
Which call on the last line turns this into a program that actually computes something?
lines = sc.textFile("data.txt")
lineLengths = lines.map(lambda s: len(s))
totalLength = lineLengths. ? (lambda a, b: a + b)reduce is the action that returns one total to the driver. map and reduceByKey are both transformations that return a new distributed dataset, so nothing would run.
Source: spark.apache.orgCheckpoint 4 of 6· Put it in order
Put the events of the line-length program in the order they happen
- 1.Spark breaks the computation into tasks on separate machines
- 2.map defines lineLengths without computing it
- 3.reduce is called as an action
- 4.Each machine returns only its local answer to the driver
- 5.textFile defines a base RDD that only points to the file
The two transformations only build up the plan. The action is the point where Spark creates tasks, and those tasks send their partial results back to the driver.
“At this point Spark breaks the computation into tasks to run on separate machines”Source: spark.apache.org
Checkpoint 5 of 6· Exam question
A data engineer runs the following code in a Databricks notebook cell: ```python df = spark.read.parquet("/mnt/sales/transactions") filtered = df.filter(df.amount > 1000) grouped = filtered.groupBy("region").sum("amount") ``` No further code is executed in this cell. What happens when this cell finishes running?
Correct answer: A — No Spark job is submitted to the cluster; Spark only records `filtered` and `grouped` as logical plans built on top of `df`, since every call in the cell is a transformation.
- A. This is correct: `read`, `filter`, and `groupBy`/`sum` are all transformations, so the cell only builds up a logical plan referencing `df`. No job is submitted to the cluster because no action was called.
- B. This is incorrect because no action appears anywhere in the cell, so no job is submitted at all, let alone two separate ones. Filtering and grouping stay purely as plan-building steps until something forces execution.
- C. This is incorrect because being a wide transformation does not make `groupBy` eager. Wide transformations still only add to the logical plan and only cause a shuffle when a later action forces the plan to run.
- D. This is incorrect because leaving a DataFrame unused by an action is not an error condition in Spark; the cell simply finishes with `grouped` holding an unexecuted plan, and no exception is raised.
- E. This is incorrect because Spark does not selectively execute some transformations eagerly based on file metadata. Every transformation in the chain, including the filter, remains unexecuted until an action is called.
3.Why waiting is faster: whole-plan optimization and immutability
Laziness is a deliberate design choice. Because Spark has seen every transformation by the time an action arrives, it can plan the whole chain at once. Databricks puts it this way: rather than evaluating each transformation in the exact order specified, Spark waits until an action triggers computation on all transformations, and picks the most efficient physical plan for that logic. The RDD guide gives a concrete payoff. Spark can see that a mapped dataset only feeds a reduce, so it returns only the result of the reduce to the driver, rather than the larger mapped dataset.
This also changes what a DataFrame is. Lazy evaluation means DataFrames store logical queries as a set of instructions against a data source, not an in-memory result. pandas DataFrames are different: they use eager execution. Spark DataFrames are also immutable. A transformation never changes the DataFrame you called it on. It returns a new one, which you have to assign to a variable to use it later.
from pyspark.sql.functions import count
df_chained = (
df_order.filter(col("o_orderstatus") == "F")
.groupBy(col("o_orderpriority"))
.agg(count(col("o_orderkey")).alias("n_orders"))
.sort(col("n_orders").desc())
)
display(df_chained)Each method in that chain returns a DataFrame. Databricks notes that lazy evaluation is what lets you chain methods for convenience and readability: assigning df_chained runs nothing, and display evaluates the filter, grouping, aggregation and sort together. If you want to see an intermediate step, the documented approach is to call an action on it.
Yes. DataFrames are immutable, so filter returned a new DataFrame that you threw away. df is unchanged. To use the result you have to keep it, for example adults = df.filter(...), and then call adults.count().
Checkpoint 6 of 6· Check yourself
Which statement best describes what a Spark DataFrame holds before any action has run?
Under lazy evaluation a DataFrame is a set of instructions, not data. That is the opposite of pandas' eager model.
“Lazy evaluation means that DataFrames store logical queries as a set of instructions against a data source rather than an in-memory result.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.spark.read...load() loads the file into memory as soon as the line runs.Why is that wrong?
read is a transformation. No data is loaded until an action such as display needs a result.
Covered in Watching laziness happen, line by line
2.Calling a transformation such as filter on a DataFrame changes that DataFrame in place.Why is that wrong?
DataFrames are immutable. Each transformation returns a new DataFrame, and you have to assign it to a variable to use it later.
Covered in Why waiting is faster: whole-plan optimization and immutability
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/spark/faqOfficial docs
“In Apache Spark, all operations are defined as either transformations or actions.”
↩︎ Every operation is a transformation or an action“Transformations: add some processing logic to the plan. Examples include reading data, joins, aggregations, and type casting.”
↩︎ Every operation is a transformation or an action“none of the logic defined by a collection of operations are evaluated until an action is triggered”
↩︎ Key concept“Examples include writes, displaying or previewing results, manual caching, or getting the count of rows.”
↩︎ Checkpoint - 2.https://spark.apache.org/docs/latest/rdd-programming-guide.htmlSecondary source
“actions, which return a value to the driver program after running a computation on the dataset”
↩︎ Every operation is a transformation or an action“although there is also a parallel reduceByKey that returns a distributed dataset”
↩︎ Every operation is a transformation or an action“lines is merely a pointer to the file”
↩︎ Watching laziness happen, line by line“each machine runs both its part of the map and a local reduction, returning only its answer to the driver program”
↩︎ Watching laziness happen, line by line“return only the result of the reduce to the driver, rather than the larger mapped dataset”
↩︎ Why waiting is faster: whole-plan optimization and immutability“The transformations are only computed when an action requires a result to be returned to the driver program.”
↩︎ Prediction“At this point Spark breaks the computation into tasks to run on separate machines”
↩︎ Checkpoint - 3.https://docs.databricks.com/aws/en/pyspark/basicsOfficial docs
“Because read is a transformation, Spark does not load data until you call an action such as display.”
↩︎ Watching laziness happen, line by line“This lazy evaluation means you can chain multiple methods for convenience and readability.”
↩︎ Why waiting is faster: whole-plan optimization and immutability“Because read is a transformation, Spark does not load data until you call an action such as display.”
↩︎ Exam trap 1 - 4.https://docs.databricks.com/aws/en/pysparkOfficial docs
“Rather than evaluating each transformation in the exact order specified, Spark waits until an action triggers computation on all transformations.”
↩︎ Why waiting is faster: whole-plan optimization and immutability“This varies drastically from eager execution, which is the model used by pandas DataFrames.”
↩︎ Why waiting is faster: whole-plan optimization and immutability“If you want to evaluate an intermediate step of your transformation, call an action.”
↩︎ Why waiting is faster: whole-plan optimization and immutability“after performing transformations, a new DataFrame is returned that has to be saved to a variable in order to access it in subsequent operations”
↩︎ Exam trap 2“Lazy evaluation means that DataFrames store logical queries as a set of instructions against a data source rather than an in-memory result.”
↩︎ Checkpoint