CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 1 · Lesson 6/32

    Spark Transformations, Actions and Lazy Evaluation

    Describe the execution patterns of the Apache SparkTM engine, including actions, transformations, and lazy evaluation.

    10 min read
    3.12% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Classify a Spark operation as a transformation or an action
    • Trace when Spark actually computes a chain of operations
    • Explain why lazy evaluation and immutable DataFrames let Spark optimize a whole query

    Key concept

    Lazy evaluation — Transformations only record logic in a plan. Spark computes nothing until an action asks for a result, and then it runs the whole recorded chain in one go.

    1.Every operation is a transformation or an action

    Spark's execution model splits everything you can do with data into two groups. Databricks says it plainly: in Apache Spark, all operations are defined as either transformations or actions. Once you know which group an operation belongs to, you can predict when Spark does any work.

    A transformation adds processing logic to the plan. It takes a dataset and describes a new one. Databricks lists reading data, joins, aggregations and type casting as examples. In the original RDD API, map is the standard transformation: it passes each element through a function and returns a new RDD with the results.

    An action triggers that processing logic to evaluate and output a result. The RDD programming guide describes actions as operations that return a value to the driver program after running a computation on the dataset. reduce is the standard example. It aggregates all elements with a function and returns the final result to the driver. Databricks' examples of actions include writes, displaying or previewing results, and counting rows.

    The two operation types side by side
    AspectTransformationAction
    What it doesAdds processing logic to the planTriggers that logic to evaluate and output a result
    What you get backA new dataset (RDD or DataFrame)A value returned to the driver, or data written out
    Examples in the sourcesreading data, joins, aggregations, type casting, map, filterwrites, display/show, count, reduce, collect

    Checkpoint 1 of 6· Check yourself

    Which of these operations is an action?

    Checkpoint 2 of 6· Exam question

    Which statement correctly describes why Apache Spark defers execution of `.filter()`, `.select()`, and other DataFrame transformations until an action such as `.count()` or `.show()` is called?

    Sources12

    2.Watching laziness happen, line by line

    The RDD guide's example: two transformations followed by one actionpython
    lines = sc.textFile("data.txt")
    lineLengths = lines.map(lambda s: len(s))
    totalLength = lineLengths.reduce(lambda a, b: a + b)

    The guide goes through it one line at a time. Line 1 defines a base RDD, but nothing is loaded into memory: lines is merely a pointer to the file. Line 2 defines lineLengths as a map transformation, and it too is not computed yet, because of laziness. Line 3 calls reduce, an action. Only now does Spark break the computation into tasks that run on separate machines. Each machine runs its part of the map plus a local reduction, and returns only its answer to the driver.

    DataFrames work the same way. Databricks says that because read is a transformation, Spark does not load data until you call an action such as display. Defining a DataFrame against a file or table costs almost nothing. The cost comes later, at the first action.

    DataFrameReader is a transformation; display is the action that loads the CSVpython
    df_csv = (spark.read
      .format("csv")
      .option("header", True)
      .option("inferSchema", True)
      .load(volume_file_path)
    )
    display(df_csv)

    Checkpoint 3 of 6· Fill the gap

    Which call on the last line turns this into a program that actually computes something?

    lines = sc.textFile("data.txt")
    lineLengths = lines.map(lambda s: len(s))
    totalLength = lineLengths. ? (lambda a, b: a + b)

    Checkpoint 4 of 6· Put it in order

    Put the events of the line-length program in the order they happen

    1. 1.Spark breaks the computation into tasks on separate machines
    2. 2.map defines lineLengths without computing it
    3. 3.reduce is called as an action
    4. 4.Each machine returns only its local answer to the driver
    5. 5.textFile defines a base RDD that only points to the file

    Checkpoint 5 of 6· Exam question

    A data engineer runs the following code in a Databricks notebook cell: ```python df = spark.read.parquet("/mnt/sales/transactions") filtered = df.filter(df.amount > 1000) grouped = filtered.groupBy("region").sum("amount") ``` No further code is executed in this cell. What happens when this cell finishes running?

    Sources23

    3.Why waiting is faster: whole-plan optimization and immutability

    Laziness is a deliberate design choice. Because Spark has seen every transformation by the time an action arrives, it can plan the whole chain at once. Databricks puts it this way: rather than evaluating each transformation in the exact order specified, Spark waits until an action triggers computation on all transformations, and picks the most efficient physical plan for that logic. The RDD guide gives a concrete payoff. Spark can see that a mapped dataset only feeds a reduce, so it returns only the result of the reduce to the driver, rather than the larger mapped dataset.

    This also changes what a DataFrame is. Lazy evaluation means DataFrames store logical queries as a set of instructions against a data source, not an in-memory result. pandas DataFrames are different: they use eager execution. Spark DataFrames are also immutable. A transformation never changes the DataFrame you called it on. It returns a new one, which you have to assign to a variable to use it later.

    Transformations chained into one plan; display is the only actionpython
    from pyspark.sql.functions import count
    
    df_chained = (
        df_order.filter(col("o_orderstatus") == "F")
        .groupBy(col("o_orderpriority"))
        .agg(count(col("o_orderkey")).alias("n_orders"))
        .sort(col("n_orders").desc())
    )
    
    display(df_chained)

    Each method in that chain returns a DataFrame. Databricks notes that lazy evaluation is what lets you chain methods for convenience and readability: assigning df_chained runs nothing, and display evaluates the filter, grouping, aggregation and sort together. If you want to see an intermediate step, the documented approach is to call an action on it.

    Checkpoint 6 of 6· Check yourself

    Which statement best describes what a Spark DataFrame holds before any action has run?

    Sources423

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.spark.read...load() loads the file into memory as soon as the line runs.Why is that wrong?

      read is a transformation. No data is loaded until an action such as display needs a result.

      Covered in Watching laziness happen, line by line

    2. 2.Calling a transformation such as filter on a DataFrame changes that DataFrame in place.Why is that wrong?

      DataFrames are immutable. Each transformation returns a new DataFrame, and you have to assign it to a variable to use it later.

      Covered in Why waiting is faster: whole-plan optimization and immutability

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “In Apache Spark, all operations are defined as either transformations or actions.”
      ↩︎ Every operation is a transformation or an action
      “Transformations: add some processing logic to the plan. Examples include reading data, joins, aggregations, and type casting.”
      ↩︎ Every operation is a transformation or an action
      “none of the logic defined by a collection of operations are evaluated until an action is triggered”
      ↩︎ Key concept
      “Examples include writes, displaying or previewing results, manual caching, or getting the count of rows.”
      ↩︎ Checkpoint
    2. 2.
      “actions, which return a value to the driver program after running a computation on the dataset”
      ↩︎ Every operation is a transformation or an action
      “although there is also a parallel reduceByKey that returns a distributed dataset”
      ↩︎ Every operation is a transformation or an action
      “lines is merely a pointer to the file”
      ↩︎ Watching laziness happen, line by line
      “each machine runs both its part of the map and a local reduction, returning only its answer to the driver program”
      ↩︎ Watching laziness happen, line by line
      “return only the result of the reduce to the driver, rather than the larger mapped dataset”
      ↩︎ Why waiting is faster: whole-plan optimization and immutability
      “The transformations are only computed when an action requires a result to be returned to the driver program.”
      ↩︎ Prediction
      “At this point Spark breaks the computation into tasks to run on separate machines”
      ↩︎ Checkpoint
    3. 3.
      “Because read is a transformation, Spark does not load data until you call an action such as display.”
      ↩︎ Watching laziness happen, line by line
      “This lazy evaluation means you can chain multiple methods for convenience and readability.”
      ↩︎ Why waiting is faster: whole-plan optimization and immutability
      “Because read is a transformation, Spark does not load data until you call an action such as display.”
      ↩︎ Exam trap 1
    4. 4.
      “Rather than evaluating each transformation in the exact order specified, Spark waits until an action triggers computation on all transformations.”
      ↩︎ Why waiting is faster: whole-plan optimization and immutability
      “This varies drastically from eager execution, which is the model used by pandas DataFrames.”
      ↩︎ Why waiting is faster: whole-plan optimization and immutability
      “If you want to evaluate an intermediate step of your transformation, call an action.”
      ↩︎ Why waiting is faster: whole-plan optimization and immutability
      “after performing transformations, a new DataFrame is returned that has to be saved to a variable in order to access it in subsequent operations”
      ↩︎ Exam trap 2
      “Lazy evaluation means that DataFrames store logical queries as a set of instructions against a data source rather than an in-memory result.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Spark Actions in Practice: collect, Recomputation and Spark Connect

    Spotted a mistake, or was something unclear? Tell us.