CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 1 · Lesson 6/32

    Spark Actions in Practice: collect, Recomputation and Spark Connect

    Describe the execution patterns of the Apache SparkTM engine, including actions, transformations, and lazy evaluation.

    8 min read
    3.12% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Name the categories of Spark actions and give an example of each
    • Explain why collect is risky and why production pipelines should limit themselves to write actions
    • Distinguish lazy query execution from eager or lazy schema analysis in Spark Classic and Spark Connect

    1.What counts as an action, and where its result goes

    Every Spark computation starts with an action, so you need to recognise actions on sight. Databricks describes actions as instructing Spark to compute a result from a series of transformations on one or more DataFrames, and groups them into four kinds: output to the console or editor (display, show), collecting data back as Row objects (take(n), first, head), writing to data sources (saveAsTable), and aggregations that trigger a computation (count).

    Checkpoint 1 of 5· Match them up

    Match each DataFrame method to the category of action it belongs to

    Tap a term, then the definition that fits it.

    Pay particular attention to the collecting actions. collect() returns every record of the DataFrame as a list of Row. That list is built on the driver, so its reference page warns it should only be used if the result is expected to be small, because all the data is loaded into the driver's memory. Look at the second call below. The filter is a transformation that runs on the cluster, and collect brings back only the rows that survived it.

    collect() as the action at the end of a filterpython
    df = spark.createDataFrame([(14, "Tom"), (23, "Alice"), (16, "Bob")], ["age", "name"])
    df.collect()
    # [Row(age=14, name='Tom'), Row(age=23, name='Alice'), Row(age=16, name='Bob')]
    
    df.filter(df.age > 15).collect()
    # [Row(age=23, name='Alice'), Row(age=16, name='Bob')]

    Checkpoint 2 of 5· Check yourself

    A colleague calls collect() on a DataFrame built from a very large table, with no filter. What is the documented risk?

    Checkpoint 3 of 5· Exam question

    A developer wants to force Spark to materialize an intermediate DataFrame so its execution time can be measured separately from the rest of the pipeline, using the notebook cell below: ```python cleaned = raw_df.dropDuplicates().filter(col("status") == "active") # ADD A LINE HERE to force execution and print row count final = cleaned.join(other_df, "id") ``` Which line, inserted where indicated, forces Spark to execute the transformations that build `cleaned` at that point in the notebook, independent of the later `join`?

    Sources12

    2.Each action recomputes: why production code should only write

    Since the plan is only a set of instructions, every action can run it again from the source. The RDD guide's way out is persist() (or cache()), which keeps a dataset in memory after the first time it is computed. Storage levels are covered in their own lesson. What matters for the execution model is that, by default, two actions mean two computations.

    This is why Databricks treats actions as optimization boundaries. Its optimizations consider all transformations triggered by a given action at once and find the best plan. Any extra action, such as a debugging count(), a preview display() or a manual cache, cuts that plan short. Databricks says: in production data pipelines, writing data is typically the only action that should be present. It also warns that manually caching data or returning preview results in production pipelines can increase cost and latency.

    Checkpoint 4 of 5· Check yourself

    A production pipeline reads three tables, joins and aggregates them, calls display() twice to sanity-check intermediate results, then writes to a target table. What does Databricks guidance recommend?

    Sources34

    3.Lazy execution vs lazy analysis: Spark Classic and Spark Connect

    "Lazy" covers two separate things: when Spark executes a query, and when it analyzes it, meaning it resolves column names and checks the plan is valid. Databricks compares the two Spark front ends. Spark Classic is the traditional in-process model. Spark Connect is a gRPC protocol in which a client sends plans to a remote Spark server, and it is used on serverless compute and Databricks Connect. On execution they agree: both follow the same lazy execution model for query execution.

    When each kind of operation executes (Databricks comparison)
    OperationSpark ClassicSpark Connect
    Transformations: df.filter(...), df.select(...), df.limit(...)Lazy executionLazy execution
    SQL queries: spark.sql("select ...")Lazy executionLazy execution
    Actions: df.collect(), df.show()Eager executionEager execution
    SQL commands: spark.sql("insert ..."), spark.sql("create ...")Eager executionEager execution

    Analysis is where they differ. Spark Classic performs analysis eagerly during logical plan construction, so a mistake shows up straight away. Its example is spark.sql("select 1 as a, 2 as b").filter("c > 1"), which throws an error at once because column c doesn't exist. In Spark Connect, transformations are built on the client side and sent as unresolved plans to the server. The same line raises nothing. The error only appears on df.columns or df.show(), when the plan reaches the server for analysis. If your code needs to catch analysis errors early, Databricks suggests forcing eager analysis with df.columns, df.schema or df.collect().

    Python UDFs are also deferred in Spark Connect. They are serialized only when an action runs, so a UDF picks up variable values from execution time:

    In Spark Connect the UDF is serialized at show(), so it sees x = 456python
    from pyspark.sql.functions import udf
    
    x = 123
    
    @udf("INT")
    def foo():
      return x
    
    
    df = spark.range(1).select(foo())
    x = 456
    df.show() # Prints 456

    Checkpoint 5 of 5· Check yourself

    On serverless compute (Spark Connect), you run df = spark.sql("select 1 as a, 2 as b").filter("c > 1"). When is the missing-column error raised?

    Sources5

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Anything passed to spark.sql() runs immediately, because SQL is executed when submitted.Why is that wrong?

      A spark.sql select query is lazy, just like a DataFrame transformation. Only SQL commands such as insert or create run eagerly.

      Covered in Lazy execution vs lazy analysis: Spark Classic and Spark Connect

    2. 2.A typo in a column name always raises an error on the line where the transformation is written.Why is that wrong?

      That is true in Spark Classic, which analyzes eagerly. Spark Connect defers analysis until something like df.columns, df.schema or an action sends the plan to the server.

      Covered in Lazy execution vs lazy analysis: Spark Classic and Spark Connect

    3. 3.Extra count() or display() calls in a production pipeline are harmless sanity checks.Why is that wrong?

      Each extra action interrupts whole-plan optimization and can add cost and latency. Writing should typically be the only action.

      Covered in Each action recomputes: why production code should only write

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Actions to output data in the console or your editor, such as display or show”
      ↩︎ What counts as an action, and where its result goes
      “Actions to write to data sources, such as saveAsTable”
      ↩︎ What counts as an action, and where its result goes
      “Aggregations that trigger a computation, such as count”
      ↩︎ What counts as an action, and where its result goes
      “All other actions interrupt query optimization and can lead to bottlenecks.”
      ↩︎ Exam trap 3
      “Actions to collect data (returns Row objects), such as take(n), and first or head”
      ↩︎ Checkpoint
      “In production data pipelines, writing data is typically the only action that should be present.”
      ↩︎ Checkpoint
    2. 2.
      “should only be used if the resulting list is expected to be small, as all the data is loaded into the driver's memory”
      ↩︎ What counts as an action, and where its result goes
    3. 3.
      “which would cause lineLengths to be saved in memory after the first time it is computed”
      ↩︎ Each action recomputes: why production code should only write
      “By default, each transformed RDD may be recomputed each time you run an action on it.”
      ↩︎ Prediction
    4. 4.
      “consider all transformations triggered by a given action at once and find the optimal plan”
      ↩︎ Each action recomputes: why production code should only write
      “Manually caching data or returning preview results in production pipelines can interrupt these optimizations and lead to increases in cost and latency.”
      ↩︎ Each action recomputes: why production code should only write
    5. 5.
      “Spark Classic performs analysis eagerly during logical plan construction.”
      ↩︎ Lazy execution vs lazy analysis: Spark Classic and Spark Connect
      “Transformations are constructed on the client side and sent as unresolved plans to the server.”
      ↩︎ Lazy execution vs lazy analysis: Spark Classic and Spark Connect
      “Both Spark Classic and Spark Connect follow the same lazy execution model for query execution.”
      ↩︎ Exam trap 1
      “you can trigger eager analysis, for example with df.columns, df.schema, or df.collect()”
      ↩︎ Exam trap 2
      “on df.columns or df.show() an error will be thrown because the unresolved plan is sent to the server for analysis”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise Databricks Certified Associate Developer for Apache Spark in quiz mode.

    Spotted a mistake, or was something unclear? Tell us.