CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 1 · Lesson 7/32

    Pandas API on Spark, Structured Streaming and MLlib

    Identify the features of the Apache Spark Modules, including Core, Spark SQL, DataFrames, Pandas API on Spark, Structured Streaming, and MLib.

    9 min read
    3.12% of exam
    6 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain the gap Pandas API on Spark fills and how to import it
    • Describe Structured Streaming's processing model and its fault-tolerance guarantees
    • Identify the algorithm families and Pipelines API that MLlib provides
    • Choose the right Spark module for a given workload

    1.Pandas API on Spark: pandas syntax at cluster scale

    pandas is a Python package that many data scientists use for data structures and analysis, but "pandas does not scale out to big data." Pandas API on Spark is the Spark module built to fill that gap. It provides "pandas equivalent APIs that work on Apache Spark." The workload runs distributed across multiple nodes, so the same code can run with pandas (for tests and smaller datasets) and with Spark (for production and distributed datasets).

    PySpark users get something out of it too. Pandas API on Spark supports tasks that are difficult in plain PySpark, "for example plotting data directly from a PySpark DataFrame." Version history is the other exam-relevant detail. The module is available starting in Apache Spark 3.2, which ships in Databricks Runtime 10.0 and above. On Databricks Runtime 9.1 LTS and below, you use its predecessor, Koalas. You import it like this:

    Checkpoint 1 of 5· Fill the gap

    Which module name completes the import for Pandas API on Spark?

    import pyspark. ?  as ps

    Because the module runs on Spark, "Various configurations in PySpark could be applied internally in pandas API on Spark." The quickstart's example is Arrow optimization, which speeds up conversion to pandas:

    Enabling Arrow and timing a conversion from pandas-on-Spark to pandaspython
    spark.conf.set("spark.sql.execution.arrow.pyspark.enabled", True)
    %timeit ps.range(300000).to_pandas()

    The Databricks documentation also points to a user guide, a reference and a migration guide for moving from Koalas to Pandas API on Spark. In the quickstart, familiar pandas-style calls such as fillna, mean and groupby work on a pandas-on-Spark DataFrame, which is why existing pandas knowledge carries over so directly.

    Sources123

    2.Structured Streaming: batch logic over arriving data

    "Apache Spark Structured Streaming is a near real-time processing engine that offers end-to-end fault tolerance with exactly-once processing guarantees using familiar Spark APIs." Its main feature is that you do not learn a separate streaming language. You write the computation the way you would write a batch computation on static data, and the engine "performs the computation incrementally and continuously updates the result as streaming data arrives." The PySpark overview adds that it is the Spark SQL engine that runs it incrementally.

    The concepts page lists the features that make this work. Checkpoints store processing state to enable fault tolerance and exactly-once delivery. Output modes (append, update, complete) control what is emitted from stateful queries. Trigger intervals balance latency against cost. Queries are either stateless, processing rows without keeping state, or stateful, keeping intermediate state for aggregations, joins and deduplication. Watermarks control how long a stateful operation waits for late-arriving data.

    The same page also describes where streaming data comes from and goes to. Auto Loader incrementally and efficiently processes new data files as they arrive in cloud storage. Delta Lake tables can serve as streaming sources and sinks with exactly-once processing guarantees, and standard connectors reach message buses, queues and enterprise applications. A micro-batch size setting limits input rates so batches stay a consistent size.

    Checkpoint 2 of 5· Check yourself

    Which Structured Streaming feature stores processing state so a query can recover with exactly-once semantics?

    Checkpoint 3 of 5· Exam question

    A PySpark job registers a standard Python UDF that multiplies every value in a numeric column by a tax rate, but profiling shows the row-by-row Python UDF is the bottleneck on a 200-million-row DataFrame. The team wants to keep the same per-value logic while cutting the serialization overhead between the JVM and the Python worker. Which change achieves that?

    Sources42

    3.MLlib: Spark's machine learning library

    MLlib is the Apache Spark machine learning library. It provides common learning algorithms and utilities for classification, regression, clustering, collaborative filtering and dimensionality reduction, along with "underlying optimization primitives." It is built on Spark and "provides a uniform set of APIs that help users create and tune practical machine learning pipelines."

    In Python, the DataFrame-based library is the pyspark.ml package, which Databricks supports on serverless, standard and dedicated compute. The Databricks example notebooks use the MLlib Pipelines API: a binary classification application, decision-tree classifiers, a gradient-boosted-tree regression that predicts hourly bike rentals, and a custom transformer.

    For reference information about MLlib features, Databricks recommends the MLlib Programming Guide along with the Python, Scala and Java API references. The PySpark overview lists MLlib alongside Spark SQL and DataFrames, Structured Streaming, Pandas API on Spark and GraphX, so it is one module among several that share the same Spark foundation rather than a separate system.

    Checkpoint 4 of 5· Check yourself

    Which of these is NOT listed as an MLlib capability in the Databricks documentation?

    Sources52

    4.Choosing the right module

    Exam questions on this objective usually describe a workload and ask which module fits. The PySpark overview gives the mapping directly. Each module is named for the kind of work it supports, and the table below lines them up with the identifiers you will see in code.

    Spark modules mapped to the workload each one supports
    ModuleWorkloadHow you reach it in Python
    Spark SQL and DataFramesStructured data with relational queriesSparkSession and the DataFrame API
    Structured StreamingScalable processing of streamsThe same DataFrame APIs as batch
    Pandas API on Sparkpandas data structures and analysis on Sparkimport pyspark.pandas as ps
    MLlibMachine learning algorithms and pipelinespyspark.ml

    These modules are not isolated products. Spark SQL uses the same execution engine whichever API or language you use, and Structured Streaming is run by the Spark SQL engine. Read the workload in a question for its key phrase: relational queries point to Spark SQL and DataFrames, streams that arrive continuously point to Structured Streaming, pandas code at scale points to Pandas API on Spark, and learning algorithms point to MLlib.

    Checkpoint 5 of 5· Exam question

    A developer creates a temp view from a DataFrame and runs an aggregation two different ways: ```python df.createOrReplaceTempView("orders") sql_result = spark.sql("SELECT customer_id, SUM(amount) AS total FROM orders GROUP BY customer_id") df_result = df.groupBy("customer_id").agg(sum("amount").alias("total")) ``` What is true about `sql_result` and `df_result`?

    Sources26

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Pandas API on Spark is available on every Databricks Runtime, so Koalas is never needed.Why is that wrong?

      Pandas API on Spark arrived in Apache Spark 3.2 and Databricks Runtime 10.0. Clusters on Databricks Runtime 9.1 LTS and below use Koalas instead.

      Covered in Pandas API on Spark: pandas syntax at cluster scale

    2. 2.Structured Streaming needs its own streaming-specific programming model, separate from batch DataFrame code.Why is that wrong?

      You express a streaming computation the same way as a batch computation on static data, and the engine runs it incrementally as data arrives.

      Covered in Structured Streaming: batch logic over arriving data

    Practise it for real

    See a PySpark configuration take effect inside Pandas API on Spark by timing a conversion to pandas with Arrow on and off

    1. 1.In a notebook with a SparkSession, run import pyspark.pandas as ps and save the current setting with prev = spark.conf.get("spark.sql.execution.arrow.pyspark.enabled").

      Why: You will change a session setting and should restore it afterwards.

      You should see: The import succeeds on Spark 3.2+ / Databricks Runtime 10.0+, and prev holds the current value.

    2. 2.Run spark.conf.set("spark.sql.execution.arrow.pyspark.enabled", True) and then %timeit ps.range(300000).to_pandas().

      Why: Arrow optimization speeds up the internal pandas conversion.

      You should see: A timing result. The quickstart reported about 900 ms per loop.

    3. 3.Run spark.conf.set("spark.sql.execution.arrow.pyspark.enabled", False) and repeat the same %timeit line.

      Why: This isolates the effect of the Arrow setting.

      You should see: A slower timing. The quickstart reported about 3.08 s per loop.

    4. 4.Restore the original value with spark.conf.set("spark.sql.execution.arrow.pyspark.enabled", prev).

      Why: This leaves the session as you found it.

      You should see: The setting returns to its earlier value.

    Stuck? Get a nudge

    Exact timings depend on your cluster. What matters is that a PySpark setting changed how the pandas-on-Spark call performed.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Pandas API on Spark fills this gap by providing pandas equivalent APIs that work on Apache Spark.”
      ↩︎ Pandas API on Spark: pandas syntax at cluster scale
      “Pandas API on Spark is available beginning in Apache Spark 3.2 (which is included beginning in Databricks Runtime 10.0)”
      ↩︎ Pandas API on Spark: pandas syntax at cluster scale
      “pandas API on Spark supports many tasks that are difficult to do with PySpark, for example plotting data directly from a PySpark DataFrame.”
      ↩︎ Pandas API on Spark: pandas syntax at cluster scale
      “For clusters that run Databricks Runtime 9.1 LTS and below, use Koalas instead.”
      ↩︎ Exam trap 1
      “Pandas API on Spark is useful not only for pandas users but also PySpark users”
      ↩︎ Prediction
    2. 2.
      “with a single codebase that works with pandas (tests, smaller datasets) and with Spark (production, distributed datasets)”
      ↩︎ Pandas API on Spark: pandas syntax at cluster scale
      “the Spark SQL engine runs it incrementally and continuously as streaming data continues to arrive.”
      ↩︎ Structured Streaming: batch logic over arriving data
      “MLlib is a scalable machine learning library built on Spark”
      ↩︎ MLlib: Spark's machine learning library
      “provides a uniform set of APIs that help users create and tune practical machine learning pipelines”
      ↩︎ MLlib: Spark's machine learning library
      “Processing of structured data with relational queries with Spark SQL and DataFrames.”
      ↩︎ Choosing the right module
      “Scalable processing of streams with Structured Streaming.”
      ↩︎ Choosing the right module
      “Pandas data structures and data analysis tools that work on Apache Spark with Pandas API on Spark.”
      ↩︎ Choosing the right module
      “Machine learning algorithms with Machine Learning (MLLib).”
      ↩︎ Choosing the right module
    3. 3.
      “Various configurations in PySpark could be applied internally in pandas API on Spark.”
      ↩︎ Pandas API on Spark: pandas syntax at cluster scale
    4. 4.
      “Apache Spark Structured Streaming is a near real-time processing engine that offers end-to-end fault tolerance with exactly-once processing guarantees using familiar Spark APIs.”
      ↩︎ Structured Streaming: batch logic over arriving data
      “Stateless queries process rows without retaining state. Stateful queries maintain intermediate state for aggregations, joins, and deduplication.”
      ↩︎ Structured Streaming: batch logic over arriving data
      “Output mode Choose between append, update, and complete modes for stateful streaming queries.”
      ↩︎ Structured Streaming: batch logic over arriving data
      “Trigger intervals Set trigger intervals to balance latency and cost for your processing requirements.”
      ↩︎ Structured Streaming: batch logic over arriving data
      “Watermarks Control how long Structured Streaming waits for late-arriving data in stateful operations.”
      ↩︎ Structured Streaming: batch logic over arriving data
      “Auto Loader Incrementally and efficiently process new data files as they arrive in cloud storage.”
      ↩︎ Structured Streaming: batch logic over arriving data
      “Structured Streaming lets you express computation on streaming data in the same way you express a batch computation on static data.”
      ↩︎ Exam trap 2
      “Checkpoints Store processing state to enable fault tolerance and exactly-once delivery semantics.”
      ↩︎ Checkpoint
    5. 5.
      “The pyspark.ml package from Apache Spark MLlib is supported on serverless, standard, and dedicated compute.”
      ↩︎ MLlib: Spark's machine learning library
      “This notebook shows you how to build a binary classification application using the Apache Spark MLlib Pipelines API.”
      ↩︎ MLlib: Spark's machine learning library
      “Apache Spark MLlib is the Apache Spark machine learning library consisting of common learning algorithms and utilities, including classification, regression, clustering, collaborative filtering”
      ↩︎ Checkpoint
    6. 6.
      “When computing a result, the same execution engine is used, independent of which API/language you are using to express the computation.”
      ↩︎ Choosing the right module

    Spotted a mistake, or was something unclear? Tell us.