CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 7 · Lesson 31/32

    Pandas API on Spark: Advantages of pyspark.pandas

    Explain the advantages of using Pandas API on Spark.

    13 min read
    3.12% of exam
    6 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain the scaling gap in pandas that Pandas API on Spark fills, and say who it suits best
    • Describe how one pandas-style codebase can run on small local data and on distributed production data
    • Name advantages that matter to PySpark users as well as pandas users, such as plotting, broad I/O support and Arrow-optimized conversion
    • Weigh Pandas API on Spark against plain pandas and PySpark, including its flexibility limits, row-order behaviour and how it replaces Koalas

    Key concept

    Pandas API on Spark (pyspark.pandas) — A pandas-compatible API that runs on Apache Spark. You keep writing pandas-style code, and Spark spreads the work across a cluster instead of one machine's memory.

    1.The gap: pandas is familiar, but it does not scale

    Every advantage of Pandas API on Spark comes back to one problem. pandas is the Python package data scientists reach for first. It is easy to use, it has rich data structures, and most analysts already know it. But it was built to run on one machine. The Databricks documentation puts it bluntly: pandas does not scale out to big data.

    That leaves teams with two bad options. They can keep working in pandas until the data outgrows one machine. Or they can rewrite everything in a new API, PySpark, and retrain the people who write it. Pandas API on Spark removes that choice. It offers pandas-equivalent APIs, but the work runs on Apache Spark, so the familiar syntax gets Spark's distributed scale behind it.

    Databricks names two ways to run Python. For single-machine computing, you use Python libraries such as pandas as usual. For distributed Python workloads, it offers two APIs out of the box: PySpark and Pandas API on Spark. So the exam treats Pandas API on Spark as a distributed tool, not a faster single-machine pandas. Databricks also says who it suits best: data scientists who know pandas but not Apache Spark. They can move to distributed processing without first learning a new DataFrame API.

    The import statement for Pandas API on Spark (available beginning in Apache Spark 3.2)python
    import pyspark.pandas as ps

    Checkpoint 1 of 7· Check yourself

    A data science team knows pandas well but has never used Apache Spark. Their datasets have outgrown a single machine. According to Databricks, which option fits them best?

    Checkpoint 2 of 7· Exam question

    What is the primary limitation of native pandas that the Pandas API on Spark (`pyspark.pandas`) is designed to address?

    Sources12

    2.One codebase, from laptop-sized tests to distributed production

    The second advantage follows from the first. Because the API mirrors pandas, the same code can serve small and large data. Databricks says Pandas API on Spark lets you scale a pandas workload to any size by running it across multiple nodes. It does this with a single codebase that works with pandas for tests and smaller datasets, and with Spark for production and distributed datasets.

    That changes how a project grows. You write and unit-test the logic on a small sample, then run the same logic on the full dataset on a cluster. You do not keep a separate PySpark version of the pipeline. The reference documentation backs this up: pandas API on Spark follows the API specifications of the latest pandas release. Databricks describes it as familiar pandas commands on top of PySpark DataFrames.

    Creating a pandas-on-Spark DataFrame uses the same dict-plus-index constructor you would use in pandaspython
    psdf = ps.DataFrame(
        {'a': [1, 2, 3, 4, 5, 6],
         'b': [100, 200, 300, 400, 500, 600],
         'c': ["one", "two", "three", "four", "five", "six"]},
        index=[10, 20, 30, 40, 50, 60])

    The quickstart shows how close the match is. You can convert an existing pandas DataFrame into a pandas-on-Spark DataFrame. The result has type pyspark.pandas.frame.DataFrame, and the docs say it looks and behaves the same as a pandas DataFrame. Everyday pandas calls carry over unchanged: describe() for summary statistics, sort_values(by='B'), dropna(how='any'), fillna(value=5), mean(), and split-apply-combine aggregations such as the one below. You can also convert DataFrames between pandas and PySpark, so existing pandas data is not stranded. (The conversion methods are covered in a separate lesson.)

    A pandas-style group-by aggregation, executed by Sparkpython
    psdf.groupby('A').sum()

    Checkpoint 3 of 7· Fill the gap

    The quickstart turns an existing pandas DataFrame pdf into a pandas-on-Spark DataFrame. Which function completes the line?

    psdf = ps. ? (pdf)

    Checkpoint 4 of 7· Exam question

    A data scientist has an existing pandas script that cleans a DataFrame using `df.dropna()` followed by `df.groupby('region')['sales'].sum()`. They want to reuse this exact logic against a 500 GB Delta table without rewriting the pandas calls. Which change lets them do this with the pandas API on Spark?

    Sources3456

    3.Advantages for PySpark users too

    It is easy to assume Pandas API on Spark only helps people who know pandas. Databricks says otherwise. It supports many tasks that are difficult to do with PySpark, and the example it gives is plotting data directly from a PySpark DataFrame. In the quickstart, plotting is one method call. On a DataFrame, the plot() method plots all of the columns with labels, the way a pandas user would expect.

    The API is also broad. The reference covers input/output (Delta Lake, Parquet, ORC, CSV, JSON, Excel, SQL, Spark metastore tables and generic Spark I/O). It also covers Series and DataFrame operations, indexes, window functions, GroupBy, resampling, options, MLflow utilities and testing assertions. Reading and writing Spark data sources needs no new API. Pandas-style readers and writers work directly, as in this Parquet round trip from the quickstart:

    Writing and reading Parquet with pandas-style methodspython
    psdf.to_parquet('bar.parquet')
    ps.read_parquet('bar.parquet').head(10)

    Pandas API on Spark also gets Spark's tuning for free. The quickstart says various PySpark configurations can be applied internally in pandas API on Spark. For example, you can enable Arrow optimization to hugely speed up internal pandas conversion. In the quickstart's own timing, converting ps.range(300000) to pandas took about 900 ms with Arrow enabled and about 3.08 s with it disabled. Databricks makes a related point at platform level: Spark's pandas UDFs and pandas function APIs use Arrow optimizations as well. Those are covered in their own lesson.

    Enabling Arrow optimization, a PySpark setting that speeds up pandas conversion in Pandas API on Sparkpython
    spark.conf.set("spark.sql.execution.arrow.pyspark.enabled", True)
    %timeit ps.range(300000).to_pandas()

    Checkpoint 5 of 7· Match them up

    Match each advantage to the evidence the documentation gives for it

    Tap a term, then the definition that fits it.

    Checkpoint 6 of 7· Exam question

    An engineer benchmarks converting a 300,000-row Spark DataFrame to a pandas-on-Spark DataFrame and finds the operation takes several seconds. A teammate suggests a configuration change that previously cut a similar conversion time from about 3 seconds to under 1 second. Which setting most likely explains that speedup?

    Sources164

    4.Knowing where the advantages stop

    An exam answer that lists only benefits misses half the objective. Explaining the advantages also means knowing when Pandas API on Spark is the wrong tool. Databricks is explicit that PySpark is more flexible. PySpark is the official Python API for Spark, with extensive support for Spark SQL, Structured Streaming, MLlib and GraphX. Pandas API on Spark wins on familiarity and a shared codebase. It does not replace PySpark for everything.

    Where each Python DataFrame option fits, per the Databricks documentation
    OptionRuns onMain advantageMain limitation
    pandasA single machineEasy-to-use, widely known data structuresDoes not scale out to big data
    Pandas API on Spark (pyspark.pandas)Spark, distributed across multiple nodespandas-equivalent APIs at Spark scale; one codebase for tests and productionLess flexible than PySpark
    PySparkSpark, distributedMost flexible: Spark SQL, Structured Streaming, MLlib, GraphXNot pandas syntax, so pandas users must learn a new API

    "Behaves like pandas" has limits too. The quickstart warns that data in a Spark DataFrame does not keep its natural order by default. In pandas, head() returns the first rows in their original order. In Pandas API on Spark, that is only guaranteed if you set the compute.ordered_head option, and that option adds the cost of an internal sort. The data is distributed, so some pandas assumptions cost extra to keep.

    Finally, the history behind the name. Pandas API on Spark grew out of the open-source Koalas project, and Koalas now recommends switching to Pandas API on Spark. It is available beginning in Apache Spark 3.2 through import pyspark.pandas as ps. Only clusters on Databricks Runtime 9.1 LTS and below should still use Koalas.

    Checkpoint 7 of 7· Check yourself

    Which statement about Pandas API on Spark is accurate?

    Sources261

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Pandas API on Spark only benefits people who already know pandas; PySpark users gain nothing from it.Why is that wrong?

      Databricks says it also helps PySpark users, because it makes tasks that are hard in PySpark easy, such as plotting directly from a PySpark DataFrame.

      Covered in Advantages for PySpark users too

    2. 2.Because it scales like Spark, Pandas API on Spark is a full replacement for PySpark and is just as flexible.Why is that wrong?

      Databricks says PySpark is the more flexible API, with extensive support for Spark SQL, Structured Streaming, MLlib and GraphX. Pandas API on Spark's advantage is pandas familiarity at scale.

      Covered in Knowing where the advantages stop

    3. 3.Since pandas-on-Spark behaves like pandas, head() returns the first rows in their original order, just as in pandas.Why is that wrong?

      Spark DataFrames do not keep natural order by default. Keeping it requires the compute.ordered_head option, which adds the cost of an internal sort.

      Covered in Knowing where the advantages stop

    4. 4.Koalas is still the package to use for pandas-style code on current Databricks runtimes.Why is that wrong?

      The Koalas project now recommends switching to Pandas API on Spark. Koalas is only for clusters on Databricks Runtime 9.1 LTS and below.

      Covered in Knowing where the advantages stop

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “However, pandas does not scale out to big data.”
      ↩︎ The gap: pandas is familiar, but it does not scale
      “pandas API on Spark supports many tasks that are difficult to do with PySpark, for example plotting data directly from a PySpark DataFrame.”
      ↩︎ Advantages for PySpark users too
      “Pandas API on Spark is available beginning in Apache Spark 3.2”
      ↩︎ Knowing where the advantages stop
      “For clusters that run Databricks Runtime 9.1 LTS and below, use Koalas instead.”
      ↩︎ Knowing where the advantages stop
      “Pandas API on Spark fills this gap by providing pandas equivalent APIs that work on Apache Spark.”
      ↩︎ Key concept
      “Pandas API on Spark is useful not only for pandas users but also PySpark users”
      ↩︎ Exam trap 1
      “Pandas API on Spark is useful not only for pandas users but also PySpark users”
      ↩︎ Prediction
    2. 2.
      “For distributed Python workloads, Databricks offers two popular APIs out of the box: PySpark and Pandas API on Spark.”
      ↩︎ The gap: pandas is familiar, but it does not scale
      “For single-machine computing, you can use Python APIs and libraries as usual”
      ↩︎ The gap: pandas is familiar, but it does not scale
      “PySpark is more flexible than the Pandas API on Spark”
      ↩︎ Knowing where the advantages stop
      “PySpark is more flexible than the Pandas API on Spark”
      ↩︎ Exam trap 2
      “The Koalas open-source project now recommends switching to the Pandas API on Spark.”
      ↩︎ Exam trap 4
      “This open-source API is an ideal choice for data scientists who are familiar with pandas but not Apache Spark.”
      ↩︎ Checkpoint
    3. 3.
      “Pandas API on Spark allows you to scale your pandas workload to any size by running it distributed across multiple nodes”
      ↩︎ One codebase, from laptop-sized tests to distributed production
      “with a single codebase that works with pandas (tests, smaller datasets) and with Spark (production, distributed datasets)”
      ↩︎ One codebase, from laptop-sized tests to distributed production
    4. 4.
      “Pandas API on Spark provides familiar pandas commands on top of PySpark DataFrames.”
      ↩︎ One codebase, from laptop-sized tests to distributed production
      “You can also convert DataFrames between pandas and PySpark.”
      ↩︎ One codebase, from laptop-sized tests to distributed production
      “Apache Spark also supports pandas UDFs, which use similar Arrow-optimizations for arbitrary user functions defined in Python.”
      ↩︎ Advantages for PySpark users too
    5. 6.
      “It looks and behaves the same as a pandas DataFrame.”
      ↩︎ One codebase, from laptop-sized tests to distributed production
      “the plot() method is a convenience to plot all of the columns with labels”
      ↩︎ Advantages for PySpark users too
      “Various configurations in PySpark could be applied internally in pandas API on Spark.”
      ↩︎ Advantages for PySpark users too
      “you can enable Arrow optimization to hugely speed up internal pandas conversion”
      ↩︎ Advantages for PySpark users too
      “The natural order can be preserved by setting compute.ordered_head option but it causes a performance overhead with sorting internally.”
      ↩︎ Knowing where the advantages stop
      “Note that the data in a Spark dataframe does not preserve the natural order by default.”
      ↩︎ Exam trap 3

    Ready to test yourself?

    Practise the 11 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.