CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 3 · Lesson 18/32

    Converting PySpark DataFrames to Lists and Iterating Rows

    Perform operations on DataFrames such as sorting, iterating, printing schema, and conversion between DataFrame and sequence/list formats.

    9 min read
    3.12% of exam
    10 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Build a DataFrame from a Python list of tuples, with an inferred or explicit schema
    • Turn a DataFrame into a list of Row objects with collect, take, head, first and tail, and read values from each Row
    • Iterate rows with toLocalIterator or foreach, and explain the memory difference from collect
    • Decide when toPandas is safe to use

    1.From a Python list to a DataFrame

    Conversion works in both directions, and the simplest starting point is a Python list. spark.createDataFrame takes the data as a list of tuples, one tuple per row, plus a schema. The schema can be just a list of column names.

    Creating a DataFrame from a list of tuples and a list of column namespython
    df_children = spark.createDataFrame(
      data = [("Mikhail", 15), ("Zaky", 13), ("Zoya", 8)],
      schema = ['name', 'age'])

    If you pass only names, Spark infers each column's type from the Python values. To fix the types yourself, pass a StructType built from StructFields. Each field gives a name, a data type such as StringType() or IntegerType(), and a nullable flag. You import these from pyspark.sql.types.

    Checkpoint 1 of 6· Check yourself

    You build a DataFrame with schema = ['name', 'age'] and no types. What decides the type of the age column?

    Sources1

    2.From a DataFrame to a list: collect and Row objects

    Going the other way, collect() returns every record in the DataFrame as a Python list of Row objects, for example [Row(age=14, name='Tom'), ...]. You can filter or select first, and only the rows that remain are collected. Each Row lets you read a field by name with row["name"], and row.asDict() turns it into a plain dictionary. Combined with a list comprehension, that turns a DataFrame into an ordinary Python list of values or dicts.

    Turning collected Rows into a list of values and a list of dictspython
    rows = df.collect()
    [row["name"] for row in rows]
    # ['Tom', 'Alice', 'Bob']
    
    [row.asDict() for row in rows]
    # [{'age': 14, 'name': 'Tom'}, {'age': 23, 'name': 'Alice'}, {'age': 16, 'name': 'Bob'}]

    That convenience has a cost. The DataFrame lives on the cluster, but the list lives in one Python process, the driver. The reference warns that collect should only be used when the result is expected to be small, because all of the data is loaded into the driver's memory.

    Checkpoint 2 of 6· Check yourself

    Why does the reference say collect() should only be used on small results?

    Checkpoint 3 of 6· Exam question

    A pipeline needs to hand a downstream JSON serializer a Python list of dictionaries, one per row, mapping each column name to its value. Given: ```python data = df.collect() records = [___ for row in data] ``` Which expression completes the list comprehension?

    Sources2

    3.Bringing back only some rows: take, head, first, tail

    Often you only need a few rows. take(num) returns the first num rows as a list of Row, or all of them if the DataFrame is smaller. tail(num) does the same from the end. The reference warns that tail moves data into the driver, and a very large num can crash the driver with OutOfMemoryError. head and first look alike but return different things.

    head without n returns a Row. head with n returns a list.python
    df.head()
    # Row(age=2, name='Alice')
    df.head(1)
    # [Row(age=2, name='Alice')]
    df.head(0)
    # []

    first() returns the first row as a single Row. If the DataFrame is empty it returns None instead of raising an error, so code that calls first() on a possibly empty result should check for None.

    Which row-retrieval call returns what
    CallReturnsNote
    collect()list of Row (all records)All data loaded into driver memory
    take(num)list of Row (first num)All records if fewer than num exist
    tail(num)list of Row (last num)Very large num can cause OutOfMemoryError on the driver
    head() / head(n)Row / list of Rown defaults to 1
    first()RowNone if the DataFrame is empty

    Checkpoint 4 of 6· Check yourself

    A filter leaves the DataFrame empty. What does df.first() return?

    Sources3

    4.Iterating rows: toLocalIterator and foreach

    To loop over every row without building the whole list at once, use toLocalIterator(). It returns an iterator over all the rows. Its memory footprint is the key difference from collect: the iterator uses about as much memory as the largest partition, not the whole DataFrame. With prefetchPartitions=True, Spark fetches the next partition before it is needed, which can use the memory of the two largest partitions.

    Checkpoint 5 of 6· Fill the gap

    Which method returns an iterator whose memory use is bounded by the largest partition?

    list(df. ? ())

    Sometimes you want to run a function on each row rather than loop yourself. foreach(f) applies f to every Row of the DataFrame. f takes one parameter, the row being processed, and returns nothing. foreachPartition(f) is the per-partition version: f receives an iterator of the rows in one partition.

    Applying a function to every Row with foreachpython
    def func(person):
        print(person.name)
    
    df.foreach(func)

    Sources456

    5.Converting to pandas with toPandas

    toPandas() returns the contents of the DataFrame as a pandas.DataFrame, with a pandas index (0, 1, ...) next to the Spark columns. It has the same limitation as collect: use it only when the result is expected to be small, because all the data is loaded into the driver's memory. It also needs pandas to be installed. The reference says that using it with spark.sql.execution.arrow.pyspark.enabled=True is experimental.

    Checkpoint 6 of 6· Match them up

    Match each call to the shape of what it gives you

    Tap a term, then the definition that fits it.

    Sources7

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.toLocalIterator() is just collect() under another name, so it needs enough driver memory for the whole DataFrame.Why is that wrong?

      The iterator uses about as much memory as the largest partition, or the two largest with prefetching. It does not hold the whole DataFrame at once.

      Covered in Iterating rows: toLocalIterator and foreach

    2. 2.df.head() and df.head(1) return the same thing.Why is that wrong?

      Without n, head returns a single Row. With n, it returns a list of Row, even when n is 1.

      Covered in Bringing back only some rows: take, head, first, tail

    3. 3.tail(num) is always safe because it only returns the last few rows.Why is that wrong?

      tail moves data into the driver process, and a very large num can crash the driver with OutOfMemoryError.

      Covered in Bringing back only some rows: take, head, first, tail

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Schemas are defined using the StructType which is made up of StructFields”
      ↩︎ From a Python list to a DataFrame
      “Notice in the output that the data types of columns of df_children are automatically inferred.”
      ↩︎ Checkpoint
    2. 2.
      “Returns all the records in the DataFrame as a list of Row.”
      ↩︎ From a DataFrame to a list: collect and Row objects
      “as all the data is loaded into the driver's memory.”
      ↩︎ Checkpoint
    3. 3.
      “Will return this number of records or all records if the DataFrame contains less than this number of records.”
      ↩︎ Bringing back only some rows: take, head, first, tail
    4. 4.
      “With prefetch it may consume up to the memory of the 2 largest partitions.”
      ↩︎ Iterating rows: toLocalIterator and foreach
      “The iterator will consume as much memory as the largest partition in this DataFrame.”
      ↩︎ Exam trap 1
    5. 5.
      “A function that accepts one parameter which will receive each row to process.”
      ↩︎ Iterating rows: toLocalIterator and foreach
    6. 6.
      “A function that accepts one parameter which will receive each partition to process.”
      ↩︎ Iterating rows: toLocalIterator and foreach
    7. 7.
      “Usage with spark.sql.execution.arrow.pyspark.enabled=True is experimental.”
      ↩︎ Converting to pandas with toPandas
      “This method should only be used if the resulting Pandas pandas.DataFrame is expected to be small”
      ↩︎ Checkpoint

    Also cited

    Ready to test yourself?

    Practise Databricks Certified Associate Developer for Apache Spark in quiz mode.

    Spotted a mistake, or was something unclear? Tell us.