CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 3 · Lesson 18/32

    Sorting DataFrames and Printing Schemas in PySpark

    Perform operations on DataFrames such as sorting, iterating, printing schema, and conversion between DataFrame and sequence/list formats.

    9 min read
    3.12% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Sort a DataFrame by one or more columns in ascending or descending order using sort or orderBy
    • Use column ordinals, including negative ordinals, to sort
    • Print a nested schema with printSchema and control its depth with the level parameter
    • Pick between schema, dtypes and columns when you need schema information as data

    Key concept

    DataFrame as a distributed, named-column table — A DataFrame is a table spread across a cluster. Each column has a name and a type, and the rows are partitioned instead of stored in one place. Sorting sets an order over those distributed rows, and the schema describes the named, typed columns.

    1.Sorting with sort and orderBy

    PySpark has two names for one operation. sort returns a new DataFrame sorted by the column(s) you pass in, and orderBy is an alias for it. Neither changes the original DataFrame. Each returns a new one. You can pass columns as strings, as Column objects, as a list, or as integer ordinals. The order is ascending unless you ask for something else.

    You can make a sort descending in three ways. You can call .desc() on a column (df.age.desc()). You can use the desc/asc helpers from pyspark.sql.functions (sf.desc("age")). Or you can pass the ascending keyword. ascending takes one boolean or a list of booleans. If you pass a list, it must be the same length as the list of columns, so each sort key gets its own direction.

    Sorting on two columns with mixed directions. orderBy and sort give the same result.python
    df_sorted = df_customer.orderBy(col("c_acctbal").desc(), col("c_custkey").asc())
    df_sorted = df_customer.sort(col("c_acctbal").desc(), col("c_custkey").asc())

    The reference example df.orderBy(sf.desc("age"), "name") shows how ties are broken. Rows are ordered by age descending first. Rows with the same age are then ordered by name, ascending by default. In that example the two 2-year-olds come out as Alice and then Bob.

    Checkpoint 1 of 6· Fill the gap

    Which keyword argument makes this call sort by age in descending order?

    df.sort("age",  ? =False).show()

    Checkpoint 2 of 6· Exam question

    A report needs rows ordered by `department` ascending, and within each department by `salary` descending. Given: ```python from pyspark.sql.functions import col df.orderBy(___).show() ``` Which argument produces the required ordering?

    Sources123

    2.Column ordinals, and why sorted order doesn't last

    sort also accepts integers that point to columns by position. There are two catches. First, ordinals start at 1. That is different from the 0-based positions used when you index a DataFrame with df[0]. Second, a negative ordinal means "sort this column descending". With columns ["age", "name"], df.sort(1) sorts by age ascending and df.sort(-1) sorts by age descending. Neither one refers to the last column the way a negative index does in a Python list.

    Checkpoint 3 of 6· Check yourself

    A DataFrame has columns ["age", "name"]. What does df.sort(-1) do?

    Because the rows are spread across partitions, sorting is not free. The Databricks guide warns that sorting can be expensive at scale and that you should sort deliberately. Sorted order is also not something the data keeps. If you write sorted data to storage and read it back with Spark, the guide says the order is not guaranteed. If later steps need a particular order, sort again after you read the data. A common pattern is to sort and then call limit, for example display(df_sorted.limit(10)), to see the top results.

    Sources23

    Every DataFrame has a schema: the name, type and nullability of each column. printSchema() prints it to the console as a tree whose root is root. For a flat DataFrame built from (14, "Tom")-style tuples, the tree has one line per column, such as age: long (nullable = true) and name: string (nullable = true). The types were inferred from the Python values.

    printSchema with a nested struct column, at depth 1 and depth 2python
    df = spark.createDataFrame([(1, (2, 2))], ["a", "b"])
    df.printSchema(1)
    # root
    #  |-- a: long (nullable = true)
    #  |-- b: struct (nullable = true)
    
    df.printSchema(2)
    # root
    #  |-- a: long (nullable = true)
    #  |-- b: struct (nullable = true)
    #  |    |-- _1: long (nullable = true)
    #  |    |-- _2: long (nullable = true)

    The optional level parameter (an int) sets how many levels of a nested schema to print. At level 1 you only see that b is a struct. At level 2 its fields _1 and _2 appear, indented one level further. This helps with deeply nested data such as parsed JSON, where the full tree can be long.

    Checkpoint 4 of 6· Check yourself

    You want to see a struct column's type but not its child fields. Which call does that, for the DataFrame above?

    Checkpoint 5 of 6· Exam question

    A `df` has columns `name` and `age`, and its first row holds `Alice` and `30`. Given: ```python rows = df.select("name", "age").collect() print(rows[0].name, rows[0]["age"]) ``` What does this print?

    Sources4

    4.Schema information as data: schema, dtypes, columns

    printSchema only prints. It is for reading the schema yourself, not for code to inspect. When code needs the schema, use one of three DataFrame properties. They are properties, so you don't call them with parentheses. Each one gives the schema in a different shape.

    Ways to inspect a DataFrame's structure and what each returns
    MemberKindWhat you get
    printSchema(level)MethodPrints the schema as a tree to the console
    schemaPropertyThe schema as a StructType
    dtypesPropertyAll column names and their data types, as a list
    columnsPropertyThe names of all columns, as a list

    Checkpoint 6 of 6· Match them up

    Match each DataFrame member to what it gives you

    Tap a term, then the definition that fits it.

    Sources5

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Sort ordinals work like Python indexes: df.sort(0) sorts by the first column, and -1 means the last column.Why is that wrong?

      Sort ordinals start at 1. A negative ordinal means sort that column descending, not count back from the end.

      Covered in Column ordinals, and why sorted order doesn't last

    2. 2.If you sort a DataFrame before writing it, Spark will read the rows back in that sorted order.Why is that wrong?

      The Databricks guide says order is not guaranteed after you store sorted data and reload it with Spark. Sort again when you need the order.

      Covered in Column ordinals, and why sorted order doesn't last

    3. 3.With ascending=[True, False] you can sort by any number of columns, and the list is applied as far as it reaches.Why is that wrong?

      When ascending is a list, it must be exactly as long as the list of columns.

      Covered in Sorting with sort and orderBy

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 2.
      “Returns a new DataFrame sorted by the specified column(s).”
      ↩︎ Sorting with sort and orderBy
      “If a column ordinal is negative, it means sort descending.”
      ↩︎ Column ordinals, and why sorted order doesn't last
      “A column ordinal starts from 1, which is different from the 0-based __getitem__.”
      ↩︎ Exam trap 1
      “If a list is specified, the length of the list must equal the length of the cols.”
      ↩︎ Exam trap 3
    2. 3.
      “By default these methods sort in ascending order”
      ↩︎ Sorting with sort and orderBy
      “Sorting can be expensive at scale”
      ↩︎ Column ordinals, and why sorted order doesn't last
      “if you store sorted data and reload the data with Spark, order is not guaranteed.”
      ↩︎ Exam trap 2
    3. 4.
      “Prints out the schema in the tree format.”
      ↩︎ Printing the schema as a tree
      “Optionally allows to specify how many levels to print if schema is nested.”
      ↩︎ Checkpoint
    4. 5.
      “Returns the schema of this DataFrame as a StructType.”
      ↩︎ Schema information as data: schema, dtypes, columns
      “Retrieves the names of all columns in the DataFrame as a list.”
      ↩︎ Schema information as data: schema, dtypes, columns
      “A distributed collection of data grouped into named columns.”
      ↩︎ Key concept
      “Returns all column names and their data types as a list.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Converting PySpark DataFrames to Lists and Iterating Rows

    Spotted a mistake, or was something unclear? Tell us.