What you will be able to do
- Sort a DataFrame by one or more columns in ascending or descending order using sort or orderBy
- Use column ordinals, including negative ordinals, to sort
- Print a nested schema with printSchema and control its depth with the level parameter
- Pick between schema, dtypes and columns when you need schema information as data
Key concept
DataFrame as a distributed, named-column table — A DataFrame is a table spread across a cluster. Each column has a name and a type, and the rows are partitioned instead of stored in one place. Sorting sets an order over those distributed rows, and the schema describes the named, typed columns.
1.Sorting with sort and orderBy
PySpark has two names for one operation. sort returns a new DataFrame sorted by the column(s) you pass in, and orderBy is an alias for it. Neither changes the original DataFrame. Each returns a new one. You can pass columns as strings, as Column objects, as a list, or as integer ordinals. The order is ascending unless you ask for something else.
You can make a sort descending in three ways. You can call .desc() on a column (df.age.desc()). You can use the desc/asc helpers from pyspark.sql.functions (sf.desc("age")). Or you can pass the ascending keyword. ascending takes one boolean or a list of booleans. If you pass a list, it must be the same length as the list of columns, so each sort key gets its own direction.
df_sorted = df_customer.orderBy(col("c_acctbal").desc(), col("c_custkey").asc())
df_sorted = df_customer.sort(col("c_acctbal").desc(), col("c_custkey").asc())The reference example df.orderBy(sf.desc("age"), "name") shows how ties are broken. Rows are ordered by age descending first. Rows with the same age are then ordered by name, ascending by default. In that example the two 2-year-olds come out as Alice and then Bob.
Checkpoint 1 of 6· Fill the gap
Which keyword argument makes this call sort by age in descending order?
df.sort("age", ? =False).show()sort and orderBy take an ascending keyword that defaults to True. Setting it to False sorts in descending order. There is no descending or reverse parameter.
Checkpoint 2 of 6· Exam question
A report needs rows ordered by `department` ascending, and within each department by `salary` descending. Given: ```python from pyspark.sql.functions import col df.orderBy(___).show() ``` Which argument produces the required ordering?
Correct answer: A — `col("department").asc(), col("salary").desc()`
- A. Passing `col("department").asc()` followed by `col("salary").desc()` sorts primarily on department in ascending order and breaks ties within each department by descending salary, matching the requirement exactly.
- B. Plain string column names default to ascending order for both columns, so salary within each department would sort lowest-to-highest instead of the required descending order.
- C. This reverses the department direction to descending, which sorts departments Z-to-A instead of A-to-Z, and also sorts salary ascending, so neither direction matches the requirement.
- D. `orderBy` does not interpret a leading `-` on a string as a descending marker; this list form only accepts plain column-name strings and always sorts them ascending, so salary would not be descending.
- E. This sorts primarily by salary descending and only uses department as the tiebreaker, which reorders the whole result by salary first instead of grouping rows by department as the primary key.
2.Column ordinals, and why sorted order doesn't last
sort also accepts integers that point to columns by position. There are two catches. First, ordinals start at 1. That is different from the 0-based positions used when you index a DataFrame with df[0]. Second, a negative ordinal means "sort this column descending". With columns ["age", "name"], df.sort(1) sorts by age ascending and df.sort(-1) sorts by age descending. Neither one refers to the last column the way a negative index does in a Python list.
Checkpoint 3 of 6· Check yourself
A DataFrame has columns ["age", "name"]. What does df.sort(-1) do?
Ordinals are 1-based, so the 1 in -1 refers to the first column, age. The minus sign makes the sort descending. It does not count back from the end the way Python list indexing does.
“If a column ordinal is negative, it means sort descending.”Source: docs.databricks.com
Because the rows are spread across partitions, sorting is not free. The Databricks guide warns that sorting can be expensive at scale and that you should sort deliberately. Sorted order is also not something the data keeps. If you write sorted data to storage and read it back with Spark, the guide says the order is not guaranteed. If later steps need a particular order, sort again after you read the data. A common pattern is to sort and then call limit, for example display(df_sorted.limit(10)), to see the top results.
3.Printing the schema as a tree
Every DataFrame has a schema: the name, type and nullability of each column. printSchema() prints it to the console as a tree whose root is root. For a flat DataFrame built from (14, "Tom")-style tuples, the tree has one line per column, such as age: long (nullable = true) and name: string (nullable = true). The types were inferred from the Python values.
df = spark.createDataFrame([(1, (2, 2))], ["a", "b"])
df.printSchema(1)
# root
# |-- a: long (nullable = true)
# |-- b: struct (nullable = true)
df.printSchema(2)
# root
# |-- a: long (nullable = true)
# |-- b: struct (nullable = true)
# | |-- _1: long (nullable = true)
# | |-- _2: long (nullable = true)The optional level parameter (an int) sets how many levels of a nested schema to print. At level 1 you only see that b is a struct. At level 2 its fields _1 and _2 appear, indented one level further. This helps with deeply nested data such as parsed JSON, where the full tree can be long.
Checkpoint 4 of 6· Check yourself
You want to see a struct column's type but not its child fields. Which call does that, for the DataFrame above?
In the reference example, level 1 prints b: struct without _1 and _2. The parameter is called level, not nested, and schema is a property, not a method.
“Optionally allows to specify how many levels to print if schema is nested.”Source: docs.databricks.com
Checkpoint 5 of 6· Exam question
A `df` has columns `name` and `age`, and its first row holds `Alice` and `30`. Given: ```python rows = df.select("name", "age").collect() print(rows[0].name, rows[0]["age"]) ``` What does this print?
Correct answer: A — Alice 30
- A. `collect()` returns a Python list of `Row` objects, and `Row` supports both attribute access (`rows[0].name`) and dictionary-style indexing (`rows[0]["age"]`), so this prints the two underlying values, Alice and 30.
- B. Dot notation on a `Row` returns the individual field value, not the whole `Row` object, so `rows[0].name` evaluates to the string `Alice` rather than a repr of the entire row.
- C. `Row` is a subclass of `tuple` that specifically adds named attribute access, so `rows[0].name` is valid syntax and does not raise an error.
- D. Indexing a `Row` with a column name string returns that single field's value, not another `Row` wrapping the remaining fields, so `rows[0]["age"]` evaluates to the integer 30.
- E. `collect()` materializes data rows, not a header of column names, so this prints the row's actual values rather than the schema's field names.
Sources4
4.Schema information as data: schema, dtypes, columns
printSchema only prints. It is for reading the schema yourself, not for code to inspect. When code needs the schema, use one of three DataFrame properties. They are properties, so you don't call them with parentheses. Each one gives the schema in a different shape.
| Member | Kind | What you get |
|---|---|---|
| printSchema(level) | Method | Prints the schema as a tree to the console |
| schema | Property | The schema as a StructType |
| dtypes | Property | All column names and their data types, as a list |
| columns | Property | The names of all columns, as a list |
Checkpoint 6 of 6· Match them up
Match each DataFrame member to what it gives you
Tap a term, then the definition that fits it.
Only printSchema prints. schema gives the structured StructType, dtypes pairs each name with its type, and columns gives just the names.
“Returns all column names and their data types as a list.”Source: docs.databricks.com
Sources5
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Sort ordinals work like Python indexes: df.sort(0) sorts by the first column, and -1 means the last column.Why is that wrong?
Sort ordinals start at 1. A negative ordinal means sort that column descending, not count back from the end.
Covered in Column ordinals, and why sorted order doesn't last
2.If you sort a DataFrame before writing it, Spark will read the rows back in that sorted order.Why is that wrong?
The Databricks guide says order is not guaranteed after you store sorted data and reload it with Spark. Sort again when you need the order.
Covered in Column ordinals, and why sorted order doesn't last
3.With ascending=[True, False] you can sort by any number of columns, and the list is applied as far as it reaches.Why is that wrong?
When ascending is a list, it must be exactly as long as the list of columns.
Covered in Sorting with sort and orderBy
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“orderBy is an alias for sort.”
↩︎ Sorting with sort and orderBy - 2.
“Returns a new DataFrame sorted by the specified column(s).”
↩︎ Sorting with sort and orderBy“If a column ordinal is negative, it means sort descending.”
↩︎ Column ordinals, and why sorted order doesn't last“A column ordinal starts from 1, which is different from the 0-based __getitem__.”
↩︎ Exam trap 1“If a list is specified, the length of the list must equal the length of the cols.”
↩︎ Exam trap 3 - 3.https://docs.databricks.com/aws/en/pyspark/basicsOfficial docs
“By default these methods sort in ascending order”
↩︎ Sorting with sort and orderBy“Sorting can be expensive at scale”
↩︎ Column ordinals, and why sorted order doesn't last“if you store sorted data and reload the data with Spark, order is not guaranteed.”
↩︎ Exam trap 2 - 4.
“Prints out the schema in the tree format.”
↩︎ Printing the schema as a tree“Optionally allows to specify how many levels to print if schema is nested.”
↩︎ Checkpoint - 5.
“Returns the schema of this DataFrame as a StructType.”
↩︎ Schema information as data: schema, dtypes, columns“Retrieves the names of all columns in the DataFrame as a list.”
↩︎ Schema information as data: schema, dtypes, columns“A distributed collection of data grouped into named columns.”
↩︎ Key concept“Returns all column names and their data types as a list.”
↩︎ Checkpoint