What you will be able to do
- Add a new column or replace an existing one with withColumn, and know when select is the better choice
- Rename one column with withColumnRenamed, several with withColumnsRenamed, or rename while projecting with alias and selectExpr
- Drop columns by name or by Column expression and predict what happens when the column does not exist
Key concept
Transformations return a new DataFrame — withColumn, withColumnRenamed, withColumnsRenamed and drop never change the DataFrame you call them on. Each one returns a new DataFrame with a different set of columns, so you have to keep the return value (or chain the next call onto it), or the change is lost.
1.Adding or replacing a column with withColumn
A DataFrame's columns are its structure, and most of that structure gets built with one method. withColumn(colName, col) takes a string name and a Column expression and returns a new DataFrame. If no column has that name, the column is added at the end. If one does, it is replaced. The expression can use existing columns, as in the example below, where age2 is computed from age.
df = spark.createDataFrame([(2, "Alice"), (5, "Bob")], schema=["age", "name"])
df.withColumn('age2', df.age + 2).show()
# +---+-----+----+
# |age| name|age2|
# +---+-----+----+
# | 2|Alice| 4|
# | 5| Bob| 7|
# +---+-----+----+Because a matching name means replacement, withColumn is also how you change a column in place, for example to convert its type. The Databricks basics guide reuses the name c_custkey, so the integer column is overwritten by its string cast instead of sitting next to it. The expression can just as well be a comparison. The same guide makes a boolean balance_flag with col("c_acctbal") > 1000.
df_casted = df_customer.withColumn("c_custkey", col("c_custkey").cast(StringType()))Checkpoint 1 of 5· Check yourself
df_customer has an integer column c_custkey. After df_customer.withColumn("c_custkey", col("c_custkey").cast(StringType())), what happens to c_custkey?
When the name passed to withColumn matches an existing column, that column is replaced in the new DataFrame that withColumn returns. Nothing is added alongside it, and the original DataFrame is not modified.
“Returns a new DataFrame by adding a column or replacing the existing column that has the same name.”Source: docs.databricks.com
withColumn adds a projection internally, so calling it again and again stacks projections on top of each other. The reference page says this can cause performance problems and even a StackOverflowException, and it gives the fix directly: "To avoid this, use select with multiple columns at once." One withColumn is fine. A loop that generates dozens of them is the pattern to avoid.
2.Renaming columns: withColumnRenamed, withColumnsRenamed and alias
withColumnRenamed(existing, new) takes two strings, the current name and the new one. If the existing name isn't in the schema, it does nothing and raises no error. That is convenient in pipelines that may run on slightly different schemas, but it also means a typo in the old name fails silently. The column keeps its original name, and you only find out later.
df.withColumnRenamed("non_existing", "new_name").show()
# +---+-----+
# |age| name|
# +---+-----+
# | 2|Alice|
# | 5| Bob|
# +---+-----+You can chain withColumnRenamed calls to rename several columns. The plural form, withColumnsRenamed(colsMap), does it in one call. It takes a Python dict that maps existing names to new ones, and it is also a no-op for any key that is not in the schema. The reference notes that only a single map is supported at the moment.
Checkpoint 2 of 5· Fill the gap
Which method renames both columns in a single call?
df. ? ({"age": "age2", "name": "name2"}).show()withColumnsRenamed takes one dict of existing-to-new names. withColumnRenamed takes two strings and renames a single column.
Source: docs.databricks.comYou can also rename while you project. alias gives a name to a computed Column, and the Databricks guide singles it out for naming aggregation results, such as avg(...).alias("avg_account_balance"). selectExpr takes SQL expressions, so you rename with an as clause, as below.
df_customer.selectExpr(
"c_custkey as key",
"round(c_acctbal) as account_rounded"
)3.Dropping columns, by name or by Column
There are two ways to remove columns. You can leave them out of a select, or you can call drop. drop(*cols) accepts one or more names, and like the rename methods it is a no-op for names that are not in the schema.
df_customer_flag_renamed.drop("c_phone", "balance_flag_renamed")Each argument can be a string or a Column, and the reference says the two are handled differently. A string is treated literally as a column name. A Column is matched as an expression. So drop('age') and drop(df.age) are not the same operation, even though both remove age from a simple DataFrame. The difference shows up after a join, when two columns can share a name. In the reference example below, dropping by the string 'name' removes the name column from both inputs.
df2 = spark.createDataFrame([(80, "Tom"), (85, "Bob")], ["height", "name"])
df.join(df2, df.name == df2.name).drop('name').sort('age').show()
# +---+------+
# |age|height|
# +---+------+
# | 14| 80|
# | 16| 85|
# +---+------+When you need to point at one specific DataFrame's column, for example after a join, write df["col"] or df.col. The dot form has a limit. It can't reach columns whose names start with a digit or contain a space or special character, so bracket notation is the safer default.
Checkpoint 3 of 5· Check yourself
A DataFrame has the columns id and name. What does df.drop('nmae') (note the typo) return?
drop is a no-op for names that are not in the schema, so the typo causes no error and removes nothing.
“Returns a new DataFrame without specified columns. This is a no-op if the schema doesn't contain the given column name(s).”Source: docs.databricks.com
| Method | What it changes | Arguments |
|---|---|---|
| withColumn | Adds a column, or replaces the one with the same name | a name string and a Column expression |
| withColumnRenamed | Renames one column (no-op if the column is missing) | the existing name and the new name |
| withColumnsRenamed | Renames several columns in one call | one dict of existing to new names |
| drop | Removes one or more columns (no-op if a column is missing) | column names or Column objects |
| alias | Names a computed Column, often an aggregate | the new name |
Checkpoint 4 of 5· Match them up
Match each call to its effect
Tap a term, then the definition that fits it.
withColumn adds or replaces, the two rename methods change names (one or many), and drop removes columns. A string passed to drop is treated as a literal name.
“A dict of existing column names and corresponding desired column names. Currently, only a single map is supported.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
An analyst has a DataFrame `orders` with columns `order_id`, `amount`, and `region`. They need to add a new column `amount_tier` that labels rows as `"high"` when `amount` is greater than 1000 and `"standard"` otherwise, without mutating the original DataFrame: ```python orders_tiered = orders.____( "amount_tier", when(col("amount") > 1000, "high").otherwise("standard") ) ``` Which method correctly completes this task?
Correct answer: C — withColumn
- A. select projects a set of column expressions but does not take a target column name and a value expression as two positional arguments the way this call is written, so this signature does not fit.
- B. withColumnRenamed only accepts an existing column name and a new name string; it cannot evaluate a when/otherwise expression, so it cannot produce a computed column.
- C. withColumn takes a target column name and a Column expression, adding the result as a new column while returning a new DataFrame, which matches this call exactly.
- D. filter restricts which rows are kept based on a boolean condition; it does not add or compute a new column, so it does not fit this signature.
- E. withColumnsRenamed takes a dictionary mapping existing names to new names for bulk renaming, not a name plus a computed expression, so it does not match this call.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.withColumnRenamed raises an error if the old column name does not exist.Why is that wrong?
It is a no-op. It returns a DataFrame with the same columns, so a misspelled source name fails silently.
Covered in Renaming columns: withColumnRenamed, withColumnsRenamed and alias
2.drop('col') and drop(col('col')) always do exactly the same thing.Why is that wrong?
A string is treated literally as a name, and a Column is matched as an expression. The documentation describes these as different semantics, which matters when names are ambiguous, for example after a join.
Covered in Dropping columns, by name or by Column
3.Calling withColumn in a loop is the recommended way to add many columns.Why is that wrong?
Each call adds a projection, so a loop builds a large plan. Use a single select with all the new columns instead.
Covered in Adding or replacing a column with withColumn
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Returns a new DataFrame by adding a column or replacing the existing column that has the same name.”
↩︎ Adding or replacing a column with withColumn“To avoid this, use select with multiple columns at once.”
↩︎ Adding or replacing a column with withColumn“Returns a new DataFrame by adding a column or replacing the existing column that has the same name.”
↩︎ Key concept“To avoid this, use select with multiple columns at once.”
↩︎ Exam trap 3“calling it multiple times, for instance, via loops in order to add multiple columns can generate big plans”
↩︎ Prediction - 2.https://docs.databricks.com/aws/en/pyspark/basicsOfficial docs
“To create a new column, use the withColumn method.”
↩︎ Adding or replacing a column with withColumn“The alias method is especially helpful when you want to rename your columns as part of aggregations”
↩︎ Renaming columns: withColumnRenamed, withColumnsRenamed and alias“To remove columns, you can omit columns during a select or select(*) except or you can use the drop method”
↩︎ Dropping columns, by name or by Column“The . operator cannot be used to select columns starting with an integer, or ones that contain a space or special character.”
↩︎ Dropping columns, by name or by Column - 3.https://docs.databricks.com/aws/en/pyspark/reference/classes/dataframe/withColumnRenamedOfficial docs
“Returns a new DataFrame by renaming an existing column. This is a no-op if the schema doesn't contain the given column name.”
↩︎ Renaming columns: withColumnRenamed, withColumnsRenamed and alias“This is a no-op if the schema doesn't contain the given column name.”
↩︎ Exam trap 1 - 4.https://docs.databricks.com/aws/en/pyspark/reference/classes/dataframe/withColumnsRenamedOfficial docs
“A dict of existing column names and corresponding desired column names. Currently, only a single map is supported.”
↩︎ Renaming columns: withColumnRenamed, withColumnsRenamed and alias“A dict of existing column names and corresponding desired column names. Currently, only a single map is supported.”
↩︎ Checkpoint - 5.
“So dropping a column by its name drop(colName) has a different semantic with directly dropping the column drop(col(colName)).”
↩︎ Dropping columns, by name or by Column“So dropping a column by its name drop(colName) has a different semantic with directly dropping the column drop(col(colName)).”
↩︎ Exam trap 2“Returns a new DataFrame without specified columns. This is a no-op if the schema doesn't contain the given column name(s).”
↩︎ Checkpoint