CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 3 · Lesson 12/32

    Add, Rename and Drop DataFrame Columns in PySpark

    Manipulate columns, rows, and table structures by adding, dropping, splitting, renaming column names, applying filters, and exploding arrays.

    10 min read
    3.12% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Add a new column or replace an existing one with withColumn, and know when select is the better choice
    • Rename one column with withColumnRenamed, several with withColumnsRenamed, or rename while projecting with alias and selectExpr
    • Drop columns by name or by Column expression and predict what happens when the column does not exist

    Key concept

    Transformations return a new DataFrame — withColumn, withColumnRenamed, withColumnsRenamed and drop never change the DataFrame you call them on. Each one returns a new DataFrame with a different set of columns, so you have to keep the return value (or chain the next call onto it), or the change is lost.

    1.Adding or replacing a column with withColumn

    A DataFrame's columns are its structure, and most of that structure gets built with one method. withColumn(colName, col) takes a string name and a Column expression and returns a new DataFrame. If no column has that name, the column is added at the end. If one does, it is replaced. The expression can use existing columns, as in the example below, where age2 is computed from age.

    withColumn adds age2, computed from the existing age columnpython
    df = spark.createDataFrame([(2, "Alice"), (5, "Bob")], schema=["age", "name"])
    df.withColumn('age2', df.age + 2).show()
    # +---+-----+----+
    # |age| name|age2|
    # +---+-----+----+
    # |  2|Alice|   4|
    # |  5|  Bob|   7|
    # +---+-----+----+

    Because a matching name means replacement, withColumn is also how you change a column in place, for example to convert its type. The Databricks basics guide reuses the name c_custkey, so the integer column is overwritten by its string cast instead of sitting next to it. The expression can just as well be a comparison. The same guide makes a boolean balance_flag with col("c_acctbal") > 1000.

    Reusing the existing name makes withColumn replace the column instead of adding onepython
    df_casted = df_customer.withColumn("c_custkey", col("c_custkey").cast(StringType()))

    Checkpoint 1 of 5· Check yourself

    df_customer has an integer column c_custkey. After df_customer.withColumn("c_custkey", col("c_custkey").cast(StringType())), what happens to c_custkey?

    withColumn adds a projection internally, so calling it again and again stacks projections on top of each other. The reference page says this can cause performance problems and even a StackOverflowException, and it gives the fix directly: "To avoid this, use select with multiple columns at once." One withColumn is fine. A loop that generates dozens of them is the pattern to avoid.

    Sources12

    2.Renaming columns: withColumnRenamed, withColumnsRenamed and alias

    withColumnRenamed(existing, new) takes two strings, the current name and the new one. If the existing name isn't in the schema, it does nothing and raises no error. That is convenient in pipelines that may run on slightly different schemas, but it also means a typo in the old name fails silently. The column keeps its original name, and you only find out later.

    Renaming a column that does not exist leaves the DataFrame unchangedpython
    df.withColumnRenamed("non_existing", "new_name").show()
    # +---+-----+
    # |age| name|
    # +---+-----+
    # |  2|Alice|
    # |  5|  Bob|
    # +---+-----+

    You can chain withColumnRenamed calls to rename several columns. The plural form, withColumnsRenamed(colsMap), does it in one call. It takes a Python dict that maps existing names to new ones, and it is also a no-op for any key that is not in the schema. The reference notes that only a single map is supported at the moment.

    Checkpoint 2 of 5· Fill the gap

    Which method renames both columns in a single call?

    df. ? ({"age": "age2", "name": "name2"}).show()

    You can also rename while you project. alias gives a name to a computed Column, and the Databricks guide singles it out for naming aggregation results, such as avg(...).alias("avg_account_balance"). selectExpr takes SQL expressions, so you rename with an as clause, as below.

    selectExpr renames with SQL 'as' clauses while it projectspython
    df_customer.selectExpr(
      "c_custkey as key",
      "round(c_acctbal) as account_rounded"
    )

    Sources342

    3.Dropping columns, by name or by Column

    There are two ways to remove columns. You can leave them out of a select, or you can call drop. drop(*cols) accepts one or more names, and like the rename methods it is a no-op for names that are not in the schema.

    drop accepts several column names at oncepython
    df_customer_flag_renamed.drop("c_phone", "balance_flag_renamed")

    Each argument can be a string or a Column, and the reference says the two are handled differently. A string is treated literally as a column name. A Column is matched as an expression. So drop('age') and drop(df.age) are not the same operation, even though both remove age from a simple DataFrame. The difference shows up after a join, when two columns can share a name. In the reference example below, dropping by the string 'name' removes the name column from both inputs.

    After a join, drop('name') removes the name column from both sidespython
    df2 = spark.createDataFrame([(80, "Tom"), (85, "Bob")], ["height", "name"])
    df.join(df2, df.name == df2.name).drop('name').sort('age').show()
    # +---+------+
    # |age|height|
    # +---+------+
    # | 14|    80|
    # | 16|    85|
    # +---+------+

    When you need to point at one specific DataFrame's column, for example after a join, write df["col"] or df.col. The dot form has a limit. It can't reach columns whose names start with a digit or contain a space or special character, so bracket notation is the safer default.

    Checkpoint 3 of 5· Check yourself

    A DataFrame has the columns id and name. What does df.drop('nmae') (note the typo) return?

    Column-structure methods covered on this page
    MethodWhat it changesArguments
    withColumnAdds a column, or replaces the one with the same namea name string and a Column expression
    withColumnRenamedRenames one column (no-op if the column is missing)the existing name and the new name
    withColumnsRenamedRenames several columns in one callone dict of existing to new names
    dropRemoves one or more columns (no-op if a column is missing)column names or Column objects
    aliasNames a computed Column, often an aggregatethe new name

    Checkpoint 4 of 5· Match them up

    Match each call to its effect

    Tap a term, then the definition that fits it.

    Checkpoint 5 of 5· Exam question

    An analyst has a DataFrame `orders` with columns `order_id`, `amount`, and `region`. They need to add a new column `amount_tier` that labels rows as `"high"` when `amount` is greater than 1000 and `"standard"` otherwise, without mutating the original DataFrame: ```python orders_tiered = orders.____( "amount_tier", when(col("amount") > 1000, "high").otherwise("standard") ) ``` Which method correctly completes this task?

    Sources52

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.withColumnRenamed raises an error if the old column name does not exist.Why is that wrong?

      It is a no-op. It returns a DataFrame with the same columns, so a misspelled source name fails silently.

      Covered in Renaming columns: withColumnRenamed, withColumnsRenamed and alias

    2. 2.drop('col') and drop(col('col')) always do exactly the same thing.Why is that wrong?

      A string is treated literally as a name, and a Column is matched as an expression. The documentation describes these as different semantics, which matters when names are ambiguous, for example after a join.

      Covered in Dropping columns, by name or by Column

    3. 3.Calling withColumn in a loop is the recommended way to add many columns.Why is that wrong?

      Each call adds a projection, so a loop builds a large plan. Use a single select with all the new columns instead.

      Covered in Adding or replacing a column with withColumn

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Returns a new DataFrame by adding a column or replacing the existing column that has the same name.”
      ↩︎ Adding or replacing a column with withColumn
      “To avoid this, use select with multiple columns at once.”
      ↩︎ Adding or replacing a column with withColumn
      “Returns a new DataFrame by adding a column or replacing the existing column that has the same name.”
      ↩︎ Key concept
      “To avoid this, use select with multiple columns at once.”
      ↩︎ Exam trap 3
      “calling it multiple times, for instance, via loops in order to add multiple columns can generate big plans”
      ↩︎ Prediction
    2. 2.
      “To create a new column, use the withColumn method.”
      ↩︎ Adding or replacing a column with withColumn
      “The alias method is especially helpful when you want to rename your columns as part of aggregations”
      ↩︎ Renaming columns: withColumnRenamed, withColumnsRenamed and alias
      “To remove columns, you can omit columns during a select or select(*) except or you can use the drop method”
      ↩︎ Dropping columns, by name or by Column
      “The . operator cannot be used to select columns starting with an integer, or ones that contain a space or special character.”
      ↩︎ Dropping columns, by name or by Column
    3. 3.
      “Returns a new DataFrame by renaming an existing column. This is a no-op if the schema doesn't contain the given column name.”
      ↩︎ Renaming columns: withColumnRenamed, withColumnsRenamed and alias
      “This is a no-op if the schema doesn't contain the given column name.”
      ↩︎ Exam trap 1
    4. 4.
      “A dict of existing column names and corresponding desired column names. Currently, only a single map is supported.”
      ↩︎ Renaming columns: withColumnRenamed, withColumnsRenamed and alias
      “A dict of existing column names and corresponding desired column names. Currently, only a single map is supported.”
      ↩︎ Checkpoint
    5. 5.
      “So dropping a column by its name drop(colName) has a different semantic with directly dropping the column drop(col(colName)).”
      ↩︎ Dropping columns, by name or by Column
      “So dropping a column by its name drop(colName) has a different semantic with directly dropping the column drop(col(colName)).”
      ↩︎ Exam trap 2
      “Returns a new DataFrame without specified columns. This is a no-op if the schema doesn't contain the given column name(s).”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Filter Rows, Split Strings and Explode Arrays in PySpark

    Spotted a mistake, or was something unclear? Tell us.