CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 1 · Lesson 5/32

    Repartition vs Coalesce: Controlling DataFrame Partitions

    Configure Spark partitioning in distributed data processing, including shuffles and partitions

    9 min read
    3.12% of exam
    6 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Use repartition() to hash-partition a DataFrame by a partition count, by columns, or both
    • Use repartitionByRange() and explain why its output can differ from run to run
    • Choose between coalesce() and repartition() based on whether a shuffle is acceptable
    • Write the same partitioning intent as Spark SQL hints
    • Tell DataFrameWriter.partitionBy() apart from in-memory repartitioning

    1.repartition() and repartitionByRange()

    Configuration settings shape partitions across a whole session. The DataFrame API lets you set the partitioning of one DataFrame at a specific point in your code. repartition() returns a new DataFrame that is hash partitioned. Its first argument can be an integer target count or a column. If you pass a column there, Spark uses it as the first partitioning column. When you give no number, Spark uses the default number of partitions. The RDD guide states the cost directly: repartition always shuffles all data over the network.

    repartition(10) produces partition ids 0–9; repartition(7, "age") hash-partitions by age into 7 partitionspython
    df.repartition(10).select(
        sf.spark_partition_id().alias("partition")
    ).distinct().sort("partition").show()

    Checkpoint 1 of 5· Fill the gap

    Which method hash-partitions this DataFrame into 7 partitions by the age column?

    df. ? (7, "age").select(
        sf.spark_partition_id().alias("partition")
    ).distinct().sort("partition").show()

    repartitionByRange() has the same signature but produces a range partitioned DataFrame. Rows are assigned by value ranges of the columns you name, and you must name at least one column. If you specify no sort order, Spark assumes ascending with nulls first. In the example below, ages 14 and 16 land in partition 0 and age 23 in partition 1.

    repartitionByRange(2, "age") groups neighbouring age values into the same partitionpython
    spark.createDataFrame(
        [(14, "Tom"), (23, "Alice"), (16, "Bob")], ["age", "name"]
    ).repartitionByRange(2, "age").select(
        "age", "name", sf.spark_partition_id()
    ).show()

    Sources123

    2.coalesce(): fewer partitions without a shuffle

    coalesce(numPartitions) takes only a number, with no columns. It creates a narrow dependency: when you go from 1000 partitions to 100, each new partition simply claims 10 of the existing ones, and no shuffle happens. That makes it the cheap way to reduce partitions, for example after a filter has made a large dataset much smaller. It can only reduce the count, though. Asking for more partitions leaves the DataFrame at its current number.

    The narrow dependency also has a downside. Because no shuffle separates the upstream work from the coalesced partitions, a drastic coalesce such as coalesce(1) can push the computation onto fewer nodes than you want, in that case one node. The reference suggests calling repartition() instead. It adds a shuffle step, but the upstream partitions still run in parallel.

    Checkpoint 2 of 5· Fill the gap

    This sample collapses a 3-partition range into a single partition without a shuffle. Which method fills the blank?

    from pyspark.sql import functions as sf
    spark.range(0, 10, 1, 3). ? (1).select(
        sf.spark_partition_id().alias("partition")
    ).distinct().sort("partition").show()
    coalesce() vs repartition() on a DataFrame
    Behaviourcoalesce(n)repartition(n, *cols)
    ShuffleNo: narrow dependency, new partitions claim existing onesYes: adds a shuffle step
    Requesting more partitions than existStays at the current number of partitionsSupported (e.g. 9 input partitions repartitioned to 10)
    Partitioning columnsNot accepted, only numPartitionsOptional; result is hash partitioned
    Drastic reduction (e.g. to 1)Work may run on fewer nodes than you likeUpstream partitions still execute in parallel

    Checkpoint 3 of 5· Exam question

    By default, before any Adaptive Query Execution coalescing takes place, how many partitions does Spark use for the shuffle produced by a `groupBy` or `join` operation, per the `spark.sql.shuffle.partitions` configuration?

    Sources4

    3.The same controls in SQL: partitioning hints

    SQL users get the same controls through hints. Spark SQL's coalesce hints work just like coalesce, repartition and repartitionByRange in the Dataset API. They are used for performance tuning and to reduce the number of output files. The hints differ in which arguments they accept:

    - COALESCE takes only a partition number. - REPARTITION takes a partition number, columns, both, or neither. - REPARTITION_BY_RANGE must have column names, and the partition number is optional. - REBALANCE takes an initial partition number, columns, both, or neither.

    Partitioning hints in Spark SQLsql
    SELECT /*+ COALESCE(3) */ * FROM t;
    SELECT /*+ REPARTITION(3) */ * FROM t;
    SELECT /*+ REPARTITION(c) */ * FROM t;
    SELECT /*+ REPARTITION(3, c) */ * FROM t;
    SELECT /*+ REPARTITION */ * FROM t;
    SELECT /*+ REPARTITION_BY_RANGE(c) */ * FROM t;
    SELECT /*+ REPARTITION_BY_RANGE(3, c) */ * FROM t;

    Checkpoint 4 of 5· Check yourself

    Which of these hints is NOT valid according to the hint rules?

    Sources53

    4.Not to be confused: DataFrameWriter.partitionBy()

    The word *partition* also shows up when you write data. DataFrameWriter.partitionBy(*cols) partitions the output on the file system by the given columns. It lays the files out in a Hive-style scheme with one directory per value, such as name=Alice. Later reads can then target a single directory. This is a storage layout. It does not change how many in-memory partitions or tasks the DataFrame has. That is what repartition() and coalesce() control.

    Checkpoint 5 of 5· Check yourself

    A colleague writes df.write.partitionBy("name").parquet(path) to cut the number of tasks in the next stage. What does partitionBy actually do?

    Sources6

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.coalesce(n) can raise the partition count when n is larger than the current count. It just avoids the shuffle.Why is that wrong?

      coalesce only merges existing partitions. Asking for more leaves the DataFrame at its current partition count, so use repartition to increase it.

      Covered in coalesce(): fewer partitions without a shuffle

    2. 2.coalesce(1) is always better than repartition(1) because it avoids a shuffle.Why is that wrong?

      A drastic coalesce can push the upstream computation onto very few nodes. repartition(1) adds a shuffle but keeps the upstream partitions running in parallel.

      Covered in coalesce(): fewer partitions without a shuffle

    3. 3.repartitionByRange always produces the same partition boundaries for the same data.Why is that wrong?

      The ranges are estimated by sampling, so the output may differ between runs.

      Covered in repartition() and repartitionByRange()

    Practise it for real

    See repartition, repartitionByRange and coalesce change a DataFrame's partitions by inspecting spark_partition_id()

    1. 1.Build the reference DataFrame: from pyspark.sql import functions as sf; df = spark.range(0, 64, 1, 9) with the name and age columns added as in the repartition reference example.

      Why: spark.range's fourth argument fixes the starting partition count at 9, so every later change is visible.

      You should see: A 64-row DataFrame with columns id, name and age, split across 9 partitions.

    2. 2.Run df.repartition(10).select(sf.spark_partition_id().alias("partition")).distinct().sort("partition").show()

      Why: repartition shuffles the rows into a new hash-partitioned layout and can raise the count.

      You should see: Partition ids 0 through 9.

    3. 3.Run spark.range(0, 10, 1, 3).coalesce(1) with the same spark_partition_id() select, distinct and sort.

      Why: coalesce merges the 3 existing partitions through a narrow dependency.

      You should see: A single row: partition 0.

    4. 4.Run the repartitionByRange(2, "age") example on the Tom/Alice/Bob DataFrame and show sf.spark_partition_id().

      Why: Range partitioning puts neighbouring age values in the same partition.

      You should see: Ages 14 and 16 in partition 0 and age 23 in partition 1, as in the reference output.

    Stuck? Get a nudge

    Try df.coalesce(20) on the 9-partition DataFrame and count the distinct partition ids. Is the result what the coalesce reference predicts?

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “The resulting DataFrame is hash partitioned.”
      ↩︎ repartition() and repartitionByRange()
      “If not specified, the default number of partitions is used.”
      ↩︎ repartition() and repartitionByRange()
    2. 3.
      “At least one partition-by expression must be specified.”
      ↩︎ repartition() and repartitionByRange()
      “The resulting DataFrame is range partitioned.”
      ↩︎ The same controls in SQL: partitioning hints
      “the output may not be consistent, since sampling can return different values”
      ↩︎ Exam trap 3
    3. 4.
      “there will not be a shuffle, instead each of the 100 new partitions will claim 10 of the current partitions”
      ↩︎ coalesce(): fewer partitions without a shuffle
      “This will add a shuffle step, but means the current upstream partitions will be executed in parallel (per whatever the current partitioning is).”
      ↩︎ coalesce(): fewer partitions without a shuffle
      “If a larger number of partitions is requested, it will stay at the current number of partitions.”
      ↩︎ Exam trap 1
      “this may result in your computation taking place on fewer nodes than you like”
      ↩︎ Exam trap 2
      “If a larger number of partitions is requested, it will stay at the current number of partitions.”
      ↩︎ Prediction
    4. 5.
      “Coalesce hints allow Spark SQL users to control the number of output files just like coalesce, repartition and repartitionByRange in the Dataset API”
      ↩︎ The same controls in SQL: partitioning hints
      “hint only has a partition number as a parameter.”
      ↩︎ Checkpoint
    5. 6.
      “If specified, the output is laid out on the file system similar to Hive's partitioning scheme.”
      ↩︎ Not to be confused: DataFrameWriter.partitionBy()
      “Partitions the output by the given columns on the file system.”
      ↩︎ Checkpoint

    Spotted a mistake, or was something unclear? Tell us.