CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 1 · Lesson 1/32

    Spark Challenges: Shuffles, Memory Pressure, Resource Limits and Workload Fit

    Identify the advantages and challenges of implementing Spark.

    10 min read
    3.12% of exam
    8 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why shuffles are the most expensive operations in a Spark workload
    • Recognize spill, skew and garbage-collection pressure as the main memory-related challenges and describe their symptoms
    • Identify the limits that come from Spark's cluster architecture, such as application isolation, driver memory and cluster cost
    • Decide when a workload fits Spark poorly, for example task-parallel work or Python UDFs that replace native functions

    1.The shuffle: the price of distributing data

    Spark is fast because each task works on one partition independently. That breaks down when an operation needs data that lives in many partitions. To combine all the values for one key, for example, 'Spark needs to perform an all-to-all operation'. It reads from every partition and brings matching values together. That movement of data is called the shuffle.

    A shuffle 'typically involves copying data across executors and machines, making the shuffle a complex and costly operation'. Join operations, most ByKey operations and repartitioning can all trigger one. Shuffles also stress memory. Databricks notes that the memory used for shuffles, joins, sorts and aggregations is where trouble usually starts. When that memory runs out, data moves to disk, and this 'is most common during data shuffling'. The next section covers that memory pressure.

    Shuffles also add a tuning burden. The right number of shuffle partitions depends on the data, and Databricks notes that data sizes 'may differ vastly from stage to stage, query to query, making this number hard to tune'. The Spark guide likewise says shuffle behavior is tuned by adjusting a variety of configuration parameters.

    Checkpoint 1 of 6· Check yourself

    A colleague says a groupByKey is cheap because Spark is an in-memory engine. What is the best correction?

    Sources123

    2.Memory pressure: spill, skew and garbage collection

    Running in memory is an advantage only while the data fits. Databricks defines the first failure mode directly: 'Spill is what happens when Spark runs low on execution memory.' Spark then moves data from memory to disk, which can be expensive. The RDD guide describes the same effect for shuffle data structures: 'Spark will spill these tables to disk, incurring the additional overhead of disk I/O and increased garbage collection.' So memory pressure costs you twice, once in disk I/O and again in JVM garbage collection.

    Shuffles also leave intermediate files on disk. Spark keeps them until the related RDDs are garbage collected, and that can take a long time if the application still holds references. The guide warns that 'long-running Spark jobs may consume a large amount of disk space'.

    The second failure mode is skew. 'Skew is when one or just a few tasks take much longer than the rest.' Because a stage isn't finished until its slowest task is, skew 'results in poor cluster utilization and longer jobs': most executors sit idle while one works through an oversized partition.

    How memory-related challenges show up
    ChallengeWhat is happeningWhat you see
    SpillSpark runs low on execution memory and moves data to diskSpill stats appear in the stage details
    SkewOne or a few tasks take much longer than the restMax task duration well above the 75th percentile
    Garbage collectionSpilled shuffle tables add JVM collection overheadThe RDD guide lists increased garbage collection as an extra cost of spilling, on top of disk I/O

    Checkpoint 2 of 6· Check yourself

    In a stage's summary metrics, the 75th-percentile task duration is 40 seconds and the Max is 3 minutes. What is the most likely problem?

    The advantage that memory pressure undermines is in-memory reuse. Spark lets you persist a dataset in memory, 'allowing it to be reused efficiently across parallel operations', and that 'allows future actions to be much faster (often by more than 10x)'. Caching has a cost, though. Each node keeps the partitions it computes in memory, and Spark drops old partitions in least-recently-used fashion when the cache fills. Databricks also warns that manually caching data in production pipelines 'can interrupt these optimizations and lead to increases in cost and latency'. Caching is a trade-off, not a free speedup.

    Checkpoint 3 of 6· Exam question

    An engineer migrates a nightly aggregation pipeline from disk-based MapReduce to Spark and observes a large runtime improvement on a job that repeatedly scans and re-scans the same intermediate dataset across several stages. Which characteristic of Spark's design explains this speedup?

    Sources214

    3.Limits that come from the cluster architecture

    Some challenges come from how a Spark application is built. Each application gets its own executor processes for its whole lifetime. 'This has the benefit of isolating applications from each other', since tasks from different applications run in different JVMs. The trade-off is that 'data cannot be shared across different Spark applications (instances of SparkContext) without writing it to an external storage system'. Cluster managers such as Standalone, YARN and Kubernetes allocate resources across applications, so concurrent applications compete for the same cluster resources and can't hand data to each other in memory.

    The driver brings constraints of its own. It schedules every task, so it must stay reachable: 'the driver program must be network addressable from the worker nodes', and it should run close to them, preferably on the same local area network. Its memory is also limited. Databricks warns that pulling results back from the cluster has to stay small: 'Only collect small amounts of data back to R data frames, or the Spark driver will run out of memory.'

    Checkpoint 4 of 6· Check yourself

    Two separate Spark applications on the same cluster need to use the same intermediate result. What must happen?

    Sources567

    4.When Spark is the wrong tool

    The last challenge is choosing Spark for work it doesn't suit. Spark is built for data parallelism and is recommended for large-scale data processing such as joins, filtering and aggregation. Work that consists of many independent computations is task parallelism, and Databricks points to Ray for it: 'Ray is designed for task parallelism, where multiple tasks run concurrently and independently.'

    Workload fit: Spark versus Ray
    WorkloadBetter fit
    ETL, analytics reporting, feature engineering, data preprocessingSpark
    Large-scale machine learning with MLlibSpark
    Reinforcement learning, simulation modeling, hyperparameter searchRay
    Deep learning training and high-performance computing (HPC)Ray

    Even inside Spark, some code works against the engine. Python UDFs that replace a native function are the usual example. 'Serialization is required to transfer data between Python and Spark. This significantly slows down queries.' Use native functions where they exist. When a Python UDF is unavoidable, Databricks recommends Pandas UDFs, where Apache Arrow moves data efficiently between Spark and Python.

    What Spark does handle for you on suitable workloads is recovery from failures. The RDD guide says 'RDDs automatically recover from node failures', and that a lost cached partition 'will automatically be recomputed using the transformations that originally created it'. Replicated storage levels let tasks keep running without waiting for that recomputation.

    Checkpoint 5 of 6· Match them up

    Match each situation to the recommendation the documentation gives

    Tap a term, then the definition that fits it.

    Checkpoint 6 of 6· Exam question

    During a long-running Spark job, one executor is terminated after the underlying VM is reclaimed by the cloud provider. The job continues and eventually finishes with correct results. Which two statements correctly describe how Spark achieves this fault tolerance? (Choose 2 answers)(Select 2)

    Sources871

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Spark is an in-memory engine, so a Spark job never touches disk during processing.Why is that wrong?

      When Spark runs low on execution memory it spills data to disk, most often during shuffles, which adds disk I/O and extra garbage collection.

      Covered in Memory pressure: spill, skew and garbage collection

    2. 2.Applications on the same Spark cluster share executors, so one application can reuse another's in-memory data.Why is that wrong?

      Each application has its own isolated executor processes. Data passes between applications only through external storage.

      Covered in Limits that come from the cluster architecture

    3. 3.Spark is the best choice for every distributed Python workload.Why is that wrong?

      Spark excels at data parallelism. Task-parallel workloads such as hyperparameter search or simulation suit Ray better.

      Covered in When Spark is the wrong tool

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “The Shuffle is an expensive operation since it involves disk I/O, data serialization, and network I/O.”
      ↩︎ The shuffle: the price of distributing data
      “This typically involves copying data across executors and machines, making the shuffle a complex and costly operation.”
      ↩︎ The shuffle: the price of distributing data
      “Shuffle behavior can be tuned by adjusting a variety of configuration parameters.”
      ↩︎ The shuffle: the price of distributing data
      “Spark will spill these tables to disk, incurring the additional overhead of disk I/O and increased garbage collection.”
      ↩︎ Memory pressure: spill, skew and garbage collection
      “This means that long-running Spark jobs may consume a large amount of disk space.”
      ↩︎ Memory pressure: spill, skew and garbage collection
      “allowing it to be reused efficiently across parallel operations”
      ↩︎ Memory pressure: spill, skew and garbage collection
      “This allows future actions to be much faster (often by more than 10x).”
      ↩︎ Memory pressure: spill, skew and garbage collection
      “Finally, RDDs automatically recover from node failures.”
      ↩︎ When Spark is the wrong tool
      “if any partition of an RDD is lost, it will automatically be recomputed using the transformations that originally created it.”
      ↩︎ When Spark is the wrong tool
    2. 2.
      “It starts to move data from memory to disk, which can be expensive. It is most common during data shuffling.”
      ↩︎ The shuffle: the price of distributing data
      “Spill is what happens when Spark runs low on execution memory.”
      ↩︎ Memory pressure: spill, skew and garbage collection
      “Skew is when one or just a few tasks take much longer than the rest.”
      ↩︎ Memory pressure: spill, skew and garbage collection
      “This results in poor cluster utilization and longer jobs.”
      ↩︎ Memory pressure: spill, skew and garbage collection
      “Spill is what happens when Spark runs low on execution memory.”
      ↩︎ Exam trap 1
      “If the Max duration is 50% more than the 75th percentile, you may be suffering from skew.”
      ↩︎ Checkpoint
    3. 3.
      “data sizes may differ vastly from stage to stage, query to query, making this number hard to tune”
      ↩︎ The shuffle: the price of distributing data
    4. 4.
      “Manually caching data or returning preview results in production pipelines can interrupt these optimizations and lead to increases in cost and latency.”
      ↩︎ Memory pressure: spill, skew and garbage collection
    5. 5.
      “This has the benefit of isolating applications from each other”
      ↩︎ Limits that come from the cluster architecture
      “the driver program must be network addressable from the worker nodes.”
      ↩︎ Limits that come from the cluster architecture
      “data cannot be shared across different Spark applications (instances of SparkContext) without writing it to an external storage system.”
      ↩︎ Exam trap 2
      “data cannot be shared across different Spark applications (instances of SparkContext) without writing it to an external storage system.”
      ↩︎ Checkpoint
    6. 6.
      “Only collect small amounts of data back to R data frames, or the Spark driver will run out of memory.”
      ↩︎ Limits that come from the cluster architecture
    7. 7.
      “So, if you spin up two worker clusters and it takes an hour, you are paying for those workers for the full hour.”
      ↩︎ Limits that come from the cluster architecture
      “an autoscaling cluster is usually the cheapest, but not necessarily the fastest.”
      ↩︎ Limits that come from the cluster architecture
      “Serialization is required to transfer data between Python and Spark. This significantly slows down queries.”
      ↩︎ When Spark is the wrong tool
    8. 8.
      “Ray is designed for task parallelism, where multiple tasks run concurrently and independently.”
      ↩︎ When Spark is the wrong tool
      “Ray is designed for task parallelism, where multiple tasks run concurrently and independently.”
      ↩︎ Exam trap 3
      “Use Ray for workloads where Spark is less optimized, such as reinforcement learning, hierarchical time series forecasting, simulation modeling, hyperparameter search”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 11 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.