CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 1 · Lesson 1/32

    Spark Advantages: Unified Engine, In-Memory Processing and Fault Tolerance

    Identify the advantages and challenges of implementing Spark.

    9 min read
    3.12% of exam
    7 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why Spark's single execution engine for SQL, DataFrames, streaming, pandas-style code and machine learning is an advantage
    • Describe how in-memory caching and lazy evaluation speed up Spark workloads, and when manual caching works against you
    • Explain how Spark recovers from lost partitions by recomputing them instead of relying only on stored copies

    Key concept

    Data parallelism — Spark's core strength is running the same operation on every piece of a large dataset at once, with each piece handled in parallel across the cluster. Most of its advantages, and most of its challenges, come from how well a workload fits this model.

    1.One engine for many kinds of workload

    The first advantage to understand is that Spark is a unified engine. Databricks describes it as 'a unified analytics engine for big data and machine learning'. In practice, one engine runs relational queries, streaming, pandas-style code and machine learning, so a team doesn't need a separate system for each job.

    The Spark SQL guide says the same execution engine is used whichever API or language you write in. Because of that, developers can use whichever API fits a given transformation best. The same idea applies to streaming. With Structured Streaming, 'you can express your streaming computation the same way you would express a batch computation on static data', and the Spark SQL engine runs it incrementally as new data arrives. Databricks describes the engine behind its pipelines as having 'a unified architecture for batch and streaming processing'. The engine also reaches beyond SQL. Pandas API on Spark lets you scale a pandas workload by running it distributed across multiple nodes, and MLlib provides scalable machine learning built on Spark.

    Spark modules that share the one engine
    ModuleWhat it gives you
    Spark SQL and DataFramesRelational queries on structured data; SQL can be mixed with Spark programs
    Structured StreamingStreaming computation written the same way as a batch computation on static data
    Pandas API on Sparkpandas data structures and analysis tools that run distributed across multiple nodes
    MLlibA scalable machine learning library with uniform APIs for building and tuning pipelines

    Checkpoint 1 of 4· Match them up

    Match each Spark module to the problem it solves

    Tap a term, then the definition that fits it.

    Checkpoint 2 of 4· Exam question

    A retail company runs separate Hadoop MapReduce jobs for nightly ETL, a dedicated SQL engine for analyst reporting, and a standalone cluster for demand-forecasting models, and wants to consolidate onto one platform without losing performance. Which advantage of Apache Spark most directly supports this consolidation?

    Sources123

    2.In-memory processing and lazy evaluation

    The second advantage is speed from keeping data in memory. The RDD programming guide calls persisting (or caching) a dataset in memory across operations one of Spark's most important capabilities. Each node keeps the partitions it computes and reuses them in later actions. The guide says 'this allows future actions to be much faster (often by more than 10x)' and calls caching 'a key tool for iterative algorithms and fast interactive use'. Databricks gives the general reason: caching stores frequently accessed data in a faster medium, which cuts the time needed to fetch it compared with reading the original source again.

    Lazy evaluation is the other half of Spark's speed. Every operation is either a transformation, which adds logic to a plan, or an action, which triggers evaluation. In Spark, 'none of the logic defined by a collection of operations are evaluated until an action is triggered'. Spark waits for that action so that it can look at the whole chain at once and choose 'the most efficient physical plan to evaluate logic specified by transformations', instead of running each step in the order you wrote it.

    These two points fit together once you ask where the reuse happens. Caching pays off when the same dataset feeds several actions, as in iterative algorithms and fast interactive use, where the cached partitions are read again instead of recomputed. A production pipeline is different: Databricks says writing data is typically the only action that should be present, and manually caching data or returning previews adds extra actions that interrupt the optimization of the whole chain. So the 10x benefit describes repeated reuse in iterative and interactive work, while the Databricks warning describes production pipelines.

    Checkpoint 3 of 4· Check yourself

    Why can adding manual cache and preview calls to a production pipeline make it slower and more expensive?

    Sources4561

    3.Fault tolerance through recomputation

    The third advantage is fault tolerance, and it follows from the lazy, plan-based model described in the previous section. A DataFrame doesn't hold a fixed result. It holds instructions, and 'ultimately Apache Spark resolves queries back to the original data sources, so the data itself is not changed'. As Databricks puts it, 'DataFrames are immutable'. Because Spark always knows how each piece of data was derived, it can rebuild a lost piece from the original source.

    The RDD guide says plainly that 'RDDs automatically recover from node failures'. The same applies to cached data: if any partition is lost, 'it will automatically be recomputed using the transformations that originally created it'. Replicated storage levels add speed, not correctness. They let tasks keep running without waiting for a lost partition to be recomputed, which helps when Spark serves requests that need fast recovery.

    Checkpoint 4 of 4· Check yourself

    What do replicated storage levels add, given that every storage level already recovers lost data by recomputing it?

    A model answer to the reflection: recomputing a lost partition means re-running the transformations that created it, so if those early steps are expensive, the recovery is expensive too. Tasks that need the partition wait until it has been rebuilt. Replicated storage levels avoid that wait, because tasks can keep running on the replica, and that is why they can be worth the extra memory when fast recovery matters. Spark also keeps some work around to reduce the repeat cost. Shuffle files are preserved, 'This is done so the shuffle files don’t need to be re-created if the lineage is re-computed.' The trade-off is that long-running jobs may consume a large amount of disk space, because those files stay until the RDDs are garbage collected.

    Sources14

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Because caching speeds up repeated actions, caching intermediate results and adding previews always makes a production pipeline faster.Why is that wrong?

      Manual caches and previews are actions. They break up Spark's optimization of the whole chain of transformations and can increase both cost and latency.

      Covered in In-memory processing and lazy evaluation

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Databricks is built on top of Apache Spark, a unified analytics engine for big data and machine learning.”
      ↩︎ One engine for many kinds of workload
      “You can express your streaming computation the same way you would express a batch computation on static data”
      ↩︎ One engine for many kinds of workload
      “Spark optimizes data processing by identifying the most efficient physical plan to evaluate logic specified by transformations.”
      ↩︎ In-memory processing and lazy evaluation
      “In production data pipelines, writing data is typically the only action that should be present.”
      ↩︎ In-memory processing and lazy evaluation
      “ultimately Apache Spark resolves queries back to the original data sources, so the data itself is not changed”
      ↩︎ Fault tolerance through recomputation
      “In other words, DataFrames are immutable.”
      ↩︎ Fault tolerance through recomputation
      “Pandas API on Spark allows you to scale your pandas workload to any size by running it distributed across multiple nodes”
      ↩︎ Checkpoint
    2. 2.
      “When computing a result, the same execution engine is used, independent of which API/language you are using to express the computation.”
      ↩︎ One engine for many kinds of workload
    3. 3.
      “has a unified architecture for batch and streaming processing”
      ↩︎ One engine for many kinds of workload
    4. 4.
      “This allows future actions to be much faster (often by more than 10x).”
      ↩︎ In-memory processing and lazy evaluation
      “Caching is a key tool for iterative algorithms and fast interactive use.”
      ↩︎ In-memory processing and lazy evaluation
      “Finally, RDDs automatically recover from node failures.”
      ↩︎ Fault tolerance through recomputation
      “if any partition of an RDD is lost, it will automatically be recomputed using the transformations that originally created it”
      ↩︎ Fault tolerance through recomputation
      “This is done so the shuffle files don’t need to be re-created if the lineage is re-computed.”
      ↩︎ Fault tolerance through recomputation
      “All the storage levels provide full fault tolerance by recomputing lost data”
      ↩︎ Prediction
      “the replicated ones let you continue running tasks on the RDD without waiting to recompute a lost partition”
      ↩︎ Checkpoint
    5. 5.
      “Caching stores frequently accessed data in a faster medium, reducing the time required to retrieve it compared to accessing the original data source.”
      ↩︎ In-memory processing and lazy evaluation
    6. 6.
      “none of the logic defined by a collection of operations are evaluated until an action is triggered”
      ↩︎ In-memory processing and lazy evaluation
      “Manually caching data or returning preview results in production pipelines can interrupt these optimizations and lead to increases in cost and latency.”
      ↩︎ Exam trap 1
      “Manually caching data or returning preview results in production pipelines can interrupt these optimizations and lead to increases in cost and latency.”
      ↩︎ Prediction

    Also cited

    Continue to page 2 of 2

    Spark Challenges: Shuffles, Memory Pressure, Resource Limits and Workload Fit

    Spotted a mistake, or was something unclear? Tell us.