CertSafari

    Databricks Certified Associate Developer for Apache Spark Lessons

    32 lessons, one per exam-guide subdomain, in the order the guide teaches them. Every claim is cited to the official documentation.

    Domain 1: Apache Spark Architecture and Components

    7 lessons · 22% of the exam

    1. 1.1Spark Advantages: Unified Engine, In-Memory Processing and Fault Tolerance

      Subdomain 1.1: Identify the advantages and challenges of implementing Spark.

      2 pages · 19 min read

    2. 1.2Spark Cluster Architecture: Driver, Worker Nodes and Executors

      Subdomain 1.2: Identify the role of core components of Apache SparkTM's Architecture, including cluster, driver node, worker nodes/executors, CPU cores, and memory.

      2 pages · 17 min read

    3. 1.3Spark Architecture: Driver, Executors, SparkSession and DataFrames

      Subdomain 1.3: Describe the architecture of Apache SparkTM, including DataFrame and Dataset concepts, SparkSession lifecycle, caching, storage levels, and garbage collection.

      2 pages · 20 min read

    4. 1.4Spark Execution Hierarchy: Application, Job, Stage, Task

      Subdomain 1.4: Explain the Apache SparkTM Architecture execution hierarchy.

      14 min read

    5. 1.5Spark Partitions, Shuffles and Shuffle Partition Settings

      Subdomain 1.5: Configure Spark partitioning in distributed data processing, including shuffles and partitions

      2 pages · 20 min read

    6. 1.6Spark Transformations, Actions and Lazy Evaluation

      Subdomain 1.6: Describe the execution patterns of the Apache SparkTM engine, including actions, transformations, and lazy evaluation.

      2 pages · 18 min read

    7. 1.7Spark Core, Spark SQL and DataFrames: The Unified Engine

      Subdomain 1.7: Identify the features of the Apache Spark Modules, including Core, Spark SQL, DataFrames, Pandas API on Spark, Structured Streaming, and MLib.

      2 pages · 19 min read

    Domain 2: Using Spark SQL

    4 lessons · 12% of the exam

    1. 2.1Spark File I/O: Formats, Save Modes and partitionBy

      Subdomain 2.1: Utilize common data sources such as JDBC, files, etc., to efficiently read from and write to Spark DataFrames using Spark SQL, including overwriting and partitioning by column.

      2 pages · 22 min read

    2. 2.2Spark SQL on Files: read_files, Path Queries and Save Modes

      Subdomain 2.2: Execute SQL queries directly on files, including ORC Files, JSON Files, CSV Files, Text Files, and Delta Files, and understand the different save modes for outputting data in Spark SQL.

      15 min read

    3. 2.3saveAsTable with partitionBy, bucketBy and sortBy

      Subdomain 2.3: Save data to persistent tables while applying sorting and partitioning to optimize data retrieval.

      14 min read

    4. 2.4Spark SQL Temporary Views: createOrReplaceTempView and Global Temp Views

      Subdomain 2.4: Register DataFrames as temporary views in Spark SQL, allowing them to be queried with SQL syntax.

      14 min read

    Domain 3: Developing Apache SparkTM DataFrame/DataSet API Applications

    10 lessons · 31% of the exam

    1. 3.1Add, Rename and Drop DataFrame Columns in PySpark

      Subdomain 3.1: Manipulate columns, rows, and table structures by adding, dropping, splitting, renaming column names, applying filters, and exploding arrays.

      2 pages · 19 min read

    2. 3.2Deduplicating Spark DataFrames: distinct, dropDuplicates and Watermarks

      Subdomain 3.2: Perform data deduplication and validation operations on DataFrames.

      2 pages · 22 min read

    3. 3.3Spark DataFrame Aggregations: count, mean, approx_count_distinct, summary

      Subdomain 3.3: Perform aggregate operations on DataFrames such as count, approximate count distinct, and mean, summary.

      15 min read

    4. 3.4Unix Epoch to Date String in PySpark: from_unixtime, unix_timestamp, to_date

      Subdomain 3.4: Manipulate and utilize Date data type, such as Unix epoch to date string, and extract date component.

      2 pages · 18 min read

    5. 3.5Spark DataFrame Joins and Unions: Inner, Left, Broadcast, Cross, Union

      Subdomain 3.5: Combine DataFrames with operations such as Inner join, left join, broadcast join, multiple keys, cross join, union, and union all.

      16 min read

    6. 3.6Reading and Writing DataFrames with Schemas and Save Modes

      Subdomain 3.6: Manage input and output operations by writing, overwriting, and reading DataFrames with schemas.

      17 min read

    7. 3.7Sorting DataFrames and Printing Schemas in PySpark

      Subdomain 3.7: Perform operations on DataFrames such as sorting, iterating, printing schema, and conversion between DataFrame and sequence/list formats.

      2 pages · 18 min read

    8. 3.8Python UDFs and pandas UDFs in PySpark: create, register, invoke

      Subdomain 3.8: Create and invoke user-defined functions with or without stateful operators, including StateStores.

      2 pages · 21 min read

    9. 3.9Spark Shared Variables: Broadcast Variables and Accumulators

      Subdomain 3.9: Describe different types of variables in Spark, including broadcast variables and accumulators.

      13 min read

    10. 3.10Broadcast Joins in Spark: broadcast(), Hints and autoBroadcastJoinThreshold

      Subdomain 3.10: Describe the purpose and implementation of broadcast joins

      14 min read

    Domain 4: Troubleshooting and Tuning Apache Spark DataFrame API Applications

    3 lessons · 9% of the exam

    1. 4.1Spark repartition vs coalesce: controlling partitions and shuffles

      Subdomain 4.1: Implement performance tuning strategies & optimize cluster utilization, including partitioning, repartitioning, coalescing, identifying data skew, and reducing shuffling

      2 pages · 18 min read

    2. 4.2Adaptive Query Execution: How Spark Re-optimizes Queries at Runtime

      Subdomain 4.2: Describe Adaptive Query Execution (AQE) and its benefits.

      2 pages · 20 min read

    3. 4.3Spark Driver Logs and Executor Logs on Databricks

      Subdomain 4.3: Perform logging and monitoring of Spark applications - publish, customize, and analyze Driver logs and Executor logs to diagnose out-of-memory errors, cluster underutilization, etc.

      2 pages · 18 min read

    Domain 5: Structured Streaming

    4 lessons · 12% of the exam

    1. 5.1Structured Streaming Programming Model and Micro-Batch Triggers

      Subdomain 5.1: Explain the Structured Streaming engine in Spark, including its functions, programming model, micro-batch processing, exactly-once semantics, and fault tolerance mechanisms.

      2 pages · 19 min read

    2. 5.2Create and Write Streaming DataFrames: readStream, writeStream and Output Modes

      Subdomain 5.2: Create and write Streaming DataFrames and Streaming Datasets, including the basic output modes and output sinks.

      2 pages · 19 min read

    3. 5.3Streaming DataFrames: Selection, Projection and Aggregation

      Subdomain 5.3: Perform basic operations on Streaming DataFrames and Streaming Datasets, such as selection, projection, window and aggregation.

      2 pages · 19 min read

    4. 5.4Streaming dropDuplicates: Deduplication With and Without a Watermark

      Subdomain 5.4: Perform Streaming Deduplication in Structured Streaming, both with and without watermark usage.

      2 pages · 18 min read

    Domain 6: Using Spark Connect to deploy applications

    2 lessons · 6% of the exam

    1. 6.1Spark Connect Architecture: gRPC, Unresolved Plans and Thin Clients

      Subdomain 6.1: Describe the features of Spark Connect.

      2 pages · 20 min read

    2. 6.2Spark Deploy Modes: Client, Cluster and Local

      Subdomain 6.2: Describe the different deployment mode types (Client, Cluster, Local) in the Apache SparkTM environment.

      12 min read

    Domain 7: Using Pandas API on Spark

    2 lessons · 6% of the exam

    1. 7.1Pandas API on Spark: Advantages of pyspark.pandas

      Subdomain 7.1: Explain the advantages of using Pandas API on Spark.

      13 min read

    2. 7.2Pandas UDFs: creating with pandas_udf and invoking on DataFrames

      Subdomain 7.2: Create and invoke Pandas UDF.

      2 pages · 20 min read