Databricks Certified Associate Developer for Apache Spark Lessons
32 lessons, one per exam-guide subdomain, in the order the guide teaches them. Every claim is cited to the official documentation.
Domain 1: Apache Spark Architecture and Components
7 lessons · 22% of the exam
1.1Spark Advantages: Unified Engine, In-Memory Processing and Fault Tolerance
Subdomain 1.1: Identify the advantages and challenges of implementing Spark.
2 pages · 19 min read
1.2Spark Cluster Architecture: Driver, Worker Nodes and Executors
Subdomain 1.2: Identify the role of core components of Apache SparkTM's Architecture, including cluster, driver node, worker nodes/executors, CPU cores, and memory.
2 pages · 17 min read
1.3Spark Architecture: Driver, Executors, SparkSession and DataFrames
Subdomain 1.3: Describe the architecture of Apache SparkTM, including DataFrame and Dataset concepts, SparkSession lifecycle, caching, storage levels, and garbage collection.
2 pages · 20 min read
1.4Spark Execution Hierarchy: Application, Job, Stage, Task
Subdomain 1.4: Explain the Apache SparkTM Architecture execution hierarchy.
14 min read
1.5Spark Partitions, Shuffles and Shuffle Partition Settings
Subdomain 1.5: Configure Spark partitioning in distributed data processing, including shuffles and partitions
2 pages · 20 min read
1.6Spark Transformations, Actions and Lazy Evaluation
Subdomain 1.6: Describe the execution patterns of the Apache SparkTM engine, including actions, transformations, and lazy evaluation.
2 pages · 18 min read
1.7Spark Core, Spark SQL and DataFrames: The Unified Engine
Subdomain 1.7: Identify the features of the Apache Spark Modules, including Core, Spark SQL, DataFrames, Pandas API on Spark, Structured Streaming, and MLib.
2 pages · 19 min read
Domain 2: Using Spark SQL
4 lessons · 12% of the exam
2.1Spark File I/O: Formats, Save Modes and partitionBy
Subdomain 2.1: Utilize common data sources such as JDBC, files, etc., to efficiently read from and write to Spark DataFrames using Spark SQL, including overwriting and partitioning by column.
2 pages · 22 min read
2.2Spark SQL on Files: read_files, Path Queries and Save Modes
Subdomain 2.2: Execute SQL queries directly on files, including ORC Files, JSON Files, CSV Files, Text Files, and Delta Files, and understand the different save modes for outputting data in Spark SQL.
15 min read
2.3saveAsTable with partitionBy, bucketBy and sortBy
Subdomain 2.3: Save data to persistent tables while applying sorting and partitioning to optimize data retrieval.
14 min read
2.4Spark SQL Temporary Views: createOrReplaceTempView and Global Temp Views
Subdomain 2.4: Register DataFrames as temporary views in Spark SQL, allowing them to be queried with SQL syntax.
14 min read
Domain 3: Developing Apache SparkTM DataFrame/DataSet API Applications
10 lessons · 31% of the exam
3.1Add, Rename and Drop DataFrame Columns in PySpark
Subdomain 3.1: Manipulate columns, rows, and table structures by adding, dropping, splitting, renaming column names, applying filters, and exploding arrays.
2 pages · 19 min read
3.2Deduplicating Spark DataFrames: distinct, dropDuplicates and Watermarks
Subdomain 3.2: Perform data deduplication and validation operations on DataFrames.
2 pages · 22 min read
3.3Spark DataFrame Aggregations: count, mean, approx_count_distinct, summary
Subdomain 3.3: Perform aggregate operations on DataFrames such as count, approximate count distinct, and mean, summary.
15 min read
3.4Unix Epoch to Date String in PySpark: from_unixtime, unix_timestamp, to_date
Subdomain 3.4: Manipulate and utilize Date data type, such as Unix epoch to date string, and extract date component.
2 pages · 18 min read
3.5Spark DataFrame Joins and Unions: Inner, Left, Broadcast, Cross, Union
Subdomain 3.5: Combine DataFrames with operations such as Inner join, left join, broadcast join, multiple keys, cross join, union, and union all.
16 min read
3.6Reading and Writing DataFrames with Schemas and Save Modes
Subdomain 3.6: Manage input and output operations by writing, overwriting, and reading DataFrames with schemas.
17 min read
3.7Sorting DataFrames and Printing Schemas in PySpark
Subdomain 3.7: Perform operations on DataFrames such as sorting, iterating, printing schema, and conversion between DataFrame and sequence/list formats.
2 pages · 18 min read
3.8Python UDFs and pandas UDFs in PySpark: create, register, invoke
Subdomain 3.8: Create and invoke user-defined functions with or without stateful operators, including StateStores.
2 pages · 21 min read
3.9Spark Shared Variables: Broadcast Variables and Accumulators
Subdomain 3.9: Describe different types of variables in Spark, including broadcast variables and accumulators.
13 min read
3.10Broadcast Joins in Spark: broadcast(), Hints and autoBroadcastJoinThreshold
Subdomain 3.10: Describe the purpose and implementation of broadcast joins
14 min read
Domain 4: Troubleshooting and Tuning Apache Spark DataFrame API Applications
3 lessons · 9% of the exam
4.1Spark repartition vs coalesce: controlling partitions and shuffles
Subdomain 4.1: Implement performance tuning strategies & optimize cluster utilization, including partitioning, repartitioning, coalescing, identifying data skew, and reducing shuffling
2 pages · 18 min read
4.2Adaptive Query Execution: How Spark Re-optimizes Queries at Runtime
Subdomain 4.2: Describe Adaptive Query Execution (AQE) and its benefits.
2 pages · 20 min read
4.3Spark Driver Logs and Executor Logs on Databricks
Subdomain 4.3: Perform logging and monitoring of Spark applications - publish, customize, and analyze Driver logs and Executor logs to diagnose out-of-memory errors, cluster underutilization, etc.
2 pages · 18 min read
Domain 5: Structured Streaming
4 lessons · 12% of the exam
5.1Structured Streaming Programming Model and Micro-Batch Triggers
Subdomain 5.1: Explain the Structured Streaming engine in Spark, including its functions, programming model, micro-batch processing, exactly-once semantics, and fault tolerance mechanisms.
2 pages · 19 min read
5.2Create and Write Streaming DataFrames: readStream, writeStream and Output Modes
Subdomain 5.2: Create and write Streaming DataFrames and Streaming Datasets, including the basic output modes and output sinks.
2 pages · 19 min read
5.3Streaming DataFrames: Selection, Projection and Aggregation
Subdomain 5.3: Perform basic operations on Streaming DataFrames and Streaming Datasets, such as selection, projection, window and aggregation.
2 pages · 19 min read
5.4Streaming dropDuplicates: Deduplication With and Without a Watermark
Subdomain 5.4: Perform Streaming Deduplication in Structured Streaming, both with and without watermark usage.
2 pages · 18 min read
Domain 6: Using Spark Connect to deploy applications
2 lessons · 6% of the exam
6.1Spark Connect Architecture: gRPC, Unresolved Plans and Thin Clients
Subdomain 6.1: Describe the features of Spark Connect.
2 pages · 20 min read
6.2Spark Deploy Modes: Client, Cluster and Local
Subdomain 6.2: Describe the different deployment mode types (Client, Cluster, Local) in the Apache SparkTM environment.
12 min read
Domain 7: Using Pandas API on Spark
2 lessons · 6% of the exam