What you will be able to do
- Name the three debugging surfaces for a Spark application and the three kinds of compute logging Databricks provides
- Say what lands in the driver logs and when to read them
- Navigate from a misbehaving task to the executor log4j output that explains it
- Explain who can read driver and executor logs under each access mode, and how to publish logs and metrics beyond the live compute
Key concept
Driver logs vs executor logs — The driver logs hold what your notebook, job and library code prints or logs, plus the stack traces of exceptions raised while starting work. Executor logs hold the log4j output of the worker that ran a specific task, so you read them when particular tasks are misbehaving.
1.Where Spark tells you what went wrong
When a Spark application fails or runs slowly, Databricks gives you three places to look: the Spark UI, the driver logs and the executor logs. The Spark UI shows jobs, stages, tasks and executors while the application runs. The two kinds of log hold the text output that explains why. Logs are not the only record, though. Databricks says it "provides three kinds of logging of compute-related activity", and each one answers a different question.
| Log | What it captures | Where you open it |
|---|---|---|
| Compute event log | Compute lifecycle events such as creation, termination and configuration edits | Event log tab on the compute details page |
| Apache Spark driver and worker log | Output you use for debugging the application itself | Driver logs tab (driver); Spark UI tab (workers) |
| Compute init-script logs | Output of init scripts | See Init script logging |
The event log is about the compute as a whole: it records events that are triggered manually by users or automatically by Databricks, and keeps them for 60 days. It answers questions like "was the cluster resized?" The driver and worker logs answer "what did my code do?" Keeping those two questions apart is the first step in any diagnosis.
Checkpoint 1 of 5· Check yourself
A colleague wants to know whether someone edited the cluster configuration shortly before a job started failing. Which log should they open first?
Configuration edits are compute lifecycle events, and the compute event log records those. Driver and executor logs record application output.
“Compute event logs, which capture compute lifecycle events like creation, termination, and configuration edits.”Source: docs.databricks.com
Sources1
2.Driver logs: your code's output and its exceptions
Whatever your notebooks, jobs and libraries print or log goes to the Spark driver logs. You open them from the Driver logs tab on the compute details page, and clicking a log file's name downloads it. The driver logs are split into three outputs: Standard output, Standard error and Log4j logs. The debugging guide adds that print statements that are part of the DAG appear in these logs too.
The main reason to open the driver logs is to find exceptions. If a job never got going, for example a streaming query that failed at startup so no Streaming tab appears in the Spark UI, the cause is in the driver logs: "You can drill into the Driver logs to look at the stack trace of the exception." The same applies when a job did start but its work keeps failing or never completes. The driver logs show what the underlying problem is.
Checkpoint 2 of 5· Match them up
Match each driver-log symptom or output to what it tells you
Tap a term, then the definition that fits it.
The driver logs have three outputs (standard output, standard error, Log4j). When a job fails to start, the exception's stack trace is in the driver logs.
“You can drill into the Driver logs to look at the stack trace of the exception.”Source: docs.databricks.com
3.Executor logs: following one misbehaving task
The driver doesn't see everything. Tasks run on executors, and when particular tasks misbehave (they fail, retry or run far longer than their siblings), their output is on the worker that ran them. To get there, you first find the executor in the Spark UI and then go to that worker's log. The task details page tells you which executor ran each task. From the compute UI page you click the number of nodes, then the master, which lists every worker. Pick the worker where the suspicious task ran and open its log4j output.
Checkpoint 3 of 5· Put it in order
Put these steps in order to reach the executor log for a suspicious task
- 1.On the task details page, note the executor where the task ran
- 2.Click the # nodes link
- 3.Choose the worker where the suspicious task ran and open its log4j output
- 4.Click the master to see the list of all workers
- 5.Go to the compute UI page
The executor ID comes from the task details page. Then you go from the compute UI to the nodes, the master's worker list and finally that worker's log4j output.
“You can choose the worker where the suspicious task was run and then get to the log4j output.”Source: docs.databricks.com
The driver log holds the driver's output and exceptions. The log4j output written while task 47 ran stays on the worker that executed it. You need the executor ID from the task details page to know which worker's log to open.
Sources2
4.Publishing logs: access modes, log delivery and external export
Who can read logs depends on the compute's access mode, and this is easy to overlook when you plan a debugging workflow.
| Access mode | Driver logs | Executor logs |
|---|---|---|
| Standard | Only workspace admins | Not available |
| Dedicated | The dedicated user or group, and workspace admins | The dedicated user or group, and workspace admins |
There is also the question of what survives. If you restart a terminated compute, its Spark UI shows the restarted compute, not the history of the run you wanted to investigate. To keep logs, you publish them: you can configure a log delivery location for the compute, and both worker and compute logs are delivered there. The operational-excellence guidance recommends that you "Enable cluster log delivery to persist Spark event logs to cloud storage for historical analysis". You can then analyze the event logs for long-running stages, data skew, shuffle operations and memory pressure.
For metrics rather than logs, the same guidance says to configure Spark metrics to export to external monitoring systems. One documented example installs Datadog agents on compute nodes through a compute-scoped init script, which you can apply to all compute with a compute policy. Note what the sources here do not cover: changing log4j log levels or log formats. They describe customization only as where logs are delivered and where metrics are exported.
Checkpoint 4 of 5· Exam question
A data engineer runs the following PySpark job on a Databricks all-purpose cluster and the driver fails with `java.lang.OutOfMemoryError: Java heap space`. The `transactions` table produces roughly two million distinct `region` values after the aggregation below, and the driver is provisioned with 8 GB of memory: ```python df = spark.read.parquet("/mnt/sales/transactions") results = df.groupBy("region").agg(sum("amount").alias("total")) rows = results.collect() for row in rows: print(row) ``` Which change most directly resolves the driver OutOfMemoryError while still letting the engineer inspect the aggregated totals?
Correct answer: A — Write `results` to a Delta table with `results.write.format("delta").mode("overwrite").save(path)`, then read back only a sample of rows with `spark.read.format("delta").load(path).limit(20).show()` instead of collecting everything.
- A. Persisting the aggregated result to storage and then reading back a bounded sample keeps the full two million rows out of the driver heap entirely, which is what a plain `collect()` call cannot avoid regardless of how the DataFrame was partitioned upstream.
- B. The failure happens because `collect()` serializes every row of `results` into the driver's JVM heap in one operation; raising driver memory only postpones the failure to a larger dataset and does nothing to change the underlying pattern of pulling all rows into one process.
- C. Partitioning affects how work is distributed across executors during the shuffle that produces `results`, but `collect()` still gathers every partition's rows into the driver afterward, so repartitioning the aggregated DataFrame does not reduce what lands in driver memory.
- D. Switching to the RDD API changes how the reduction executes on the cluster, but the final `collect()` call in this alternative still pulls the same two million aggregated key-value pairs into the driver, so the out-of-memory condition remains unresolved.
- E. `persist()` controls how a DataFrame's data is cached across executors for reuse in later stages, but it has no effect on what happens when `collect()` is subsequently called — that operation still copies every resulting row into the driver's memory.
Checkpoint 5 of 5· Check yourself
A team investigates a slow job only after the compute has been terminated and restarted. What will they find in the Spark UI, and what should they have set up?
After a restart the Spark UI shows the restarted compute, not the historical run. Configuring log delivery sends worker and compute logs to a location you choose.
“If you restart a terminated compute, the Spark UI displays information for the restarted compute, not the historical information for the terminated compute.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Any user on a cluster can open executor logs to debug their own failed tasks.Why is that wrong?
Executor logs are not available on Standard access mode compute at all. On Dedicated access mode, only the dedicated user or group and workspace admins can read them.
Covered in Publishing logs: access modes, log delivery and external export
2.The Spark UI keeps the history of a terminated compute, so you can investigate after restarting it.Why is that wrong?
After a restart the Spark UI shows the restarted compute. To keep logs for later analysis, configure a log delivery location.
Covered in Publishing logs: access modes, log delivery and external export
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Databricks provides three kinds of logging of compute-related activity:”
↩︎ Where Spark tells you what went wrong“The direct print and log statements from your notebooks, jobs, and libraries go to the Spark driver logs.”
↩︎ Driver logs: your code's output and its exceptions“You can install Datadog agents on compute nodes to send Datadog metrics to your Datadog account.”
↩︎ Publishing logs: access modes, log delivery and external export“You can also configure a log delivery location for the compute. Both worker and compute logs are delivered to the location you specify.”
↩︎ Exam trap 2“Compute event logs, which capture compute lifecycle events like creation, termination, and configuration edits.”
↩︎ Checkpoint“If you restart a terminated compute, the Spark UI displays information for the restarted compute, not the historical information for the terminated compute.”
↩︎ Checkpoint - 2.
“Prints: Any print statements as part of the DAG shows up in the logs too.”
↩︎ Driver logs: your code's output and its exceptions“Executor logs are helpful if you see certain tasks are misbehaving and would like to see the logs for specific tasks.”
↩︎ Executor logs: following one misbehaving task“For compute with Standard access mode, only workspace admins can access driver logs.”
↩︎ Publishing logs: access modes, log delivery and external export“Executor logs are helpful if you see certain tasks are misbehaving and would like to see the logs for specific tasks.”
↩︎ Key concept“Executor logs are not available for compute with Standard access mode.”
↩︎ Exam trap 1“You can drill into the Driver logs to look at the stack trace of the exception.”
↩︎ Checkpoint“You can choose the worker where the suspicious task was run and then get to the log4j output.”
↩︎ Checkpoint“Executor logs are not available for compute with Standard access mode.”
↩︎ Prediction - 3.https://docs.databricks.com/aws/en/lakehouse-architecture/operational-excellence/best-practicesOfficial docs
“Enable cluster log delivery to persist Spark event logs to cloud storage for historical analysis.”
↩︎ Publishing logs: access modes, log delivery and external export“Configure Spark metrics to export to external monitoring systems for centralized observability.”
↩︎ Publishing logs: access modes, log delivery and external export