What you will be able to do
- Explain that the deploy mode decides where the Spark driver process runs
- Tell client mode apart from cluster mode by where the driver is launched, and know what each one means for networking
- Describe local mode as Spark running in a single process with no worker nodes, and recognise it on Databricks as single node compute
- Explain why a Spark Connect or Databricks Connect client is not the driver, even when your code runs on your own machine
Key concept
Deploy mode (driver location) — The deploy mode answers one question: where does the driver process run? In client mode the driver runs where you submitted the application, outside the cluster. In cluster mode it runs inside the cluster. In local mode the driver and the executors share a single process on one machine.
1.Every deploy mode answers one question: where is the driver?
A Spark application has two kinds of process. The driver runs your main program and holds the SparkContext. It turns your DataFrame code into jobs, stages and tasks, and schedules those tasks. The executors run on worker nodes, do the actual computation and keep data in memory or on disk. Databricks describes its own clusters in the same terms: a driver node and zero or more worker nodes, which Databricks also calls executor nodes. The driver node keeps the state of any attached notebook and holds the SparkContext.
Once you have this split in mind, the three deploy modes are easy to place. The executors almost always live on the cluster, so the modes differ in where the driver lives. It can run on the machine that submitted the application (client), on a node inside the cluster (cluster), or in the same process as the executors on one machine (local). Before reading further, commit to an answer below.
| Term | Meaning |
|---|---|
| Driver program | The process running the main() function of the application and creating the SparkContext |
| Executor | A process launched for an application on a worker node, that runs tasks and keeps data in memory or disk storage |
| Deploy mode | Distinguishes where the driver process runs |
Checkpoint 1 of 7· Check yourself
According to the Spark glossary, which process runs the application's main() function and creates the SparkContext?
The driver program is the process that runs main() and creates the SparkContext. Executors only run the tasks the driver sends them.
“The process running the main() function of the application and creating the SparkContext”Source: spark.apache.org
2.Client mode and cluster mode
In client mode, the process that submits the application also starts the driver, outside the cluster. This could be your laptop, an edge node or a gateway machine. The executors still run on the cluster's worker nodes, but the scheduling decisions come from the submitting machine. Output the driver prints, such as the result of collect(), shows up right where you are working. That is why client mode suits interactive work, such as shells, notebooks and step-by-step exploration. The exam guide links client mode with interactive clusters for exactly this reason.
Client mode has a networking cost. The driver must accept connections from its executors for as long as it runs, so the workers must be able to reach it over the network. Because the driver schedules every task, it should run close to the workers, ideally on the same local network. A driver on a laptop far from the cluster, behind a VPN, can turn every scheduling round trip into a slow link.
In cluster mode, the cluster framework starts the driver on a node inside the cluster. Once the application is handed over, the submitting machine has no further role and can disconnect. The application keeps running even if the person who submitted it closes their laptop. This makes cluster mode the natural fit for scheduled, unattended production work. The exam guide links it with job clusters. On Databricks, the cluster definition reflects the same idea: a cluster has one Spark driver plus a number of executors, so the driver is one of the cluster's own nodes.
| Deploy mode | Who launches the driver | Where the driver runs |
|---|---|---|
| client | The submitter | Outside of the cluster |
| cluster | The framework | Inside of the cluster |
Checkpoint 2 of 7· Check yourself
A data engineer starts an application from their workstation in client mode, and the executors run on a remote cluster. Which statement is correct?
In client mode the submitter launches the driver outside the cluster, here on the workstation. The executors are still on the cluster and must be able to connect back to the driver.
“In "client" mode, the submitter launches the driver outside of the cluster.”Source: spark.apache.org
Checkpoint 3 of 7· Exam question
A data engineer submits a PySpark job to an existing standalone Spark cluster from their local machine using the following command, and the submission fails immediately: ``` spark-submit --master spark://cluster-host:7077 --deploy-mode cluster app.py ``` What is the most likely reason cluster deploy mode is rejected here?
Correct answer: A — The standalone cluster manager does not support cluster deploy mode for Python applications, so the driver cannot be launched on a worker node this way.
- A. The Spark standalone cluster manager does not support cluster deploy mode for Python applications, only for Java and Scala jars, so launching a `.py` file with `--deploy-mode cluster` against a standalone master is rejected. This is exactly the situation described, making it the correct explanation.
- B. A bad hostname would surface as a connection or DNS resolution error rather than a deploy-mode restriction, and nothing in the scenario suggests the hostname itself is invalid. This distractor misattributes the failure to network resolution instead of the actual mode restriction.
- C. The `--supervise` flag only controls automatic driver restart behavior in cluster deploy mode and has no bearing on whether cluster mode is accepted in the first place. Omitting it does not force a fallback to client mode.
- D. This scenario uses a standalone master URL, not YARN, so a YARN-specific restriction does not apply here. It also misdescribes how `--py-files` and deploy mode interact, since that flag manages dependency shipping, not mode eligibility.
- E. Spark does not require a distributed filesystem path for the application file in cluster deploy mode on every cluster manager, and standalone mode can distribute a local file to the driver node in supported cases. This restriction is not what causes the described failure.
3.Local mode: one machine, one process
Local mode removes the cluster from the picture. The driver and the execution of tasks share one JVM on one machine, and parallelism comes from threads instead of separate executor nodes. There is no cluster manager and nothing to reach over the network. You choose local mode through the master setting rather than the deploy-mode setting. In the Spark programming guide, the master is a cluster URL or the special value "local". The guide recommends not hardcoding the master in production code. Instead, pass it in through spark-submit, and keep "local" for tests.
conf = SparkConf().setAppName(appName).setMaster(master)
sc = SparkContext(conf=conf)Checkpoint 4 of 7· Fill the gap
Which method sets the value that can switch this application into local mode?
conf = SparkConf().setAppName(appName). ? (master)
sc = SparkContext(conf=conf)Local mode is selected through the master value, for example by passing "local" to setMaster. Deploy mode only chooses between client and cluster when a real cluster is involved.
Source: spark.apache.orgLocal mode also exists on Databricks. Single node compute has one driver node and no worker nodes, and Spark runs in local mode on that driver. This setup is cheaper for small workloads that don't need distributed processing. When the data outgrows one machine, you move to multi-node compute, with one driver and one or more workers.
| Compute type | Nodes | How Spark runs |
|---|---|---|
| Single node cluster | One driver node, no worker nodes | Local mode on the driver |
| Distributed cluster | One driver node and one or more worker nodes | Tasks run in parallel across worker nodes |
Checkpoint 5 of 7· Exam question
A developer opens a Databricks notebook attached to an all-purpose (interactive) cluster and runs cells one at a time, checking the output of each transformation before writing the next cell. Which description of the deployment mode behind this session is accurate?
Correct answer: A — This matches client deploy mode: the driver stays reachable to the submitting session for the whole interaction, so results from each cell return immediately for inspection.
- A. Interactive, cell-by-cell notebook sessions on an all-purpose cluster map to client deploy mode, where the driver remains a live, reachable process throughout the session so each cell's output can be inspected before the next runs. This matches the incremental workflow described.
- B. Shipping the driver to a worker node independent of the session describes cluster deploy mode, which is built for unattended runs rather than interactive, cell-by-cell inspection. That does not fit a workflow where the user actively checks each cell's result.
- C. Local mode runs an entire job inside a single JVM with no cluster involvement at all, but this scenario explicitly attaches to a cluster and relies on its executors. The notebook session does contact the cluster, so local mode does not describe it.
- D. Autoscaler settings control how many worker nodes a cluster has, not which deployment mode a notebook session uses. Toggling autoscaling does not change the location of the driver process.
- E. All-purpose clusters are long-running and shared across a session, unlike ephemeral job clusters that spin up and tear down per job run. Databricks does not create a fresh job cluster for every notebook cell, so this describes a different kind of compute entirely.
4.Where Spark Connect fits: your code is local, the driver is not
This domain is about Spark Connect, so it helps to see how it fits with the deploy modes. With classic client mode, running code on your laptop meant the driver ran on your laptop too. Spark Connect breaks that link. The client application is a separate process from the Spark driver and sends unresolved plans to a server over gRPC. The driver runs with the server, close to the executors.
Databricks Connect, which is built on Spark Connect, splits the work clearly. Your general Python or Scala code runs on your machine, so you can debug it in your IDE. DataFrame operations are turned into Spark plans and run on Databricks compute. Results only come back to your machine when you call something like collect(), show() or toPandas(). So you get the interactive feel of client mode while the driver stays next to the workers. That placement is exactly what the Spark cluster overview recommends.
Checkpoint 6 of 7· Match them up
With Databricks Connect, match each kind of work to where it runs
Tap a term, then the definition that fits it.
Only the Spark work goes to the remote compute. Your own control-flow code stays local, and results come back only when an action brings them to the client.
“All code is executed locally, while all Spark code continues to run on the remote cluster.”Source: docs.databricks.com
Checkpoint 7 of 7· Exam question
A production pipeline is scheduled as a Databricks Job that spins up a new job cluster for each run, executes a PySpark script with no notebook attached, and terminates the cluster once the script finishes. Which deployment mode characteristic applies to this run, and why?
Correct answer: A — Cluster deploy mode applies: the driver process runs on the job cluster itself, so the run does not depend on the machine or session that originally triggered the job.
- A. Unattended Databricks Jobs running on ephemeral job clusters run the driver on the cluster itself, matching cluster deploy mode, so the job's lifecycle is self-contained on that cluster once triggered. This fits a run with no notebook attached and no ongoing client session.
- B. The control plane schedules and triggers the job, but it does not act as a live client holding the driver process for the run's duration; that would defeat the purpose of an unattended, self-contained job cluster. This misplaces where the driver actually executes.
- C. The absence of a notebook does not mean Spark skips provisioning executors; a job cluster still allocates worker nodes and runs the script across them like any other Spark application. Local mode is a separate, explicit configuration, not a default for notebook-less runs.
- D. The `--supervise` flag governs automatic driver restart behavior within cluster deploy mode and is unrelated to whether a Databricks Job cluster run uses cluster or client deploy mode in the first place. Its absence does not force a different mode.
- E. Ephemeral job clusters are specifically designed so the driver runs on the cluster rather than externally, which is what lets the run continue independent of the triggering session. Describing the driver as external contradicts the self-contained nature of a job cluster run.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.In cluster deploy mode, the driver runs on the machine that submitted the job.Why is that wrong?
That describes client mode. In cluster mode the framework starts the driver inside the cluster, so the submitting machine can disconnect once the application is handed over.
Covered in Client mode and cluster mode
2.In client mode, it doesn't matter where the submitting machine is, because the executors do all the work.Why is that wrong?
The driver schedules every task and must accept connections from its executors, so it has to be reachable from the workers and should run close to them.
Covered in Client mode and cluster mode
3.Local mode still needs at least one worker node to run executors.Why is that wrong?
Local mode runs Spark in a single process on one machine. On Databricks, single node compute has a driver and no worker nodes, and Spark runs in local mode.
Covered in Local mode: one machine, one process
4.A Spark Connect client on your laptop is the same as running the driver in client mode.Why is that wrong?
A Spark Connect client is a separate process from the driver. It sends plans to a server, and the driver runs there.
Covered in Where Spark Connect fits: your code is local, the driver is not
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://spark.apache.org/docs/latest/cluster-overview.htmlSecondary source
“Distinguishes where the driver process runs.”
↩︎ Every deploy mode answers one question: where is the driver?“The process running the main() function of the application and creating the SparkContext”
↩︎ Every deploy mode answers one question: where is the driver?“In "client" mode, the submitter launches the driver outside of the cluster.”
↩︎ Client mode and cluster mode“Because the driver schedules tasks on the cluster, it should be run close to the worker nodes, preferably on the same local area network.”
↩︎ Client mode and cluster mode“Distinguishes where the driver process runs.”
↩︎ Key concept“In "cluster" mode, the framework launches the driver inside of the cluster.”
↩︎ Exam trap 1“the driver program must be network addressable from the worker nodes”
↩︎ Exam trap 2“In "cluster" mode, the framework launches the driver inside of the cluster.”
↩︎ Prediction - 2.https://docs.databricks.com/aws/en/sparkrOfficial docs
“Databricks clusters consist of an Apache Spark driver node and zero or more Spark worker (also known as executor) nodes.”
↩︎ Every deploy mode answers one question: where is the driver?“The driver node maintains attached notebook state, maintains the SparkContext, interprets notebook and library commands”
↩︎ Every deploy mode answers one question: where is the driver?“A single node cluster has one driver node and no worker nodes, with Spark running in local mode”
↩︎ Local mode: one machine, one process“A single node cluster has one driver node and no worker nodes, with Spark running in local mode”
↩︎ Exam trap 3 - 3.
“A cluster has one Spark Driver and num_workers Executors for a total of num_workers + 1 Spark nodes.”
↩︎ Client mode and cluster mode - 4.https://spark.apache.org/docs/latest/rdd-programming-guide.htmlSecondary source
“for local testing and unit tests, you can pass “local” to run Spark in-process”
↩︎ Local mode: one machine, one process“you will not want to hardcode master in the program, but rather launch the application with spark-submit”
↩︎ Local mode: one machine, one process - 5.https://spark.apache.org/docs/latest/spark-connect-overview.htmlSecondary source
“The client does not run in the same process as the Spark driver.”
↩︎ Where Spark Connect fits: your code is local, the driver is not“The client does not run in the same process as the Spark driver.”
↩︎ Exam trap 4 - 6.
“General code runs locally: Python and Scala code runs on the client side, enabling interactive debugging.”
↩︎ Where Spark Connect fits: your code is local, the driver is not“All code is executed locally, while all Spark code continues to run on the remote cluster.”
↩︎ Where Spark Connect fits: your code is local, the driver is not