What you will be able to do
- Name the places Databricks uses Spark Connect, and describe Databricks Connect as an extension of it
- State what runs locally and what runs on remote compute in a Spark Connect client
- Contrast eager analysis in Spark Classic with deferred analysis in Spark Connect
- Avoid the temp-view and UDF pitfalls that come from deferred name resolution
1.Where Spark Connect runs, and Databricks Connect
Spark Connect is an open-source, gRPC-based protocol in Apache Spark. It lets a client run Spark workloads remotely through the DataFrame API. On Databricks it powers more surfaces than many candidates expect:
- Scala notebooks on Databricks Runtime 13.3 and above, on standard compute - Python notebooks on Databricks Runtime 14.3 and above, on standard compute - serverless compute - Databricks Connect - Lakeflow pipelines with environment versions configured
Databricks Connect is the client library for connecting IDEs, notebooks and custom applications to Databricks compute. For Databricks Runtime 13.3 LTS and above, it is an extension of Spark Connect, with additions and modifications to support Databricks compute modes and Unity Catalog. You create sessions with DatabricksSession instead of SparkSession. Databricks Connect can also run against an open-source Spark Connect server, but code that uses Databricks-only features may fail there.
For exam questions, the most useful thing to know is the split between work done locally and work done remotely:
- General code runs locally. Your Python or Scala code runs on the client, so you can debug it interactively.
- DataFrame operations run on Databricks compute. Transformations become Spark plans that run remotely. Results only come back to the client when you call collect(), show() or toPandas().
- UDF code also runs on Databricks compute. UDFs you define locally are serialized and sent to the cluster. So application dependencies are installed locally, but UDF dependencies have to be installed on Databricks.
Checkpoint 1 of 5· Check yourself
A developer uses Databricks Connect from PyCharm. Their script loops over a Python list, applies a UDF in a DataFrame select, and then calls show(). Where does the UDF body execute?
Ordinary Python code runs locally, but UDFs are serialized, sent to the cluster and run there. That is also why UDF dependencies must be installed on Databricks.
“UDFs defined locally are serialized and transmitted to the cluster where it runs.”Source: docs.databricks.com
Checkpoint 2 of 5· Exam question
How does Databricks Connect relate to open-source Spark Connect, and which Databricks Runtime versions support it?
Correct answer: A — Databricks Connect extends Spark Connect with additions for Databricks compute modes and Unity Catalog, and it requires Databricks Runtime 13.3 LTS or above.
- A. Correct — Databricks Connect is built as an extension of the open-source Spark Connect protocol, adding support for Databricks compute modes and Unity Catalog, and it requires Databricks Runtime 13.3 LTS or newer.
- B. Incorrect — Databricks Connect is explicitly built on top of the Spark Connect protocol rather than being unrelated to it, and its Unity Catalog extensions require a newer runtime, not an older one.
- C. Incorrect — Databricks Connect does not replace or remove Spark Connect; it layers Databricks-specific additions onto the same protocol and targets Databricks compute, not an external standalone cluster.
- D. Incorrect — Databricks Connect uses the gRPC-based Spark Connect protocol rather than the Py4J driver bridge, and it has specific minimum runtime requirements rather than working on every version.
- E. Incorrect — Databricks Connect can target multiple Databricks compute types, including all-purpose and job clusters as well as SQL warehouses, so this restriction does not describe it.
2.Deferred analysis: errors appear later than you expect
Both Spark Classic and Spark Connect execute queries lazily. Transformations and spark.sql SELECTs build a plan, while actions and SQL commands such as INSERT or CREATE run eagerly. The key difference is analysis. Spark Connect defers analysis and name resolution until execution time.
Spark Classic resolves the plan while you build it, so a misspelled column fails on the line where you wrote it. A Spark Connect client keeps an unresolved plan locally. Anything that needs a resolved plan sends it to the server over RPC: accessing a schema, explaining the plan, persisting a DataFrame, or running an action. The server analyses the plan only at that point.
This has two practical consequences. First, a try/except around a chain of transformations catches nothing in Spark Connect. If you depend on that exception, force analysis inside the try block with df.columns, df.schema or df.collect(). Second, schema access is no longer free. The first access to a DataFrame's schema triggers an RPC, which then caches the schema, so you get better performance by avoiding analysis requests on large numbers of DataFrames.
| Aspect | Spark Classic | Spark Connect |
|---|---|---|
| Query execution | Lazy | Lazy |
| Schema analysis | Eager | Lazy |
| Schema access | Local | Triggers RPC and caches the schema on first access |
| Temporary views | Plan embedded | Name lookup |
| UDF serialization | At creation | At execution |
Checkpoint 3 of 5· Fill the gap
Which attribute access makes this try block catch an analysis error in Spark Connect?
try:
df = ...
df. ? # This will trigger eager analysis
except Exception as e:
print(f"Error: {repr(e)}")Reading df.columns needs a resolved schema, so the client sends the plan to the server for analysis and any error is raised inside the try block.
Source: docs.databricks.comSources2
3.Temp views and UDFs bind at execution time
Deferred resolution also changes what a DataFrame refers to.
Temporary views. In Spark Connect, a DataFrame holds only a reference to a temporary view by name. If you replace that view later, the DataFrame's data changes too, because the name is looked up at execution time. Spark Classic embeds the view's plan when the DataFrame is created, so replacing the view has no effect on it. The documented fix is to always create unique temporary view names, for example by including a UUID.
Python UDFs. In Spark Connect, Python UDFs are lazy. They are serialized and registered only at execution time. A UDF that reads an outer variable sees the variable's value when the action runs, not when the UDF was defined. Spark Classic captures the value at creation.
from pyspark.sql.functions import udf
x = 123
@udf("INT")
def foo():
return x
df = spark.range(1).select(foo())
x = 456
df.show() # Prints 456If a UDF has to depend on an external value, bind that value early with a function factory. Wrap the UDF creation in a helper function and pass in the current value as an argument, so each generated UDF keeps its own copy.
from pyspark.sql.functions import udf
def make_udf(value):
def foo():
return value
return udf(foo)
x = 123
foo_udf = make_udf(x)
x = 456
df = spark.range(1).select(foo_udf())
df.show() # Prints 123 as expectedCheckpoint 4 of 5· Check yourself
In a Spark Connect session, df10 is built from spark.table("tmp"), where tmp holds 10 rows. The code then runs createOrReplaceTempView("tmp") with 100 rows and calls df10.collect(). What happens, and what is the recommended fix?
Spark Connect resolves the view by name at execution time, so df10 sees the replacement. Unique view names keep earlier DataFrames stable.
“In Spark Connect, the DataFrame stores only a reference to the temporary view by name.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
A data engineering team submits a nightly ETL job as a scheduled Databricks job with no interactive user present, and needs the driver process to run inside the cluster so it keeps running even if the submitting machine goes offline: ``` spark-submit \ --deploy-mode ____ \ --class com.acme.etl.NightlyJob \ etl-job.jar ``` Which value correctly fills the blank?
Correct answer: A — cluster
- A. Correct — cluster deploy mode launches the driver process inside the cluster itself, so the job keeps running independently of whether the submitting machine stays online.
- B. Incorrect — client deploy mode launches the driver on the submitting machine, so the job would fail or stop if that machine goes offline before completion.
- C. Incorrect — local is not a `spark-submit` deploy mode value; it refers to running Spark entirely in a single JVM process without a cluster at all.
- D. Incorrect — standalone describes a cluster manager type, not a deploy mode value accepted by the `--deploy-mode` flag.
- E. Incorrect — embedded is not a recognized `spark-submit` deploy mode; only client and cluster are valid values for this flag.
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.In Spark Connect, referencing a non-existent column in filter() raises an analysis error on that line, just as in Spark Classic.Why is that wrong?
Spark Connect only builds an unresolved plan on the client. The error appears when schema access or an action sends the plan to the server for analysis.
Covered in Deferred analysis: errors appear later than you expect
2.A Python UDF captures the values of outer variables at the moment it is defined, whichever Spark mode you use.Why is that wrong?
That is true only in Spark Classic. In Spark Connect, Python UDFs are serialized at execution time, so they see the latest values unless you bind them with a function factory.
Covered in Temp views and UDFs bind at execution time
3.Databricks Connect is a separate protocol that replaces Spark Connect.Why is that wrong?
For Databricks Runtime 13.3 LTS and above, Databricks Connect is built on open-source Spark Connect and extends it for Databricks compute modes and Unity Catalog.
Practise it for real
Watch Spark Connect defer analysis by running a broken query against a local Spark Connect server
1.From the extracted Spark folder, run ./sbin/start-connect-server.sh
Why: This starts a Spark server with the Spark Connect endpoint enabled
You should see: The server is running and ready to accept Spark Connect sessions
2.Run export SPARK_REMOTE="sc://localhost" and then start ./bin/pyspark
Why: Setting SPARK_REMOTE turns the shell's session into a Spark Connect session without any code change
You should see: The welcome message says: Client connected to the Spark Connect server at localhost
3.Run type(spark)
Why: This confirms which kind of session you have
You should see: <class 'pyspark.sql.connect.session.SparkSession'>, whose path includes .connect.
4.Run df = spark.sql("select 1 as a, 2 as b").filter("c > 1")
Why: The client only builds an unresolved plan locally
You should see: No error, even though column c doesn't exist
5.Run df.columns
Why: Schema access sends the plan to the server for analysis
You should see: An error saying column c can't be found
Stuck? Get a nudge
If step 4 raises an error straight away, check type(spark). You're probably in a classic session because SPARK_REMOTE wasn't set in the shell that launched pyspark.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Spark Connect is an open-source gRPC-based protocol within Apache Spark that allows remote execution of Spark workloads using the DataFrame API.”
↩︎ Where Spark Connect runs, and Databricks Connect“General code runs locally: Python and Scala code runs on the client side, enabling interactive debugging.”
↩︎ Where Spark Connect runs, and Databricks Connect“DataFrame APIs are executed on Databricks compute.”
↩︎ Where Spark Connect runs, and Databricks Connect“Databricks Connect is an extension of Spark Connect with additions and modifications to support working with Databricks compute modes and Unity Catalog.”
↩︎ Exam trap 3“UDFs defined locally are serialized and transmitted to the cluster where it runs.”
↩︎ Checkpoint - 2.
“Python notebooks with Databricks Runtime version 14.3 and above, on standard compute”
↩︎ Where Spark Connect runs, and Databricks Connect“The key difference between Spark Connect and Spark Classic is that Spark Connect defers analysis and name resolution to execution time”
↩︎ Deferred analysis: errors appear later than you expect“you can trigger eager analysis, for example with df.columns, df.schema, or df.collect().”
↩︎ Deferred analysis: errors appear later than you expect“Performance can be improved if you avoid analysis requests on large numbers of DataFrames.”
↩︎ Deferred analysis: errors appear later than you expect“To mitigate the difference, always create unique temporary view names.”
↩︎ Temp views and UDFs bind at execution time“In Spark Classic, the value of x at the time of UDF creation is captured”
↩︎ Temp views and UDFs bind at execution time“The key difference between Spark Connect and Spark Classic is that Spark Connect defers analysis and name resolution to execution time”
↩︎ Exam trap 1“In Spark Connect, Python UDFs are lazy. Their serialization and registration are deferred until execution time.”
↩︎ Exam trap 2“on df.columns or df.show() an error will be thrown because the unresolved plan is sent to the server for analysis.”
↩︎ Prediction“In Spark Connect, the DataFrame stores only a reference to the temporary view by name.”
↩︎ Checkpoint - 3.
“You can optionally run Databricks Connect against an open source Spark Connect server.”
↩︎ Where Spark Connect runs, and Databricks Connect