What you will be able to do
- Explain how Spark Connect separates the client application from the Spark server
- Trace a DataFrame operation from the client, over gRPC, to the server and back as Arrow batches
- Identify which Spark APIs a Spark Connect client cannot use, and why
- Name the operational benefits of the decoupled design and connect a client using sc:// connection strings
Key concept
Unresolved logical plan as the protocol — A Spark Connect client never executes Spark itself. It builds an unresolved description of the DataFrame operations and sends it to a remote Spark server, which analyses, optimises and executes it. Almost every Spark Connect feature and limitation follows from this.
1.A decoupled client and server
Spark Connect arrived in Apache Spark 3.4. It splits a Spark application into two parts. The client is where your code runs. The server is where Spark runs. Databricks defines it as a gRPC-based protocol that specifies how a client application talks to a remote Spark server, and it lets you run Spark workloads remotely through the DataFrame API.
Because the client only has to describe work, not perform it, the client library can be small. The Apache docs call it "a thin API", and the Databricks docs say the same thing: it is designed to be embedded in application servers, IDEs, notebooks and other programming languages. A web service, a notebook server or a debugger session in your IDE can each hold a Spark session without running a Spark driver inside them.
Checkpoint 1 of 6· Check yourself
Which statement best describes Spark Connect?
Spark Connect is a client-server protocol over gRPC; the client talks to a remote Spark server rather than hosting a driver.
“Spark Connect is a gRPC-based protocol within Apache Spark that specifies how a client application can communicate with a remote Spark Server.”Source: docs.databricks.com
The client only translates DataFrame calls into plans and sends them over the network. The heavy part, the Spark driver and executors, stays on the remote server, so the embedding application doesn't have to host a driver JVM.
2.How a query travels: protobuf, gRPC and Arrow
Every DataFrame call goes through the same round trip. On the client, the Spark Connect library turns DataFrame operations into unresolved logical query plans and encodes them as protocol buffers. These messages go to the server over the gRPC framework.
On the server, a Spark Connect endpoint embedded in the Spark server receives the plans and turns them into Spark's own logical plan operators. The Apache docs compare this step to parsing a SQL query: attributes and relations are parsed, and an initial parse plan is built. After that, the standard Spark execution process takes over, so Spark Connect queries get the same optimizer as any other Spark query. Results don't come back as Python objects or JVM references. The server streams them to the client over gRPC as Apache Arrow-encoded row batches.
The Databricks docs give the whole protocol in one line: unresolved logical plans plus Apache Arrow, carried over gRPC. Databricks also states the transport layer: gRPC over HTTP/2.
Checkpoint 2 of 6· Put it in order
Put the stages of a Spark Connect query in order, from the client call to the result.
- 1.Results are streamed back to the client as Apache Arrow-encoded row batches
- 2.The Spark Connect endpoint on the server translates them into Spark logical plan operators
- 3.The client translates DataFrame operations into unresolved logical plans encoded as protocol buffers
- 4.The plans are sent to the server using gRPC
- 5.The standard Spark execution process runs the query
Plans are built and encoded on the client, shipped over gRPC, converted to logical operators on the server, executed by normal Spark, and returned as Arrow batches.
“Results are streamed back to the client through gRPC as Apache Arrow-encoded row batches.”Source: spark.apache.org
3.What the client can no longer reach
A main design goal of Spark Connect is full separation and isolation of the client from the server. That isolation costs you some APIs, and the exam expects you to know which ones.
First, the client runs in a different process from the Spark driver, so it can't reach into the driver JVM. In PySpark, the client doesn't use Py4J, so private fields that hold the JVM implementation, such as df._jdf, aren't available.
Second, the protocol describes work as logical plans, so it doesn't support every Spark execution API. Most importantly, it doesn't support RDDs.
Third, the client is session-based. It can't touch cluster-wide properties that affect every connected client, so it has no access to the static Spark configuration or the SparkContext.
The Databricks Connect for Python limitations list shows the same boundaries: Spark Context and RDDs are not available, and neither are libraries that depend on them or on the underlying JVM.
| Capability | Available to a Spark Connect client? | Reason |
|---|---|---|
| Accessing JVM internals such as df._jdf | No | The client does not run in the same process as the Spark driver and PySpark does not use Py4J |
| RDDs | No | The protocol uses logical plans, so it does not support all execution APIs |
| SparkContext and static Spark configuration | No | The session-based client cannot change properties that affect all connected clients |
| DataFrame API | Yes | The protocol is built on the DataFrame API and unresolved logical plans |
Checkpoint 3 of 6· Check yourself
A team moves a PySpark job to a Spark Connect session. Which part of the job is most likely to break?
RDDs are not part of the Spark Connect protocol, and the client has no SparkContext. DataFrame and SQL operations are fully supported.
“the Spark Connect protocol does not support all the execution APIs of Spark, most importantly RDDs.”Source: spark.apache.org
Checkpoint 4 of 6· Exam question
The following PySpark code runs against a Spark Connect endpoint: ```python df = spark.read.table("sales") df2 = df.filter(df.amount > 100).select("id", "amount") ``` Before `df2` reaches the Spark Connect server, how is this chain of `filter` and `select` calls represented on the wire?
Correct answer: A — As an unresolved logical plan built from the DataFrame API calls, encoded with Protocol Buffers and sent over a gRPC connection to the server.
- A. Correct — Spark Connect clients build an unresolved logical plan from DataFrame calls and encode it with Protocol Buffers before sending it over gRPC; the server resolves and optimizes it on arrival.
- B. Incorrect — the Spark Connect client has no access to internal Java references such as `_jdf`, and no Py4J socket is used at all; that direct-JVM bridge is exactly what Spark Connect removes.
- C. Incorrect — the client never computes partition or shuffle assignments itself; physical planning and execution happen entirely on the server after it receives the unresolved plan.
- D. Incorrect — DataFrame operations are not converted to SQL text on the client, and Spark Connect communicates over gRPC rather than a JDBC/Thrift connection.
- E. Incorrect — no bytecode compilation or class loading occurs on the client; the transformation chain stays as a structured plan object until the server resolves it.
4.Why the split helps: stability, upgrades and debugging
The same isolation that removes JVM access also fixes several operational problems that come up when many users share one cluster.
Stability. Each application runs in its own process, so one that uses too much memory only hurts its own environment. Users can also declare their own dependencies on the client without worrying about conflicts with the Spark driver.
Upgradability. The Spark driver can be upgraded independently of applications, for example to pick up performance improvements and security fixes. Applications can stay forward-compatible as long as the server-side RPC definitions remain backwards compatible.
Debuggability and observability. Your code runs in an ordinary local process, so you can debug it interactively from your IDE and monitor it with your application framework's own metrics and logging libraries. Databricks lists this as a headline feature of its Spark Connect–based client: interactively develop and debug from any IDE.
Checkpoint 5 of 6· Match them up
Match each benefit of the Spark Connect architecture to what it means in practice.
Tap a term, then the definition that fits it.
Each benefit comes from running the client in its own process, separate from the server-side driver.
“The Spark driver can now seamlessly be upgraded independently of applications”Source: spark.apache.org
5.Connecting a client with an sc:// connection string
A session only uses Spark Connect if you ask for it. Otherwise it behaves like classic Spark. There are three ways to opt in, and all of them use a connection string that starts with sc://.
1. Set the SPARK_REMOTE environment variable on the client machine, for example sc://localhost, then create a session as usual. This needs no code change.
2. Pass --remote when you launch the pyspark or spark-shell shell.
3. Call remote() on the session builder in your own code.
The Scala REPL connects to a local server on port 15002 by default. A connection string can also include a host, a port and a token, such as sc://myhost.com:443/;token=ABCDEFG. To confirm which kind of session you have, check its type: if the class path includes .connect., you're on Spark Connect. Databricks Connect accepts the same mechanisms, either through the remote function or the SPARK_REMOTE variable.
from pyspark.sql import SparkSession
spark = SparkSession.builder.remote("sc://localhost").getOrCreate()Checkpoint 6 of 6· Fill the gap
Which flag launches the PySpark shell as a Spark Connect client?
./bin/pyspark ? "sc://localhost"The --remote parameter, followed by the server location, starts the shell as a Spark Connect client.
Source: spark.apache.orgExam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A Spark Connect client can still use RDDs and SparkContext. They are just slower over the network.Why is that wrong?
The protocol only describes logical plans, so RDDs aren't supported, and the session-based client has no access to SparkContext or the static configuration.
Covered in What the client can no longer reach
2.Spark Connect sends results back as pickled Python objects or JVM references, the same way Py4J does.Why is that wrong?
The client doesn't use Py4J. Results are streamed back over gRPC as Apache Arrow-encoded row batches.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Spark Connect is a gRPC-based protocol within Apache Spark that specifies how a client application can communicate with a remote Spark Server.”
↩︎ A decoupled client and server“Transformations are constructed on the client side and sent as unresolved plans to the server.”
↩︎ Key concept - 2.
“The client API is designed to be thin, so that it can be embedded everywhere: in application servers, IDEs, notebooks, and programming languages.”
↩︎ A decoupled client and server“The underlying protocol uses Spark unresolved logical plans and Apache Arrow on top of gRPC.”
↩︎ How a query travels: protobuf, gRPC and Arrow“Interactively develop and debug from any IDE.”
↩︎ Why the split helps: stability, upgrades and debugging“The underlying protocol uses Spark unresolved logical plans and Apache Arrow on top of gRPC.”
↩︎ Exam trap 2 - 3.https://spark.apache.org/docs/latest/spark-connect-overview.htmlSecondary source
“In Apache Spark 3.4, Spark Connect introduced a decoupled client-server architecture”
↩︎ A decoupled client and server“The Spark Connect client translates DataFrame operations into unresolved logical query plans which are encoded using protocol buffers.”
↩︎ How a query travels: protobuf, gRPC and Arrow“This is similar to parsing a SQL query, where attributes and relations are parsed and an initial parse plan is built.”
↩︎ How a query travels: protobuf, gRPC and Arrow“in PySpark, the client does not use Py4J”
↩︎ What the client can no longer reach“Applications that use too much memory will now only impact their own environment as they can run in their own processes.”
↩︎ Why the split helps: stability, upgrades and debugging“By default, the REPL will attempt to connect to a local Spark Server on port 15002.”
↩︎ Connecting a client with an sc:// connection string“Most importantly, the client does not have access to the static Spark configuration or the SparkContext.”
↩︎ Exam trap 1“Spark Connect introduced a decoupled client-server architecture”
↩︎ Prediction“Results are streamed back to the client through gRPC as Apache Arrow-encoded row batches.”
↩︎ Checkpoint“the Spark Connect protocol does not support all the execution APIs of Spark, most importantly RDDs.”
↩︎ Checkpoint“The Spark driver can now seamlessly be upgraded independently of applications”
↩︎ Checkpoint - 4.
“Databricks Connect communicates with the Databricks Clusters via gRPC over HTTP/2.”
↩︎ How a query travels: protobuf, gRPC and Arrow“You can pass the string in the remote function or set the SPARK_REMOTE environment variable.”
↩︎ Connecting a client with an sc:// connection string - 5.
“Libraries that use RDDs, Spark Context, or access the underlying Spark JVM”
↩︎ What the client can no longer reach