CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 6 · Lesson 29/32

    Spark Connect Architecture: gRPC, Unresolved Plans and Thin Clients

    Describe the features of Spark Connect.

    10 min read
    3.12% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain how Spark Connect separates the client application from the Spark server
    • Trace a DataFrame operation from the client, over gRPC, to the server and back as Arrow batches
    • Identify which Spark APIs a Spark Connect client cannot use, and why
    • Name the operational benefits of the decoupled design and connect a client using sc:// connection strings

    Key concept

    Unresolved logical plan as the protocol — A Spark Connect client never executes Spark itself. It builds an unresolved description of the DataFrame operations and sends it to a remote Spark server, which analyses, optimises and executes it. Almost every Spark Connect feature and limitation follows from this.

    1.A decoupled client and server

    Spark Connect arrived in Apache Spark 3.4. It splits a Spark application into two parts. The client is where your code runs. The server is where Spark runs. Databricks defines it as a gRPC-based protocol that specifies how a client application talks to a remote Spark server, and it lets you run Spark workloads remotely through the DataFrame API.

    Because the client only has to describe work, not perform it, the client library can be small. The Apache docs call it "a thin API", and the Databricks docs say the same thing: it is designed to be embedded in application servers, IDEs, notebooks and other programming languages. A web service, a notebook server or a debugger session in your IDE can each hold a Spark session without running a Spark driver inside them.

    Checkpoint 1 of 6· Check yourself

    Which statement best describes Spark Connect?

    Sources123

    2.How a query travels: protobuf, gRPC and Arrow

    Every DataFrame call goes through the same round trip. On the client, the Spark Connect library turns DataFrame operations into unresolved logical query plans and encodes them as protocol buffers. These messages go to the server over the gRPC framework.

    On the server, a Spark Connect endpoint embedded in the Spark server receives the plans and turns them into Spark's own logical plan operators. The Apache docs compare this step to parsing a SQL query: attributes and relations are parsed, and an initial parse plan is built. After that, the standard Spark execution process takes over, so Spark Connect queries get the same optimizer as any other Spark query. Results don't come back as Python objects or JVM references. The server streams them to the client over gRPC as Apache Arrow-encoded row batches.

    The Databricks docs give the whole protocol in one line: unresolved logical plans plus Apache Arrow, carried over gRPC. Databricks also states the transport layer: gRPC over HTTP/2.

    Checkpoint 2 of 6· Put it in order

    Put the stages of a Spark Connect query in order, from the client call to the result.

    1. 1.Results are streamed back to the client as Apache Arrow-encoded row batches
    2. 2.The Spark Connect endpoint on the server translates them into Spark logical plan operators
    3. 3.The client translates DataFrame operations into unresolved logical plans encoded as protocol buffers
    4. 4.The plans are sent to the server using gRPC
    5. 5.The standard Spark execution process runs the query

    Sources243

    3.What the client can no longer reach

    A main design goal of Spark Connect is full separation and isolation of the client from the server. That isolation costs you some APIs, and the exam expects you to know which ones.

    First, the client runs in a different process from the Spark driver, so it can't reach into the driver JVM. In PySpark, the client doesn't use Py4J, so private fields that hold the JVM implementation, such as df._jdf, aren't available.

    Second, the protocol describes work as logical plans, so it doesn't support every Spark execution API. Most importantly, it doesn't support RDDs.

    Third, the client is session-based. It can't touch cluster-wide properties that affect every connected client, so it has no access to the static Spark configuration or the SparkContext.

    The Databricks Connect for Python limitations list shows the same boundaries: Spark Context and RDDs are not available, and neither are libraries that depend on them or on the underlying JVM.

    Capabilities a Spark Connect client gives up, and the design reason
    CapabilityAvailable to a Spark Connect client?Reason
    Accessing JVM internals such as df._jdfNoThe client does not run in the same process as the Spark driver and PySpark does not use Py4J
    RDDsNoThe protocol uses logical plans, so it does not support all execution APIs
    SparkContext and static Spark configurationNoThe session-based client cannot change properties that affect all connected clients
    DataFrame APIYesThe protocol is built on the DataFrame API and unresolved logical plans

    Checkpoint 3 of 6· Check yourself

    A team moves a PySpark job to a Spark Connect session. Which part of the job is most likely to break?

    Checkpoint 4 of 6· Exam question

    The following PySpark code runs against a Spark Connect endpoint: ```python df = spark.read.table("sales") df2 = df.filter(df.amount > 100).select("id", "amount") ``` Before `df2` reaches the Spark Connect server, how is this chain of `filter` and `select` calls represented on the wire?

    Sources53

    4.Why the split helps: stability, upgrades and debugging

    The same isolation that removes JVM access also fixes several operational problems that come up when many users share one cluster.

    Stability. Each application runs in its own process, so one that uses too much memory only hurts its own environment. Users can also declare their own dependencies on the client without worrying about conflicts with the Spark driver.

    Upgradability. The Spark driver can be upgraded independently of applications, for example to pick up performance improvements and security fixes. Applications can stay forward-compatible as long as the server-side RPC definitions remain backwards compatible.

    Debuggability and observability. Your code runs in an ordinary local process, so you can debug it interactively from your IDE and monitor it with your application framework's own metrics and logging libraries. Databricks lists this as a headline feature of its Spark Connect–based client: interactively develop and debug from any IDE.

    Checkpoint 5 of 6· Match them up

    Match each benefit of the Spark Connect architecture to what it means in practice.

    Tap a term, then the definition that fits it.

    Sources23

    5.Connecting a client with an sc:// connection string

    A session only uses Spark Connect if you ask for it. Otherwise it behaves like classic Spark. There are three ways to opt in, and all of them use a connection string that starts with sc://.

    1. Set the SPARK_REMOTE environment variable on the client machine, for example sc://localhost, then create a session as usual. This needs no code change. 2. Pass --remote when you launch the pyspark or spark-shell shell. 3. Call remote() on the session builder in your own code.

    The Scala REPL connects to a local server on port 15002 by default. A connection string can also include a host, a port and a token, such as sc://myhost.com:443/;token=ABCDEFG. To confirm which kind of session you have, check its type: if the class path includes .connect., you're on Spark Connect. Databricks Connect accepts the same mechanisms, either through the remote function or the SPARK_REMOTE variable.

    Creating a Spark Connect session programmatically in Pythonpython
    from pyspark.sql import SparkSession
    spark = SparkSession.builder.remote("sc://localhost").getOrCreate()

    Checkpoint 6 of 6· Fill the gap

    Which flag launches the PySpark shell as a Spark Connect client?

    ./bin/pyspark  ?  "sc://localhost"

    Sources43

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A Spark Connect client can still use RDDs and SparkContext. They are just slower over the network.Why is that wrong?

      The protocol only describes logical plans, so RDDs aren't supported, and the session-based client has no access to SparkContext or the static configuration.

      Covered in What the client can no longer reach

    2. 2.Spark Connect sends results back as pickled Python objects or JVM references, the same way Py4J does.Why is that wrong?

      The client doesn't use Py4J. Results are streamed back over gRPC as Apache Arrow-encoded row batches.

      Covered in How a query travels: protobuf, gRPC and Arrow

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Spark Connect is a gRPC-based protocol within Apache Spark that specifies how a client application can communicate with a remote Spark Server.”
      ↩︎ A decoupled client and server
      “Transformations are constructed on the client side and sent as unresolved plans to the server.”
      ↩︎ Key concept
    2. 2.
      “The client API is designed to be thin, so that it can be embedded everywhere: in application servers, IDEs, notebooks, and programming languages.”
      ↩︎ A decoupled client and server
      “The underlying protocol uses Spark unresolved logical plans and Apache Arrow on top of gRPC.”
      ↩︎ How a query travels: protobuf, gRPC and Arrow
      “Interactively develop and debug from any IDE.”
      ↩︎ Why the split helps: stability, upgrades and debugging
      “The underlying protocol uses Spark unresolved logical plans and Apache Arrow on top of gRPC.”
      ↩︎ Exam trap 2
    3. 3.
      “In Apache Spark 3.4, Spark Connect introduced a decoupled client-server architecture”
      ↩︎ A decoupled client and server
      “The Spark Connect client translates DataFrame operations into unresolved logical query plans which are encoded using protocol buffers.”
      ↩︎ How a query travels: protobuf, gRPC and Arrow
      “This is similar to parsing a SQL query, where attributes and relations are parsed and an initial parse plan is built.”
      ↩︎ How a query travels: protobuf, gRPC and Arrow
      “in PySpark, the client does not use Py4J”
      ↩︎ What the client can no longer reach
      “Applications that use too much memory will now only impact their own environment as they can run in their own processes.”
      ↩︎ Why the split helps: stability, upgrades and debugging
      “By default, the REPL will attempt to connect to a local Spark Server on port 15002.”
      ↩︎ Connecting a client with an sc:// connection string
      “Most importantly, the client does not have access to the static Spark configuration or the SparkContext.”
      ↩︎ Exam trap 1
      “Spark Connect introduced a decoupled client-server architecture”
      ↩︎ Prediction
      “Results are streamed back to the client through gRPC as Apache Arrow-encoded row batches.”
      ↩︎ Checkpoint
      “the Spark Connect protocol does not support all the execution APIs of Spark, most importantly RDDs.”
      ↩︎ Checkpoint
      “The Spark driver can now seamlessly be upgraded independently of applications”
      ↩︎ Checkpoint
    4. 4.
      “Databricks Connect communicates with the Databricks Clusters via gRPC over HTTP/2.”
      ↩︎ How a query travels: protobuf, gRPC and Arrow
      “You can pass the string in the remote function or set the SPARK_REMOTE environment variable.”
      ↩︎ Connecting a client with an sc:// connection string
    5. 5.
      “Libraries that use RDDs, Spark Context, or access the underlying Spark JVM”
      ↩︎ What the client can no longer reach

    Continue to page 2 of 2

    Spark Connect vs Spark Classic: Lazy Analysis, Temp Views, UDFs and Databricks Connect

    Spotted a mistake, or was something unclear? Tell us.