CertSafari
    Databricks Certified Data Analyst Associate· Lessons

    Domain 5 · Lesson 18/39

    Photon Engine: Features, Benefits and Supported Workloads

    Understand the Features, Benefits, and Supported Workloads of Photon.

    13 min read
    2.56% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain what Photon is and which part of query processing it replaces
    • Name the main performance benefits of Photon and the queries that gain the least from it
    • Identify which workloads Photon accelerates and which it does not support
    • Describe where Photon is enabled by default and what happens when a query uses an unsupported operation
    • Tell from the Spark UI or the query profile how much of a query ran on Photon

    Key concept

    Photon as the execution layer — Photon is Databricks' native, vectorized engine. It takes over query execution from the JVM-based Spark SQL engine for the operations it supports. Spark's optimizer still plans the query, and Photon runs the plan on columnar batches, falling back to Spark for anything it cannot handle.

    1.What Photon is and where it sits

    When a SQL query runs on Databricks, two separate jobs happen. First the query is planned: the optimizer decides which joins, filters and scans to use. Then the plan is executed: rows are read, filtered, joined and written. Photon is about the second job. It is the Databricks-native vectorized query engine, and it speeds up SQL workloads, DataFrame API calls, ETL pipelines and stateless streaming.

    Two design choices set Photon apart from the engine it replaces:

    - Native code instead of the JVM. For the operations it supports, Photon swaps the JVM-based Spark SQL execution engine for a native C++ runtime. - Columnar batches instead of rows. Photon processes data in columnar batches, not one row at a time.

    Photon doesn't replace everything. When it meets an unsupported operation during execution, it falls back to the Spark runtime for the rest of that operation. Photon is also compatible with Apache Spark APIs, so existing SQL and DataFrame code runs on it unchanged. You don't rewrite queries in a 'Photon dialect' to get the speedup.

    Checkpoint 1 of 8· Check yourself

    A team moves an existing DataFrame ETL job to Photon-enabled compute. What do they have to change in their code to use Photon?

    Sources1

    2.Why Photon is faster: the benefits

    The speedups come straight from the design described above. Because Photon works on batches of thousands of rows at once, modern CPUs can use SIMD instructions, which evaluate several values in one CPU cycle. Running native C++ instead of JVM code removes garbage-collection pauses, JIT warm-up delays and memory overhead. Columnar batches also allow cache-friendly sequential reads, which make the most of memory bandwidth and CPU pipelines.

    The documentation describes these gains in several areas. On TPC-DS benchmarks, Photon delivers up to 5x better price/performance than other cloud data warehouses.

    Photon's benefit areas and what Photon does in each
    Benefit areaWhat Photon does
    Optimized joins and shufflesReplaces sort-merge joins with hash joins; uses a redesigned columnar shuffle for large-scale joins
    Write performanceNative Parquet writer speeds up Delta Lake, Apache Iceberg and Parquet writes, including UPDATE, DELETE, MERGE INTO, INSERT and CREATE TABLE AS SELECT. Wide tables gain the most.
    Scan efficiencyFilter pushdown, dictionary pruning and row-group skipping reduce data read from storage, even with many small files
    Disk cache and concurrencyFaster repeat access through the disk cache; higher throughput for concurrent queries in interactive BI
    SQL and DataFrame integrationSQL and DataFrame APIs in Python, R, Scala and Java, with no code changes

    The limit is simple. Photon makes data processing faster, so it only helps when data processing takes up most of a query's runtime. If a query normally finishes in under two seconds, most of its time goes to planning and scheduling, and Photon doesn't make a meaningful difference.

    Cost is the other half of the benefit. Photon instance types use DBUs at a different rate from the same instance type on the non-Photon runtime. So the real question for a recurring workload is whether the speedup outweighs the different rate. Databricks' cost guidance is to check regularly scheduled jobs for both speed and cost.

    Checkpoint 2 of 8· Exam question

    What is Photon in the Databricks Data Intelligence Platform, and what does it replace during query execution?

    Sources12

    3.Workloads Photon accelerates

    Photon isn't only a SQL warehouse feature. It speeds up workloads across the platform, and none of them need code changes. Each workload has its own limits, though, and streaming has the most important one.

    Workloads and how Photon applies to each
    WorkloadHow Photon applies
    SQL analytics and BIDefault engine for all SQL warehouses, which run dashboards, ad hoc queries and scheduled reports
    ETL and data engineeringFaster scans, joins, aggregations and writes for SQL or DataFrame batch jobs; the native Parquet writer helps ingestion into Delta Lake, Apache Iceberg or Parquet tables
    Lakeflow pipelinesEnabling Photon in the pipeline configuration speeds up pipeline execution
    StreamingStateless streaming only, writing to a Delta or Parquet sink; sources: Delta, Parquet, CSV, JSON, Kafka, Kinesis. Stateful streaming is not supported.
    AI and machine learningSpark SQL, DataFrames, feature engineering and GraphFrames operations

    For a data analyst, the SQL row matters most. Databricks SQL uses Photon under the hood, so dashboards and ad hoc queries on any SQL warehouse already run on it. All three warehouse types (serverless, pro and classic) support the Photon engine. Even a classic warehouse, which lacks Predictive IO and Intelligent Workload Management, still runs Photon.

    The ML row comes with a warning. Photon speeds up the Spark SQL and DataFrame parts of an ML workflow. It is not expected to help code running outside the JVM engine, such as Python libraries or Pandas UDFs.

    Checkpoint 3 of 8· Match them up

    Match each workload to how Photon applies to it

    Tap a term, then the definition that fits it.

    Checkpoint 4 of 8· Exam question

    A data analyst is provisioning a new SQL warehouse in Databricks to run dashboard queries for the sales team. Without any extra configuration, what is true about Photon on this warehouse?

    Sources1345

    4.Where Photon is on, and what it doesn't support

    How Photon gets enabled depends on the type of compute:

    - Always enabled: serverless compute, SQL warehouses, and serverless Lakeflow pipelines. - Enabled by default in the UI: classic all-purpose compute, jobs compute, and classic Lakeflow pipelines. You turn it on or off with the *Use Photon Acceleration* checkbox under *Performance*. - Explicit setting required through the APIs: compute created with the Clusters API or Jobs API needs runtime_engine set to PHOTON. With the Pipelines API, set photon to true.

    Some features only work when Photon is enabled: predictive I/O for reads and writes, and dynamic file pruning in MERGE, UPDATE and DELETE statements.

    Checkpoint 5 of 8· Check yourself

    An engineer creates a jobs cluster with the Jobs API and does not mention Photon in the request. Assume the documented behaviour. What happens?

    The documented limitations are short:

    - No UDFs (user-defined functions), RDD APIs or Dataset APIs. - Stateless streaming only; no stateful streaming. - No improvement for queries that normally run in under two seconds. - Graviton-based instances with Photon don't support Databricks Container Services on dedicated compute or Databricks SQL warehouses. This doesn't apply to Container Services on standard compute.

    An unsupported operation doesn't make a query fail. The compute switches to the Spark runtime for the rest of that operation, and the query still returns correct results. You lose speed, not correctness.

    Checkpoint 6 of 8· Check yourself

    A query on Photon-enabled compute calls a Python UDF in one step. What happens?

    Checkpoint 7 of 8· Exam question

    An analyst writes a Spark SQL query that calls a Python user-defined function (UDF) to clean a text column before aggregating results on an all-purpose cluster with Photon enabled. What happens to the UDF portion of that query?

    Sources1

    5.Checking how much of a query ran on Photon

    Because of fallback, one query can run partly on Photon and partly on Spark. Photon covers a broad set of operators:

    - Scans of Parquet, Delta, CSV and JSON - Filter and Project - Hash aggregate, join and shuffle - Nested-loop and null-aware anti joins - Spatial joins - Union, Expand and ScalarSubquery - Delta/Parquet write sink - Sort, TopK and Limit - Window functions

    The expressions it covers include comparison and logic, arithmetic, conditionals such as CASE, string functions, casts, aggregates, and date/timestamp functions. The documentation calls these categories representative, not exhaustive, and notes that individual functions may have limitations. Supported data types range from integers, decimals and strings to Struct, Array, Map, Variant, Geometry, Geography and collated strings.

    The tool to use depends on the compute type:

    - Spark UI (classic all-purpose and jobs compute): open the SQL/DataFrame tab. In the query DAG, Photon operators appear in orange and standard Spark operators in blue. - Query profile (SQL warehouses and serverless compute): the Execution Details view shows the percentage of task time spent in Photon. The plan shows Photon operators in purple and standard operators in grey.

    If a query is using Photon less than you expected, check whether it uses unsupported operations, UDFs or data formats that trigger a fallback to Spark.

    Checkpoint 8 of 8· Check yourself

    On a SQL warehouse, which signal tells you directly how much of a query's work ran on Photon?

    Sources1

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.If a query uses an operation Photon doesn't support, the query fails.Why is that wrong?

      Photon falls back to the Spark runtime for the rest of that operation without any error, and the query still produces correct results. You lose speed on that part, not correctness.

      Covered in Where Photon is on, and what it doesn't support

    2. 2.Turning on Photon speeds up every query, including short interactive lookups.Why is that wrong?

      Photon helps most with long-running queries over large datasets. Queries that normally finish in under two seconds spend most of their time on planning and scheduling, so they don't get meaningfully faster.

      Covered in Why Photon is faster: the benefits

    3. 3.Photon supports streaming workloads in general.Why is that wrong?

      Photon supports only stateless streaming that writes to a Delta or Parquet sink. Stateful streaming is not supported.

      Covered in Workloads Photon accelerates

    4. 4.Photon replaces the Catalyst optimizer and plans queries itself.Why is that wrong?

      Catalyst still plans the query. Photon replaces only the execution layer, and only for the operations it supports.

      Covered in What Photon is and where it sits

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “For supported operations, Photon replaces the JVM-based Spark SQL execution engine with a native C++ runtime.”
      ↩︎ What Photon is and where it sits
      “When Photon encounters an unsupported operation during query execution, it transparently falls back to the Spark runtime for the remainder of that operation.”
      ↩︎ What Photon is and where it sits
      “Photon processes data in batches of thousands of rows at a time, enabling modern CPUs to use SIMD instructions”
      ↩︎ Why Photon is faster: the benefits
      “By executing in native C++ instead of the JVM, Photon eliminates garbage collection pauses, JIT warm-up delays, and memory overhead.”
      ↩︎ Why Photon is faster: the benefits
      “Replaces sort-merge joins with high-performance hash joins”
      ↩︎ Why Photon is faster: the benefits
      “Implements filter pushdown, dictionary pruning, and row-group skipping to reduce data read from storage”
      ↩︎ Why Photon is faster: the benefits
      “Photon instance types consume DBUs at a different rate than the same instance type running the non-Photon runtime.”
      ↩︎ Why Photon is faster: the benefits
      “Photon is the default engine for all SQL warehouses”
      ↩︎ Workloads Photon accelerates
      “Supported sources include Delta, Parquet, CSV, JSON, Kafka, and Kinesis.”
      ↩︎ Workloads Photon accelerates
      “Photon is enabled on serverless compute, SQL warehouses, and serverless Lakeflow pipelines.”
      ↩︎ Where Photon is on, and what it doesn't support
      “you must explicitly enable Photon by setting runtime_engine to PHOTON”
      ↩︎ Where Photon is on, and what it doesn't support
      “Dynamic file pruning in MERGE, UPDATE, and DELETE statements.”
      ↩︎ Where Photon is on, and what it doesn't support
      “Your query still produces correct results.”
      ↩︎ Where Photon is on, and what it doesn't support
      “Photon operators appear in orange in the query DAG visualization.”
      ↩︎ Checking how much of a query ran on Photon
      “The Execution Details view shows the percentage of task time spent in Photon.”
      ↩︎ Checking how much of a query ran on Photon
      “check whether the query uses unsupported operations, UDFs, or data formats that cause a fallback to the Spark runtime”
      ↩︎ Checking how much of a query ran on Photon
      “For supported operations, Photon replaces the JVM-based Spark SQL execution engine with a native C++ runtime.”
      ↩︎ Key concept
      “Your query still produces correct results.”
      ↩︎ Exam trap 1
      “Photon provides the greatest benefit for longer-running queries that process large datasets.”
      ↩︎ Exam trap 2
      “Photon supports stateless streaming only.”
      ↩︎ Exam trap 3
      “The Apache Spark query optimizer (Catalyst) still plans your query, but Photon takes over at the execution layer”
      ↩︎ Exam trap 4
      “The Apache Spark query optimizer (Catalyst) still plans your query, but Photon takes over at the execution layer”
      ↩︎ Prediction
      “compatible with Apache Spark APIs, so it works with your existing code with no changes required.”
      ↩︎ Checkpoint
      “Photon provides the greatest benefit for longer-running queries that process large datasets.”
      ↩︎ Prediction
      “Photon supports stateless streaming when writing to a Delta or Parquet sink.”
      ↩︎ Checkpoint
      “the compute resource transparently switches to the Spark runtime for the remainder of that operation”
      ↩︎ Checkpoint
    2. 2.
      “jobs that run regularly should be evaluated to see whether they are not only faster but also cheaper with Photon.”
      ↩︎ Why Photon is faster: the benefits
    3. 3.
      “A classic SQL warehouse supports Photon but does not support Predictive IO or Intelligent Workload Management.”
      ↩︎ Workloads Photon accelerates
    4. 4.
      “Databricks SQL uses Photon under the hood”
      ↩︎ Workloads Photon accelerates
    5. 5.
      “It is not expected to improve performance on applications using Spark RDDs, Pandas UDFs, and non-JVM languages such as Python.”
      ↩︎ Workloads Photon accelerates

    Ready to test yourself?

    Practise Databricks Certified Data Analyst Associate in quiz mode.

    Spotted a mistake, or was something unclear? Tell us.