What you will be able to do
- Explain what Photon is and which part of query processing it replaces
- Name the main performance benefits of Photon and the queries that gain the least from it
- Identify which workloads Photon accelerates and which it does not support
- Describe where Photon is enabled by default and what happens when a query uses an unsupported operation
- Tell from the Spark UI or the query profile how much of a query ran on Photon
Key concept
Photon as the execution layer — Photon is Databricks' native, vectorized engine. It takes over query execution from the JVM-based Spark SQL engine for the operations it supports. Spark's optimizer still plans the query, and Photon runs the plan on columnar batches, falling back to Spark for anything it cannot handle.
1.What Photon is and where it sits
When a SQL query runs on Databricks, two separate jobs happen. First the query is planned: the optimizer decides which joins, filters and scans to use. Then the plan is executed: rows are read, filtered, joined and written. Photon is about the second job. It is the Databricks-native vectorized query engine, and it speeds up SQL workloads, DataFrame API calls, ETL pipelines and stateless streaming.
Two design choices set Photon apart from the engine it replaces:
- Native code instead of the JVM. For the operations it supports, Photon swaps the JVM-based Spark SQL execution engine for a native C++ runtime. - Columnar batches instead of rows. Photon processes data in columnar batches, not one row at a time.
Photon doesn't replace everything. When it meets an unsupported operation during execution, it falls back to the Spark runtime for the rest of that operation. Photon is also compatible with Apache Spark APIs, so existing SQL and DataFrame code runs on it unchanged. You don't rewrite queries in a 'Photon dialect' to get the speedup.
Checkpoint 1 of 8· Check yourself
A team moves an existing DataFrame ETL job to Photon-enabled compute. What do they have to change in their code to use Photon?
Photon is compatible with Spark APIs, so existing code runs without changes. RDD APIs are actually among the things Photon does not support.
“compatible with Apache Spark APIs, so it works with your existing code with no changes required.”Source: docs.databricks.com
Sources1
2.Why Photon is faster: the benefits
The speedups come straight from the design described above. Because Photon works on batches of thousands of rows at once, modern CPUs can use SIMD instructions, which evaluate several values in one CPU cycle. Running native C++ instead of JVM code removes garbage-collection pauses, JIT warm-up delays and memory overhead. Columnar batches also allow cache-friendly sequential reads, which make the most of memory bandwidth and CPU pipelines.
The documentation describes these gains in several areas. On TPC-DS benchmarks, Photon delivers up to 5x better price/performance than other cloud data warehouses.
| Benefit area | What Photon does |
|---|---|
| Optimized joins and shuffles | Replaces sort-merge joins with hash joins; uses a redesigned columnar shuffle for large-scale joins |
| Write performance | Native Parquet writer speeds up Delta Lake, Apache Iceberg and Parquet writes, including UPDATE, DELETE, MERGE INTO, INSERT and CREATE TABLE AS SELECT. Wide tables gain the most. |
| Scan efficiency | Filter pushdown, dictionary pruning and row-group skipping reduce data read from storage, even with many small files |
| Disk cache and concurrency | Faster repeat access through the disk cache; higher throughput for concurrent queries in interactive BI |
| SQL and DataFrame integration | SQL and DataFrame APIs in Python, R, Scala and Java, with no code changes |
The limit is simple. Photon makes data processing faster, so it only helps when data processing takes up most of a query's runtime. If a query normally finishes in under two seconds, most of its time goes to planning and scheduling, and Photon doesn't make a meaningful difference.
Cost is the other half of the benefit. Photon instance types use DBUs at a different rate from the same instance type on the non-Photon runtime. So the real question for a recurring workload is whether the speedup outweighs the different rate. Databricks' cost guidance is to check regularly scheduled jobs for both speed and cost.
Checkpoint 2 of 8· Exam question
What is Photon in the Databricks Data Intelligence Platform, and what does it replace during query execution?
Correct answer: A — A native vectorized query engine written in C++ that replaces the JVM-based Spark SQL execution engine for supported SQL and DataFrame operations, processing rows in columnar batches.
- A. This is correct: Photon is Databricks' native, vectorized query engine implemented in C++ that swaps out the JVM-based Spark SQL execution engine for supported operations, processing data in columnar batches to cut garbage-collection overhead.
- B. This describes a caching mechanism, not Photon. Photon changes how query operators execute (vectorized native code), it does not replace Spark's shuffle service with a persistent cross-run cache.
- C. This describes a storage/file-format concept like Delta Lake's own file management, not Photon. Photon is an execution engine, not a table storage format, and it does not govern small-file compaction.
- D. This describes access control, which is Unity Catalog's job, not Photon's. Photon has no role in privilege enforcement or row-level security compilation.
3.Workloads Photon accelerates
Photon isn't only a SQL warehouse feature. It speeds up workloads across the platform, and none of them need code changes. Each workload has its own limits, though, and streaming has the most important one.
| Workload | How Photon applies |
|---|---|
| SQL analytics and BI | Default engine for all SQL warehouses, which run dashboards, ad hoc queries and scheduled reports |
| ETL and data engineering | Faster scans, joins, aggregations and writes for SQL or DataFrame batch jobs; the native Parquet writer helps ingestion into Delta Lake, Apache Iceberg or Parquet tables |
| Lakeflow pipelines | Enabling Photon in the pipeline configuration speeds up pipeline execution |
| Streaming | Stateless streaming only, writing to a Delta or Parquet sink; sources: Delta, Parquet, CSV, JSON, Kafka, Kinesis. Stateful streaming is not supported. |
| AI and machine learning | Spark SQL, DataFrames, feature engineering and GraphFrames operations |
For a data analyst, the SQL row matters most. Databricks SQL uses Photon under the hood, so dashboards and ad hoc queries on any SQL warehouse already run on it. All three warehouse types (serverless, pro and classic) support the Photon engine. Even a classic warehouse, which lacks Predictive IO and Intelligent Workload Management, still runs Photon.
The ML row comes with a warning. Photon speeds up the Spark SQL and DataFrame parts of an ML workflow. It is not expected to help code running outside the JVM engine, such as Python libraries or Pandas UDFs.
Checkpoint 3 of 8· Match them up
Match each workload to how Photon applies to it
Tap a term, then the definition that fits it.
SQL warehouses use Photon by default. Streaming is supported only when it is stateless and writes to a Delta or Parquet sink. The native Parquet writer speeds up ingestion into Delta Lake, Iceberg and Parquet tables.
“Photon supports stateless streaming when writing to a Delta or Parquet sink.”Source: docs.databricks.com
The DataFrame feature engineering. Photon improves Spark SQL, DataFrames, feature engineering and GraphFrames. It is not expected to help non-JVM code such as Python, so packages like PyTorch see no gain.
Checkpoint 4 of 8· Exam question
A data analyst is provisioning a new SQL warehouse in Databricks to run dashboard queries for the sales team. Without any extra configuration, what is true about Photon on this warehouse?
Correct answer: A — Photon is enabled by default on SQL warehouses, so the analyst's queries are automatically eligible for vectorized execution without toggling any setting.
- A. This is correct: Photon is the default engine on all SQL warehouses, so dashboard and BI queries run against the vectorized engine without the analyst needing to change any configuration.
- B. This is incorrect because there is no manual enablement step for SQL warehouses; unlike all-purpose or jobs compute, warehouses ship with Photon turned on out of the box.
- C. This is incorrect because Photon eligibility on SQL warehouses does not depend on cluster size; even the smallest warehouse sizes run on the native vectorized engine by default.
- D. This is incorrect because SQL warehouses do not expose a runtime-version picker the way all-purpose clusters do; Photon is built into the warehouse engine itself, not selected as a separate image.
4.Where Photon is on, and what it doesn't support
How Photon gets enabled depends on the type of compute:
- Always enabled: serverless compute, SQL warehouses, and serverless Lakeflow pipelines.
- Enabled by default in the UI: classic all-purpose compute, jobs compute, and classic Lakeflow pipelines. You turn it on or off with the *Use Photon Acceleration* checkbox under *Performance*.
- Explicit setting required through the APIs: compute created with the Clusters API or Jobs API needs runtime_engine set to PHOTON. With the Pipelines API, set photon to true.
Some features only work when Photon is enabled: predictive I/O for reads and writes, and dynamic file pruning in MERGE, UPDATE and DELETE statements.
Checkpoint 5 of 8· Check yourself
An engineer creates a jobs cluster with the Jobs API and does not mention Photon in the request. Assume the documented behaviour. What happens?
Photon is on by default only when you create classic compute in the UI. Through the Clusters or Jobs API you must enable it explicitly.
“you must explicitly enable Photon by setting runtime_engine to PHOTON”Source: docs.databricks.com
The documented limitations are short:
- No UDFs (user-defined functions), RDD APIs or Dataset APIs. - Stateless streaming only; no stateful streaming. - No improvement for queries that normally run in under two seconds. - Graviton-based instances with Photon don't support Databricks Container Services on dedicated compute or Databricks SQL warehouses. This doesn't apply to Container Services on standard compute.
An unsupported operation doesn't make a query fail. The compute switches to the Spark runtime for the rest of that operation, and the query still returns correct results. You lose speed, not correctness.
Checkpoint 6 of 8· Check yourself
A query on Photon-enabled compute calls a Python UDF in one step. What happens?
Photon doesn't support UDFs, but the fallback is transparent: Spark runs the unsupported part, and results stay correct.
“the compute resource transparently switches to the Spark runtime for the remainder of that operation”Source: docs.databricks.com
Checkpoint 7 of 8· Exam question
An analyst writes a Spark SQL query that calls a Python user-defined function (UDF) to clean a text column before aggregating results on an all-purpose cluster with Photon enabled. What happens to the UDF portion of that query?
Correct answer: A — The UDF call falls back to standard JVM-based Spark execution because Photon does not support user-defined functions, while other parts of the query can still run on the native engine.
- A. This is correct: Photon does not support UDFs, so operators involving the UDF fall back to the standard JVM Spark execution path, while eligible operators elsewhere in the same query plan can still benefit from Photon.
- B. This is incorrect because Photon has no mechanism to auto-translate arbitrary Python UDF logic into native vectorized expressions; UDFs remain outside Photon's supported operator set.
- C. This is incorrect because Databricks does not reject queries for mixing UDFs with Photon-eligible operators; Spark's planner simply routes unsupported operators to the fallback engine instead of failing the query.
- D. This is incorrect because there is no double-processing step; the UDF operator executes once under the fallback engine, and Photon does not reprocess its output a second time.
Sources1
5.Checking how much of a query ran on Photon
Because of fallback, one query can run partly on Photon and partly on Spark. Photon covers a broad set of operators:
- Scans of Parquet, Delta, CSV and JSON - Filter and Project - Hash aggregate, join and shuffle - Nested-loop and null-aware anti joins - Spatial joins - Union, Expand and ScalarSubquery - Delta/Parquet write sink - Sort, TopK and Limit - Window functions
The expressions it covers include comparison and logic, arithmetic, conditionals such as CASE, string functions, casts, aggregates, and date/timestamp functions. The documentation calls these categories representative, not exhaustive, and notes that individual functions may have limitations. Supported data types range from integers, decimals and strings to Struct, Array, Map, Variant, Geometry, Geography and collated strings.
The tool to use depends on the compute type:
- Spark UI (classic all-purpose and jobs compute): open the SQL/DataFrame tab. In the query DAG, Photon operators appear in orange and standard Spark operators in blue. - Query profile (SQL warehouses and serverless compute): the Execution Details view shows the percentage of task time spent in Photon. The plan shows Photon operators in purple and standard operators in grey.
If a query is using Photon less than you expected, check whether it uses unsupported operations, UDFs or data formats that trigger a fallback to Spark.
Checkpoint 8 of 8· Check yourself
On a SQL warehouse, which signal tells you directly how much of a query's work ran on Photon?
On SQL warehouses and serverless compute, the query profile's Execution Details view reports the share of task time spent in Photon. The orange/blue colours belong to the Spark UI on classic compute, and in the Spark UI blue marks standard Spark operators, not Photon.
“The Execution Details view shows the percentage of task time spent in Photon.”Source: docs.databricks.com
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If a query uses an operation Photon doesn't support, the query fails.Why is that wrong?
Photon falls back to the Spark runtime for the rest of that operation without any error, and the query still produces correct results. You lose speed on that part, not correctness.
2.Turning on Photon speeds up every query, including short interactive lookups.Why is that wrong?
Photon helps most with long-running queries over large datasets. Queries that normally finish in under two seconds spend most of their time on planning and scheduling, so they don't get meaningfully faster.
Covered in Why Photon is faster: the benefits
3.Photon supports streaming workloads in general.Why is that wrong?
Photon supports only stateless streaming that writes to a Delta or Parquet sink. Stateful streaming is not supported.
Covered in Workloads Photon accelerates
4.Photon replaces the Catalyst optimizer and plans queries itself.Why is that wrong?
Catalyst still plans the query. Photon replaces only the execution layer, and only for the operations it supports.
Covered in What Photon is and where it sits
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/compute/photonOfficial docs
“For supported operations, Photon replaces the JVM-based Spark SQL execution engine with a native C++ runtime.”
↩︎ What Photon is and where it sits“When Photon encounters an unsupported operation during query execution, it transparently falls back to the Spark runtime for the remainder of that operation.”
↩︎ What Photon is and where it sits“Photon processes data in batches of thousands of rows at a time, enabling modern CPUs to use SIMD instructions”
↩︎ Why Photon is faster: the benefits“By executing in native C++ instead of the JVM, Photon eliminates garbage collection pauses, JIT warm-up delays, and memory overhead.”
↩︎ Why Photon is faster: the benefits“Replaces sort-merge joins with high-performance hash joins”
↩︎ Why Photon is faster: the benefits“Implements filter pushdown, dictionary pruning, and row-group skipping to reduce data read from storage”
↩︎ Why Photon is faster: the benefits“Photon instance types consume DBUs at a different rate than the same instance type running the non-Photon runtime.”
↩︎ Why Photon is faster: the benefits“Photon is the default engine for all SQL warehouses”
↩︎ Workloads Photon accelerates“Supported sources include Delta, Parquet, CSV, JSON, Kafka, and Kinesis.”
↩︎ Workloads Photon accelerates“Photon is enabled on serverless compute, SQL warehouses, and serverless Lakeflow pipelines.”
↩︎ Where Photon is on, and what it doesn't support“you must explicitly enable Photon by setting runtime_engine to PHOTON”
↩︎ Where Photon is on, and what it doesn't support“Dynamic file pruning in MERGE, UPDATE, and DELETE statements.”
↩︎ Where Photon is on, and what it doesn't support“Your query still produces correct results.”
↩︎ Where Photon is on, and what it doesn't support“Photon operators appear in orange in the query DAG visualization.”
↩︎ Checking how much of a query ran on Photon“The Execution Details view shows the percentage of task time spent in Photon.”
↩︎ Checking how much of a query ran on Photon“check whether the query uses unsupported operations, UDFs, or data formats that cause a fallback to the Spark runtime”
↩︎ Checking how much of a query ran on Photon“For supported operations, Photon replaces the JVM-based Spark SQL execution engine with a native C++ runtime.”
↩︎ Key concept“Your query still produces correct results.”
↩︎ Exam trap 1“Photon provides the greatest benefit for longer-running queries that process large datasets.”
↩︎ Exam trap 2“Photon supports stateless streaming only.”
↩︎ Exam trap 3“The Apache Spark query optimizer (Catalyst) still plans your query, but Photon takes over at the execution layer”
↩︎ Exam trap 4“The Apache Spark query optimizer (Catalyst) still plans your query, but Photon takes over at the execution layer”
↩︎ Prediction“compatible with Apache Spark APIs, so it works with your existing code with no changes required.”
↩︎ Checkpoint“Photon provides the greatest benefit for longer-running queries that process large datasets.”
↩︎ Prediction“Photon supports stateless streaming when writing to a Delta or Parquet sink.”
↩︎ Checkpoint“the compute resource transparently switches to the Spark runtime for the remainder of that operation”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/lakehouse-architecture/cost-optimization/best-practicesOfficial docs
“jobs that run regularly should be evaluated to see whether they are not only faster but also cheaper with Photon.”
↩︎ Why Photon is faster: the benefits - 3.
“A classic SQL warehouse supports Photon but does not support Predictive IO or Intelligent Workload Management.”
↩︎ Workloads Photon accelerates - 4.https://docs.databricks.com/aws/en/spark/faqOfficial docs
“Databricks SQL uses Photon under the hood”
↩︎ Workloads Photon accelerates - 5.
“It is not expected to improve performance on applications using Spark RDDs, Pandas UDFs, and non-JVM languages such as Python.”
↩︎ Workloads Photon accelerates