What you will be able to do
- Register a DataFrame as a session-scoped temporary view and query it with spark.sql
- Choose between createTempView and createOrReplaceTempView based on what happens when the name already exists
- Create and query global temporary views through the global_temp schema, and know where they are not supported
- Read and drop temporary views with spark.table, spark.catalog.dropTempView and spark.catalog.dropGlobalTempView
- Recognise the SQL CREATE TEMPORARY VIEW syntax that matches the DataFrame methods
Key concept
Temporary view — A name registered in the catalog that points at a DataFrame, so SQL queries can refer to that DataFrame like a table. It is not stored permanently: a local temporary view exists only as long as the SparkSession that created it.
1.From DataFrame to SQL: registering a local temporary view
A DataFrame built in Python has no name that SQL can see. Calling spark.sql("SELECT * FROM people") only works once something called people exists in the catalog. The usual way to give a DataFrame that name is createOrReplaceTempView(name), which "creates or replaces a local temporary view with this DataFrame." It takes one argument, name, a string naming the view.
After that, spark.sql runs the query and returns a DataFrame with the result. You can move freely between the two APIs: build a DataFrame, register it, query it with SQL, and keep working on the result as a DataFrame.
df = spark.createDataFrame([(2, "Alice"), (5, "Bob")], schema=["age", "name"])
df.createOrReplaceTempView("people")
df2 = df.filter(df.age > 3)
df2.createOrReplaceTempView("people")
df3 = spark.sql("SELECT * FROM people")
assert sorted(df3.collect()) == sorted(df2.collect())
spark.catalog.dropTempView("people")
# TrueNotice what the assertion proves. The second createOrReplaceTempView("people") call points the name at df2, the filtered DataFrame, so the SQL query returns only Bob's row. The view is a name that refers to a DataFrame, not a copy of the data taken at registration time, and re-registering the name moves it to a different DataFrame.
The word *local* is about scope. The view's lifetime "is tied to the SparkSession that was used to create this DataFrame." The SQL reference describes the same behaviour from the SQL side: temporary views "are visible only to the session that created them and are dropped when the session ends." Nothing is written to storage, and another session can't see the name.
Checkpoint 1 of 8· Check yourself
A notebook registers df.createOrReplaceTempView("sales"). Which statement about the view sales is correct?
createOrReplaceTempView creates a local temporary view. It is scoped to the creating session and goes away when that session ends. It is not a persisted table, and it is not a global view.
“TEMPORARY views are visible only to the session that created them and are dropped when the session ends.”Source: docs.databricks.com
Checkpoint 2 of 8· Exam question
A data engineer builds a PySpark DataFrame `sales_df` from a CSV file and needs to run ad-hoc SQL aggregations against it using `spark.sql()` within the same notebook session. Which line of code registers `sales_df` so it can be queried by name in Spark SQL?
Correct answer: A — sales_df.createOrReplaceTempView("sales_df")
- A. This call registers the DataFrame as a session-scoped temporary view so `spark.sql()` can resolve `sales_df` by name in a query. It is the current, non-deprecated API for exactly this purpose.
- B. This writes the DataFrame out as a persistent metastore table, which is a heavier operation that durably stores data on disk rather than simply registering a name for ad-hoc SQL access.
- C. `cacheTable` marks an already-registered table or view for in-memory caching; it does not register a DataFrame as a queryable name, so calling it here without a prior registration would fail.
- D. `registerTempTable` was the 1.x-era API for this purpose and has been removed from current Spark releases, so relying on it will raise an `AttributeError` on a modern cluster.
- E. `persist` controls the storage level used when the DataFrame is later computed, but it does not give Spark SQL a name to resolve, so `spark.sql()` still could not find `sales_df`.
2.createTempView vs createOrReplaceTempView: what happens on a name clash
Spark has two methods for registering a local view, and the only difference between them is what happens on a name clash. createTempView(name) "creates a local temporary view with this DataFrame" and has the same session-tied lifetime, but it refuses to overwrite: it "throws TempTableAlreadyExistsException, if the view name already exists in the catalog." createOrReplaceTempView(name) replaces the existing view instead.
This is why createOrReplaceTempView is the common choice in notebooks. If you re-run a cell that calls createTempView, it fails the second time. If you re-run one that calls createOrReplaceTempView, it just re-points the name. To reuse a name with createTempView, drop the view first, which is exactly what the documentation example does.
Checkpoint 3 of 8· Check yourself
A notebook cell calls df.createTempView("orders"). You run the same cell a second time in the same session without dropping the view. What happens?
createTempView never overwrites. On the second run the name orders is already in the catalog, so the call throws; createOrReplaceTempView would have re-pointed the name instead.
“throws TempTableAlreadyExistsException, if the view name already exists in the catalog.”Source: docs.databricks.com
df.createTempView("people") # doctest: +IGNORE_EXCEPTION_DETAIL
# Traceback (most recent call last):
# ...
# AnalysisException: "Temporary table 'people' already exists;"
spark.catalog.dropTempView("people")
# True
df.createTempView("people")Sources3
3.Global temporary views and the global_temp schema
A local view belongs to one SparkSession. To share a view across sessions in the same application, you register a *global* temporary view with createGlobalTempView(name) or createOrReplaceGlobalTempView(name). Its lifetime "is tied to this Spark application" rather than to a single session.
The naming difference matters on the exam. Global temporary views "are tied to a system preserved temporary schema global_temp", so you query them with that prefix: global_temp.people, not people. The create/replace split works exactly as it does for local views. createGlobalTempView throws TempTableAlreadyExistsException if the name already exists, and createOrReplaceGlobalTempView replaces it. Global views also have their own drop call, spark.catalog.dropGlobalTempView.
| Method | Lifetime tied to | If the name already exists | Query it as | Drop with |
|---|---|---|---|---|
| createTempView(name) | SparkSession | Throws TempTableAlreadyExistsException | people | spark.catalog.dropTempView |
| createOrReplaceTempView(name) | SparkSession | Replaces the view | people | spark.catalog.dropTempView |
| createGlobalTempView(name) | Spark application | Throws TempTableAlreadyExistsException | global_temp.people | spark.catalog.dropGlobalTempView |
| createOrReplaceGlobalTempView(name) | Spark application | Replaces the view | global_temp.people | spark.catalog.dropGlobalTempView |
df = spark.createDataFrame([(2, "Alice"), (5, "Bob")], schema=["age", "name"])
df.createOrReplaceGlobalTempView("people")
df2 = df.filter(df.age > 3)
df2.createOrReplaceGlobalTempView("people")
df3 = spark.table("global_temp.people")
sorted(df3.collect()) == sorted(df2.collect())
# TrueCheckpoint 4 of 8· Fill the gap
Which schema name completes the query against this global temporary view?
df = spark.createDataFrame([(2, "Alice"), (5, "Bob")], schema=["age", "name"])
df.createGlobalTempView("people")
df2 = spark.sql("SELECT * FROM ? .people")Global temporary views live in the system-preserved global_temp schema, so the query has to qualify the name as global_temp.people.
Source: docs.databricks.comCheckpoint 5 of 8· Match them up
Match each method or name to what it does
Tap a term, then the definition that fits it.
Both global methods register the view under global_temp for the life of the application. They differ only in whether a name clash raises an error or replaces the view, and global views are dropped with dropGlobalTempView.
“GLOBAL TEMPORARY views are tied to a system preserved temporary schema global_temp.”Source: docs.databricks.com
Checkpoint 6 of 8· Exam question
A notebook runs the following code in a single Spark session: ``` df1 = spark.range(5) df1.createOrReplaceTempView("nums") df2 = spark.range(5, 10) df2.createOrReplaceTempView("nums") result = spark.sql("SELECT COUNT(*) AS cnt FROM nums") ``` What does `result` contain after this code runs?
Correct answer: A — A single row with `cnt` equal to 5, because the second `createOrReplaceTempView` call overwrote the temporary view named `nums` with `df2`.
- A. `createOrReplaceTempView` overwrites any existing view of the same name in the current session, so after the second call `nums` refers only to `df2`, which has 5 rows (5 through 9).
- B. Reusing a temporary view name does not union the prior and new DataFrames; the earlier binding is discarded entirely, so no union ever happens.
- C. Raising an exception on reuse is the behavior of `createTempView`, not `createOrReplaceTempView`, which is designed specifically to replace an existing view without error.
- D. Replacing a temporary view does not clear or null out the referenced data; the view simply points at whichever DataFrame was registered most recently, which still has 5 rows.
- E. Spark SQL does not version temporary view registrations; a name can be bound to only one DataFrame at a time, so no combined two-row result is produced.
4.Querying views with spark.sql and spark.table, and dropping them
Once a view is registered, there are two ways to read it. spark.sql(sqlQuery) "returns a DataFrame representing the result of the given query", so the view can appear anywhere a table name can, including joins, filters and aggregations. spark.table(name) returns the view directly as a DataFrame, without any SQL. The documentation uses both: spark.sql("SELECT * FROM global_temp.people") for one global view and spark.table("global_temp.people") for the other. The example below combines two local views with the DataFrame API.
df1 = spark.createDataFrame([(1, "John"), (2, "Jane")], schema=["id", "name"])
df2 = spark.createDataFrame([(3, "Jake"), (4, "Jill")], schema=["id", "name"])
df1.createTempView("table1")
df2.createTempView("table2")
result_df = spark.table("table1").union(spark.table("table2"))*When* the view name gets resolved depends on the deployment mode. The spark.sql reference states that "in Spark Classic, a temporary view referenced in spark.sql is resolved immediately." Under Spark Connect, the same reference is analysed lazily, so if a view is dropped, modified, or replaced after spark.sql, the execution may fail or generate different results.
To clean up, use the catalog. spark.catalog.dropTempView("people") removes a local view and spark.catalog.dropGlobalTempView("people") removes a global one. In the documentation examples, both return True when they drop the view.
Under Spark Connect, the view reference in spark.sql is analysed lazily, not when spark.sql is called. Replacing the view before the query executes can change the output or make it fail. In Spark Classic, the view would already have been resolved when spark.sql ran.
Checkpoint 7 of 8· Exam question
A data engineer registers a temporary view in Notebook A on a shared interactive cluster: `orders_df.createOrReplaceTempView("orders_view")`. A colleague opens Notebook B, attaches it to the same cluster, and runs `spark.sql("SELECT * FROM orders_view")`, receiving `Table or view not found: orders_view`. What is the most likely cause?
Correct answer: A — `orders_view` was registered as a session-scoped temporary view, and Notebook B runs in a separate `SparkSession` that cannot see it.
- A. `createOrReplaceTempView` scopes the view to the `SparkSession` that created it. Each notebook attached to a cluster typically runs its own session, so a session-scoped view registered in one notebook is invisible to another.
- B. Registering a temporary view does not trigger a write to storage at all; the view is a name bound to an in-memory DataFrame lineage, so write timing is not the cause of a not-found error.
- C. Unity Catalog governs permanent catalog objects and does not control the visibility of session-scoped temporary views, so its absence is not what produces this error.
- D. `orders_view` is not a reserved Spark SQL keyword, and keyword collisions produce a parse error rather than a table-or-view-not-found error.
- E. Temporary view visibility depends on the `SparkSession` that created it, not on the Databricks Runtime version, so a runtime mismatch would not explain this specific error.
Sources6
5.The SQL equivalent: CREATE TEMPORARY VIEW
The DataFrame methods have a SQL counterpart. The CREATE VIEW statement accepts TEMPORARY and GLOBAL TEMPORARY modifiers, so you can register views from SQL with the same local and global scopes.
-- Query-backed or metric view
CREATE [ OR REPLACE ] [ [ GLOBAL ] TEMPORARY ] VIEW [ IF NOT EXISTS ] view_nameThe clauses line up with the DataFrame methods. OR REPLACE plays the role of createOrReplace...: "CREATE OR REPLACE VIEW view_name is equivalent to DROP VIEW IF EXISTS view_name followed by CREATE VIEW view_name." SQL also has an option the DataFrame API doesn't offer. With IF NOT EXISTS, if a view by this name already exists the CREATE VIEW statement is ignored, so nothing is raised and nothing is replaced. You may specify at most one of IF NOT EXISTS or OR REPLACE.
One naming rule applies to every temporary view: "A temporary view's name must not be qualified." You write CREATE TEMPORARY VIEW people, not CREATE TEMPORARY VIEW db.people. Global views are still read back through global_temp.
Checkpoint 8 of 8· Check yourself
Which CREATE VIEW statement is NOT valid for a temporary view?
OR REPLACE and IF NOT EXISTS conflict with each other, and the syntax allows at most one of them. Every other combination shown follows the documented grammar.
“You may specify at most one of IF NOT EXISTS or OR REPLACE.”Source: docs.databricks.com
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.createTempView quietly replaces a view that already has the same name.Why is that wrong?
Only the createOrReplace... methods replace an existing view. createTempView and createGlobalTempView raise TempTableAlreadyExistsException, which appears in PySpark as an AnalysisException, when the name is already taken.
Covered in createTempView vs createOrReplaceTempView: what happens on a name clash
2.After createGlobalTempView("people"), you can query it with SELECT * FROM people.Why is that wrong?
Global temporary views are registered in the system-preserved global_temp schema, so they have to be referenced as global_temp.people.
Covered in Global temporary views and the global_temp schema
3.A view created with createOrReplaceTempView stays available to later sessions, like a table.Why is that wrong?
A local temporary view is visible only to the session that created it and is dropped when that session ends. Nothing is persisted.
Covered in From DataFrame to SQL: registering a local temporary view
4.Global temporary views are the portable choice and work on every Databricks compute type.Why is that wrong?
Global temp views are not supported on Databricks serverless compute. Session-scoped views created with createOrReplaceTempView are the recommended alternative there.
Covered in Global temporary views and the global_temp schema
Practise it for real
Register a DataFrame as a local temporary view, replace it, query it with SQL, and drop it
1.Run df = spark.createDataFrame([(2, "Alice"), (5, "Bob")], schema=["age", "name"]) and then df.createOrReplaceTempView("people").
Why: This registers the DataFrame under a name that SQL can resolve.
You should see: No output. spark.sql("SELECT * FROM people") now returns both rows.
2.Run df2 = df.filter(df.age > 3) followed by df2.createOrReplaceTempView("people").
Why: This shows that createOrReplaceTempView re-points an existing name instead of failing.
You should see: No error, even though the name people already existed.
3.Run df3 = spark.sql("SELECT * FROM people") and compare sorted(df3.collect()) with sorted(df2.collect()).
Why: This confirms the SQL query now reads the filtered DataFrame.
You should see: The two lists are equal and contain only Bob's row.
4.Run df.createTempView("people").
Why: This shows that createTempView refuses to overwrite an existing view.
You should see: An AnalysisException saying the temporary table 'people' already exists.
5.Run spark.catalog.dropTempView("people").
Why: This removes the local view from the session's catalog.
You should see: Returns True.
Stuck? Get a nudge
If the SQL query in step 3 still returns Alice, check that you called createOrReplaceTempView on df2 and not on df.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/pyspark/reference/classes/dataframe/createOrReplaceTempViewOfficial docs
“Creates or replaces a local temporary view with this DataFrame.”
↩︎ From DataFrame to SQL: registering a local temporary view“The lifetime of this temporary table is tied to the SparkSession that was used to create this DataFrame.”
↩︎ From DataFrame to SQL: registering a local temporary view“The lifetime of this temporary table is tied to the SparkSession that was used to create this DataFrame.”
↩︎ Key concept - 2.
“TEMPORARY views are visible only to the session that created them and are dropped when the session ends.”
↩︎ From DataFrame to SQL: registering a local temporary view“GLOBAL TEMPORARY views are tied to a system preserved temporary schema global_temp.”
↩︎ Global temporary views and the global_temp schema“CREATE OR REPLACE VIEW view_name is equivalent to DROP VIEW IF EXISTS view_name followed by CREATE VIEW view_name.”
↩︎ The SQL equivalent: CREATE TEMPORARY VIEW“If a view by this name already exists the CREATE VIEW statement is ignored.”
↩︎ The SQL equivalent: CREATE TEMPORARY VIEW“A temporary view's name must not be qualified.”
↩︎ The SQL equivalent: CREATE TEMPORARY VIEW“GLOBAL TEMPORARY views are tied to a system preserved temporary schema global_temp.”
↩︎ Exam trap 2“TEMPORARY views are visible only to the session that created them and are dropped when the session ends.”
↩︎ Exam trap 3“You may specify at most one of IF NOT EXISTS or OR REPLACE.”
↩︎ Checkpoint - 3.
“throws TempTableAlreadyExistsException, if the view name already exists in the catalog.”
↩︎ createTempView vs createOrReplaceTempView: what happens on a name clash“Creates a local temporary view with this DataFrame.”
↩︎ createTempView vs createOrReplaceTempView: what happens on a name clash“throws TempTableAlreadyExistsException, if the view name already exists in the catalog.”
↩︎ Exam trap 1 - 4.https://docs.databricks.com/aws/en/pyspark/reference/classes/dataframe/createGlobalTempViewOfficial docs
“The lifetime of this temporary view is tied to this Spark application.”
↩︎ Global temporary views and the global_temp schema - 5.https://docs.databricks.com/aws/en/pyspark/reference/classes/dataframe/createOrReplaceGlobalTempViewOfficial docs
“Databricks recommends using session-scoped temp views (createOrReplaceTempView) on serverless compute, since global_temp views are not supported.”
↩︎ Global temporary views and the global_temp schema“Databricks recommends using session-scoped temp views (createOrReplaceTempView) on serverless compute, since global_temp views are not supported.”
↩︎ Exam trap 4 - 6.
“Returns a DataFrame representing the result of the given query.”
↩︎ Querying views with spark.sql and spark.table, and dropping them“In Spark Classic, a temporary view referenced in spark.sql is resolved immediately.”
↩︎ Querying views with spark.sql and spark.table, and dropping them“if a view is dropped, modified, or replaced after spark.sql, the execution may fail or generate different results.”
↩︎ Querying views with spark.sql and spark.table, and dropping them