What you will be able to do
- Read files into a DataFrame with an explicit schema, given as a StructType or a DDL string, using spark.read
- Explain what a schema mismatch does to CSV parsing under PERMISSIVE, DROPMALFORMED and FAILFAST modes
- Write DataFrames with df.write using format shortcuts, options and partitionBy
- Choose the right save mode (append, overwrite, error/errorifexists, ignore) and know which one is the default
- Tell saveAsTable apart from insertInto when writing to or overwriting a table
Key concept
Reader/writer builder chain — All file I/O in the DataFrame API goes through two builders: spark.read (DataFrameReader) and df.write (DataFrameWriter). You set the format, options, schema (on read) and save mode (on write), and nothing happens until a terminal call such as load(), save() or a format shortcut like csv() or parquet() runs.
1.Two entry points: spark.read and df.write
Every input and output question on the exam starts from one of two objects, and it matters which object owns which call. You read through SparkSession.read, which returns a DataFrameReader. You write through DataFrame.write, which returns a DataFrameWriter. A common distractor reverses them, for example offering df.read or spark.write. Neither exists, because reading creates a DataFrame from the session and writing takes an existing DataFrame out to storage.
Both builders give you two equivalent ways to state the format. You can use a generic chain such as spark.read.format("json").load(path) or df.write.format("parquet").save(path). Or you can call a format shortcut such as spark.read.json(path) or df.write.parquet(path). The shortcut does the same thing as setting format() and then calling load() or save(). Options work the same way in both styles: option(key, value) adds one setting, and options(**options) adds several at once.
| Method | Role in the chain |
|---|---|
| format(source) | Specifies the input data source format |
| schema(schema) | Specifies the input schema |
| option(key, value) / options(**options) | Adds input options for the underlying data source |
| load(path, format, schema, **options) | Loads data from a data source and returns it as a DataFrame |
| csv / json / parquet / orc / text | Format shortcuts that load directly and return a DataFrame |
| table(tableName) | Returns the specified table as a DataFrame |
Each format shortcut has a fixed output shape worth knowing. text() is the odd one out. It does not split lines into fields. Instead it returns a DataFrame with one string column per line, and that column has a fixed name.
Checkpoint 1 of 8· Check yourself
You call spark.read.text("logs/") on a folder of plain-text log files. What is the schema of the resulting DataFrame?
The text reader does not parse fields. Each line becomes a row in a single string column called value.
“Loads text files and returns a DataFrame whose schema starts with a string column named "value".”Source: docs.databricks.com
2.Reading DataFrames with an explicit schema
Without a schema, Spark has to work one out. Some sources, JSON among them, infer it from the data, and for CSV you can set option("inferSchema", "true"). Inference costs time, because Spark has to look at the data before the real read. DataFrameReader.schema() lets you declare the schema up front, so the source can skip that step. The method takes one of two forms: a StructType built from StructField(name, type, nullable) entries, or a DDL-formatted string such as "name STRING, age INT". Both forms describe the same thing, and the exam may show either.
from pyspark.sql.types import StructType, StructField, StringType, IntegerType
# Define schema
schema = StructType([
StructField("name", StringType(), True),
StructField("age", IntegerType(), True)
])
# Read CSV with schema
df = spark.read.schema(schema).csv("path/to/file.csv")
# Read CSV with DDL-formatted string schema
df = spark.read.schema("name STRING, age INT").csv("path/to/file.csv")You can check what you got with printSchema(). A DDL string of "col0 INT, col1 DOUBLE" prints as col0: integer and col1: double, and both are marked nullable. The same schema() call works for Parquet too. The Databricks docs recommend passing one there to avoid the overhead of schema inference.
Matching by position is the CSV-specific catch. A mismatched schema does not always fail. It can quietly shift values into the wrong fields, and the docs warn that results can differ considerably depending on which columns you access. The docs also show a partial schema, listing only review_id and rating, used to read just a subset of columns.
So what happens when a value cannot be parsed into the declared type, such as a city name in an IntegerType field? That depends on the parser mode, which you set with .option("mode", ...). This read option has nothing to do with the writer's save mode, which comes up later.
| Mode | What happens to a row with an unparseable field |
|---|---|
| PERMISSIVE (default) | The row is kept, and nulls are inserted for fields that could not be parsed |
| DROPMALFORMED | Lines that contain fields that could not be parsed are dropped |
| FAILFAST | The read is aborted if any malformed data is found |
PERMISSIVE replaces bad values with nulls without telling you, so you may want to see which rows were affected. One way is to add a _corrupt_record string column to the schema you pass in, then filter for rows where it is not null. (The badRecordsPath option takes precedence over it: rows written to that path do not appear in the DataFrame.)
schema = StructType([
StructField("review_id", StringType(), True),
StructField("rating", IntegerType(), True),
StructField("comment", StringType(), True),
StructField("_corrupt_record", StringType(), True)
])
df = (spark.read
.format("csv")
.option("header", "true")
.option("mode", "PERMISSIVE")
.schema(schema)
.load("/Volumes/<catalog>/<schema>/<volume>/reviews_csv")
)
display(df.filter(df["_corrupt_record"].isNotNull()))Checkpoint 2 of 8· Match them up
Match each CSV parser mode to its behaviour when a field does not fit the schema
Tap a term, then the definition that fits it.
PERMISSIVE is the default and keeps the row with nulls. DROPMALFORMED removes the row. FAILFAST stops the whole read.
“PERMISSIVE (default): nulls are inserted for fields that could not be parsed correctly”Source: docs.databricks.com
Checkpoint 3 of 8· Exam question
A data engineer reads a daily CSV extract with no header row. The `order_id` column must be read as a long (not the int Spark would infer) to avoid downstream overflow, and schema inference is too slow on the multi-gigabyte file. Which line fills the blank so the file is read with the required types without triggering an inference scan? ```python from pyspark.sql.types import StructType, StructField, LongType, StringType order_schema = StructType([ StructField("order_id", LongType(), True), StructField("customer", StringType(), True) ]) df = spark.read.____.csv("/mnt/raw/orders/") ```
Correct answer: B — schema(order_schema)
- A. Turning on `inferSchema` makes Spark scan the file to guess column types, which is exactly the slow, unreliable path the engineer needs to avoid, and it would not guarantee `order_id` comes back as a long.
- B. Passing the `StructType` directly forces every row to be parsed according to the declared field names and types, so `order_id` is read as a `LongType` with no inference scan required.
- C. Declaring a header only tells Spark to treat the first line as column names; it says nothing about column types and does nothing to prevent inference from running.
- D. Chaining `format` and `load` is just an alternate way to specify the CSV source path and does not attach a schema, so the file would still fall back to default typing behavior.
3.Writing DataFrames to files
Writing works like reading, but starts from the DataFrame. df.write returns a DataFrameWriter. You set the format with a shortcut (json(), csv(), parquet(), orc(), text()) or with format(...) followed by save(path). Options go in with option(), for example header for CSV or compression for Parquet. The writer has no schema() method, because the schema it writes is the DataFrame's own. If you want a different output schema, you change the DataFrame before writing it.
The writer also controls the directory layout. partitionBy(*cols) creates a separate directory level on the file system for each value of the named columns. The Databricks Parquet example derives year and month columns from a date column with withColumn and then partitions on them. The writer settings can be chained before the final save(), as below.
# Chain multiple configuration methods
df.write \
.format("parquet") \
.mode("overwrite") \
.option("compression", "snappy") \
.partitionBy("year", "month") \
.save("path/to/output")Checkpoint 4 of 8· Check yourself
You want Parquet output where every distinct value of the year column has its own directory on the file system. Which DataFrameWriter method do you use?
partitionBy splits the output on the file system by the given columns. bucketBy and sortBy organise data within buckets of a saved table, so they do not give you per-value directories.
“Partitions the output by the given columns on the file system.”Source: docs.databricks.com
4.Overwriting and appending: save modes
mode(saveMode) sets what the writer does when data or a table already exists at the target. Mode only matters when there is existing data. Writing to a new path works the same way under all four. The exam tests whether you know the default. Many candidates guess overwrite or append, but the writer refuses to touch existing data unless you tell it to.
| Mode string | Behaviour when data already exists |
|---|---|
| append | Append contents of this DataFrame to existing data |
| overwrite | Overwrite existing data |
| error or errorifexists | Throw an exception (this is the default) |
| ignore | Silently ignore this operation |
Be careful with ignore. The job reports success, but nothing is written, so stale data can stay in place without anyone noticing. The reference example below writes to one directory twice. The first write uses overwrite, the second uses append, and then the directory is read back. The resulting DataFrame contains both Alice's and Sue's rows.
import tempfile
with tempfile.TemporaryDirectory(prefix="mode") as d:
# Overwrite the path with a new Parquet file
spark.createDataFrame(
[{"age": 100, "name": "Alice"}]
).write.mode("overwrite").format("parquet").save(d)
# Append another DataFrame into the Parquet file
spark.createDataFrame(
[{"age": 120, "name": "Sue"}]
).write.mode("append").format("parquet").save(d)
# Read the Parquet file as a DataFrame.
spark.read.parquet(d).show()Checkpoint 5 of 8· Fill the gap
A nightly job must replace whatever Parquet data is already at the output path. Which mode string completes the call?
df.write.mode(" ? ").parquet("path/to/output.parquet")overwrite replaces the existing data. append would add to it, ignore would skip the write, and errorifexists, the default, would raise an exception.
Source: docs.databricks.comCheckpoint 6 of 8· Exam question
A dataset arrives as newline-delimited JSON where every record must expose `event_id` as a string and `event_ts` as a timestamp. The team wants to define this schema inline as a compact string rather than build a `StructType` object. Which line reads the files with that requirement? ```python df = spark.read.____.json("/mnt/raw/events/") ```
Correct answer: A — schema("event_id STRING, event_ts TIMESTAMP")
- A. The reader's `schema()` method accepts a DDL-style string directly, so this line applies the compact string schema exactly as the team wants, typing `event_id` and `event_ts` without building a `StructType` object.
- B. There is no generic `schema` key for the reader's `option()` method; DataSource options configure things like delimiters and parsing modes, not the column schema itself, so this call has no effect on typing.
- C. `StructType.add()` expects an actual `DataType` instance such as `StringType()`, not the bare word `"STRING"`, so this raises an error rather than producing the required inline string schema.
- D. Disabling `inferSchema` only stops Spark from guessing types on its own; it does not supply the specific `event_id`/`event_ts` types the team needs, so columns would fall back to generic defaults.
5.Writing and overwriting tables: saveAsTable vs insertInto
The writer can also target a table instead of a path. Two methods do this, and they behave differently in ways the exam likes to test.
saveAsTable(name, format, mode, partitionBy, **options) saves the DataFrame as the named table. If the table already exists, the save mode decides what happens, and as with files the default is to throw an exception. Two details matter. First, in overwrite mode the DataFrame's schema does not have to match the existing table's schema, so overwrite can replace the table's structure as well as its rows. Second, in append mode the existing table's format and options are used, and columns are matched by name.
insertInto(tableName, overwrite) inserts into a table that must already exist. It takes an overwrite boolean instead of relying on a mode string. Unlike saveAsTable, it does not use column names to find column positions.
spark.sql("DROP TABLE IF EXISTS tblA")
spark.createDataFrame([
(100, "Alice"), (120, "Bob"), (140, "Tom")],
schema=["age", "name"]
).write.saveAsTable("tblA")# Insert into existing table
df.write.insertInto("existing_table")
# Insert into existing table with overwrite
df.write.insertInto("existing_table", overwrite=True)Checkpoint 7 of 8· Check yourself
A DataFrame has the columns (name, age), and an existing table was created with the columns (age, name). Which write lines the columns up by name?
saveAsTable uses column names to find the correct column positions. The documentation contrasts this with insertInto, which does not match by name.
“Unlike DataFrameWriter.insertInto, DataFrameWriter.saveAsTable uses column names to find the correct column positions.”Source: docs.databricks.com
df.write.mode("overwrite").saveAsTable("t"). In overwrite mode, saveAsTable does not require the DataFrame's schema to match the existing table schema.
Checkpoint 8 of 8· Exam question
A developer calls `.schema(customSchema)` on a `DataFrameReader` before loading a CSV file. What effect does this have on how Spark processes the file?
Correct answer: A — Spark skips the schema-inference scan entirely and parses each row directly according to the field names and types declared in `customSchema`.
- A. Providing an explicit schema is precisely what lets Spark bypass the extra pass over the data that inference requires, so parsing goes straight to typed values using the declared fields.
- B. This describes wasted work Spark does not do; once a schema is supplied there is no need to also infer one, since inference exists only to produce a schema when none was given.
- C. Schema application is not partition-specific: the same declared schema is used to parse every partition of the file, so there is no partial fallback to inference on other partitions.
- D. Spark does not blend an inferred schema with a supplied one; supplying `customSchema` replaces inference outright rather than combining with it.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If you don't call mode(), df.write overwrites (or appends to) data that is already at the path.Why is that wrong?
The default save mode is error/errorifexists. Writing to an existing path without mode() throws an exception, so you have to ask for overwrite or append explicitly.
Covered in Overwriting and appending: save modes
2.With header=true, an explicit CSV schema is matched to the header column names.Why is that wrong?
CSV fields are matched to schema fields by position. A schema in a different order from the file shifts values into the wrong columns.
Covered in Reading DataFrames with an explicit schema
3.Reading a CSV with an explicit schema fails by default when a value doesn't fit the declared type.Why is that wrong?
The default parser mode is PERMISSIVE. It keeps the row and puts null in the field that failed. Only FAILFAST aborts the read.
Covered in Reading DataFrames with an explicit schema
4.insertInto and saveAsTable both line up DataFrame columns with table columns by name.Why is that wrong?
Only saveAsTable uses column names to find column positions. The documentation contrasts this explicitly with insertInto.
Covered in Writing and overwriting tables: saveAsTable vs insertInto
Practise it for real
Watch the overwrite and append save modes act on the same Parquet path, then read the result back.
1.Inside a tempfile.TemporaryDirectory, write a one-row DataFrame ({"age": 100, "name": "Alice"}) with .write.mode("overwrite").format("parquet").save(d).
Why: Overwrite replaces whatever is at the path, so this write succeeds even if the directory exists.
You should see: The write finishes without an error.
2.Write a second one-row DataFrame ({"age": 120, "name": "Sue"}) to the same path with .write.mode("append").format("parquet").save(d).
Why: Append adds this DataFrame's contents to the existing data instead of replacing it.
You should see: The write finishes, and the directory now holds both writes.
3.Run spark.read.parquet(d).show().
Why: Reading back shows the combined effect of the two modes.
You should see: Two rows: Alice (100) and Sue (120).
4.Write to the same path again with .write.mode("error").format("parquet").save(d).
Why: error/errorifexists is the default mode, and it refuses to write when data already exists.
You should see: An exception saying the data already exists.
Stuck? Get a nudge
If the third step shows only Sue, check whether the second write used overwrite instead of append.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Use SparkSession.read to access this interface.”
↩︎ Two entry points: spark.read and df.write“Interface used to load a DataFrame from external storage systems (e.g. file systems, key-value stores, etc).”
↩︎ Key concept“Loads text files and returns a DataFrame whose schema starts with a string column named "value".”
↩︎ Checkpoint - 2.
“Use DataFrame.write to access this interface.”
↩︎ Two entry points: spark.read and df.write“Partitions the output by the given columns on the file system.”
↩︎ Writing DataFrames to files“ignore: Silently ignore this operation if data already exists.”
↩︎ Overwriting and appending: save modes“Inserts the content of the DataFrame to the specified table.”
↩︎ Writing and overwriting tables: saveAsTable vs insertInto“error or errorifexists: Throw an exception if data already exists (default).”
↩︎ Exam trap 1“error or errorifexists: Throw an exception if data already exists (default).”
↩︎ Prediction - 3.
“By specifying the schema here, the underlying data source can skip the schema inference step, which speeds up data loading.”
↩︎ Reading DataFrames with an explicit schema“A StructType object or a DDL-formatted string (for example, 'col0 INT, col1 DOUBLE').”
↩︎ Reading DataFrames with an explicit schema - 4.
“Specify a schema when reading Parquet files to avoid the overhead of schema inference.”
↩︎ Reading DataFrames with an explicit schema“Write partitioned Parquet files for optimized query performance on large datasets.”
↩︎ Writing DataFrames to files - 5.https://docs.databricks.com/aws/en/query/formats/csvOfficial docs
“You can add the column _corrupt_record to the schema provided to the DataFrameReader to review corrupt records in the resultant DataFrame.”
↩︎ Reading DataFrames with an explicit schema“FAILFAST: aborts the reading if any malformed data is found”
↩︎ Reading DataFrames with an explicit schema“CSV has no column-name metadata, so Spark maps schema fields to columns by position”
↩︎ Exam trap 2“PERMISSIVE (default): nulls are inserted for fields that could not be parsed correctly”
↩︎ Exam trap 3“CSV has no column-name metadata, so Spark maps schema fields to columns by position”
↩︎ Prediction“PERMISSIVE (default): nulls are inserted for fields that could not be parsed correctly”
↩︎ Checkpoint - 6.
“Specifies the behavior when data or table already exists.”
↩︎ Overwriting and appending: save modes - 7.https://docs.databricks.com/aws/en/pyspark/reference/classes/dataframewriter/saveAsTableOfficial docs
“When mode is 'overwrite', the schema of the DataFrame does not need to match the existing table schema.”
↩︎ Writing and overwriting tables: saveAsTable vs insertInto“When mode is 'append', if a table already exists, its format and options are used.”
↩︎ Writing and overwriting tables: saveAsTable vs insertInto“Unlike DataFrameWriter.insertInto, DataFrameWriter.saveAsTable uses column names to find the correct column positions.”
↩︎ Exam trap 4“Unlike DataFrameWriter.insertInto, DataFrameWriter.saveAsTable uses column names to find the correct column positions.”
↩︎ Checkpoint