CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 2 · Lesson 8/32

    Spark File I/O: Formats, Save Modes and partitionBy

    Utilize common data sources such as JDBC, files, etc., to efficiently read from and write to Spark DataFrames using Spark SQL, including overwriting and partitioning by column.

    13 min read
    3.12% of exam
    7 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Read CSV, JSON, Parquet, ORC and text files into a DataFrame with DataFrameReader, with or without an explicit schema
    • Pick the right DataFrameWriter save mode and know which one applies when you don't set one
    • Overwrite existing file output explicitly with mode("overwrite") and recognise that the default mode will not do it for you
    • Write column-partitioned output with partitionBy and predict the directory layout it produces
    • Query files from SQL with read_files, and create partitioned tables from SQL with CREATE TABLE ... USING ... PARTITIONED BY

    Key concept

    The DataFrameReader / DataFrameWriter chain — Every file read starts at spark.read, and every write starts at df.write. You set the format, options and (on the write side) the save mode and partition columns, and finish with load() or save(), or with a format shortcut such as parquet() or csv().

    1.One interface for every file format

    Spark exposes data sources through one pair of builders. spark.read returns a DataFrameReader, described in the reference as the interface "used to load a DataFrame from external storage systems". df.write returns its mirror, a DataFrameWriter. Both work the same way. You can name the format with format("csv") and finish with load(path) or save(path), or you can call a shortcut method that names the format for you, such as csv(path), json(path), parquet(path), orc(path) or text(path). Between those two ends you add option(key, value) calls, and on the read side you can also call schema(...).

    Explicit format plus options: the long form of spark.read.csv(...)python
    # With options
    df = spark.read.format("csv") \
        .option("header", "true") \
        .option("inferSchema", "true") \
        .load("path/to/file.csv")

    Inferring a schema has a cost. The Parquet guide recommends that you "Specify a schema when reading Parquet files to avoid the overhead of schema inference." You can pass the schema to schema() as a StructType or as a DDL string such as "name STRING, age INT".

    The default format also matters. When you call save() or saveAsTable() without a format, what you get depends on the platform. The documentation states that Databricks uses Delta Lake as its default protocol for reading and writing data and tables, while Apache Spark uses Parquet. The formats group into three families: table formats (Delta Lake, Iceberg), binary columnar formats (Parquet, ORC), and human-readable formats (JSON, CSV, XML, Text).

    DataFrameReader shortcut methods and what each returns
    Reader methodWhat it loads
    csv(path, schema, sep, encoding, ...)A CSV file, returned as a DataFrame
    json(path, schema, ...)JSON files, returned as a DataFrame
    parquet(*paths, **options)Parquet files, returned as a DataFrame
    orc(path, mergeSchema, pathGlobFilter, ...)ORC files, returned as a DataFrame
    text(paths, wholetext, lineSep, ...)Text files; the schema starts with a string column named "value"
    table(tableName)The named table, returned as a DataFrame

    Writing works the same way for every format, and so does overwriting. When the target location already holds data, you replace it by setting the save mode to overwrite. The writer reference defines it as "Overwrite existing data." You can chain .mode("overwrite") before save(), or pass mode="overwrite" to a format shortcut such as parquet(), orc() or jdbc(). The same call applies to a file path, to saveAsTable(), and to a database table written through JDBC or a bundled connector, where it replaces the table's contents. Without it, the default mode refuses to write over existing data (see the next section for all four modes). The Parquet guide shows the pattern: the first write creates the files, and a later write with mode("overwrite") to the same path replaces them.

    Overwrite the existing Parquet output at the same pathpython
    df.write.format("parquet").mode("overwrite").save("/Volumes/<catalog>/<schema>/<volume>/reviews_parquet")

    Checkpoint 1 of 7· Check yourself

    You load a directory of log files and get back a DataFrame with one string column named value. Which reader method did you most likely call?

    Sources12345

    2.Save modes: what happens when the target already exists

    mode(saveMode) "Specifies the behavior when data or table already exists." It accepts four values. Two of them (error and errorifexists) are spellings of the same mode. You can set the mode by chaining .mode("overwrite"), or by passing it as a keyword to a format method, as in df.write.parquet(d, mode="overwrite"). The Parquet writer's signature lists 'error' or 'errorifexists' (default), which confirms that without a mode, a write to an existing location fails. On an exam question, a missing mode() call usually signals this error behaviour.

    ignore behaves differently from error. When data already exists, ignore does nothing and raises no exception, so the job succeeds even though nothing was written.

    The four DataFrameWriter save modes
    Mode stringWhen the target already has data
    appendAdds this DataFrame's contents to the existing data
    overwriteReplaces the existing data
    error / errorifexistsThrows an exception (this is the default)
    ignoreSilently skips the write

    Checkpoint 2 of 7· Match them up

    Match each save mode to its behaviour when the path already contains data

    Tap a term, then the definition that fits it.

    Checkpoint 3 of 7· Exam question

    A data engineer is reading a table from a PostgreSQL database using Spark's JDBC data source: ```python df = spark.read.format("jdbc") \ .option("url", "jdbc:postgresql://dbhost:5432/hr") \ .option("driver", "org.postgresql.Driver") \ .option("user", "svc_reader") \ .option("password", "s3cr3t") \ .load() ``` Running this code raises an `AnalysisException` because a required option is missing. Which option must be added to identify what to read?

    Sources64

    3.partitionBy: one directory per column value

    partitionBy(*cols) sets the physical layout of a file write. In the reference's words, "the output is laid out on the file system similar to Hive's partitioning scheme". Each distinct value of each partition column gets its own column=value subdirectory. The Parquet guide gives the reason: "Write partitioned Parquet files for optimized query performance on large datasets." A query that filters on the partition column can skip the directories it doesn't need.

    The official example below writes two rows partitioned by name with mode("overwrite") and reads them back in two ways. Reading the root path returns both age and name, because Spark rebuilds name from the directory names. Reading the name=Alice directory directly returns only age. Inside that directory, the partition value is part of the path, not a column in the files.

    Partitioned Parquet write, then a read of the root and of one partition directorypython
    with tempfile.TemporaryDirectory(prefix="partitionBy") as d:
        spark.createDataFrame(
            [{"age": 100, "name": "Alice"}, {"age": 120, "name": "Ruifeng Zheng"}]
        ).write.partitionBy("name").mode("overwrite").format("parquet").save(d)
    
        spark.read.parquet(d).sort("age").show()
        # +---+-------------+
        # |age|         name|
        # +---+-------------+
        # |100| Alice|
        # |120|Ruifeng Zheng|
        # +---+-------------+
    
        # Read one partition as a DataFrame.
        spark.read.parquet(f"{d}{os.path.sep}name=Alice").show()
        # +---+
        # |age|
        # +---+
        # |100|
        # +---+

    partitionBy accepts several columns, and the order you give sets the nesting order. The Parquet guide derives year and month from a check_in date with withColumn, then calls .partitionBy("year", "month"). Partition columns must already exist in the DataFrame, so you derive them before the write. You can also pass partition columns as an argument instead of chaining the method, for example parquet(path, mode, partitionBy, compression).

    Checkpoint 4 of 7· Check yourself

    What does df.write.partitionBy("country").parquet(path) change about the output?

    Checkpoint 5 of 7· Fill the gap

    Which writer method completes this multi-column partitioned write?

    # Partition by multiple columns
    df.write. ? ("year", "month").parquet("path/to/output.parquet")

    Checkpoint 6 of 7· Exam question

    A developer writes `df.write.format("parquet").save("/tmp/output")` and reruns the exact same job a second time against a path that already contains data from the first run. No `.mode(...)` was specified. What happens on the second run?

    Sources72

    4.Reading and partitioning files from SQL

    You can also read files with Spark SQL instead of the DataFrame API. The Parquet guide says to "Use read_files to query Parquet files directly from cloud storage using SQL without creating a table." read_files is a table-valued function. It takes a path and named options such as format => 'parquet', and it supports JSON, CSV, XML, TEXT, BINARYFILE, PARQUET, AVRO and ORC.

    Query Parquet files in a volume from SQL without registering a tablesql
    SELECT * FROM read_files(
      '/Volumes/<catalog>/<schema>/<volume>/reviews_parquet',
      format => 'parquet'
    )

    For data you will query again, SQL can create a table from a query, and the same statement can partition it. PARTITIONED BY in SQL does the same job as partitionBy on the writer. Note that the derived year and month columns are computed in the SELECT, just as the Python version adds them with withColumn first.

    The exam guide also mentions the shorthand `SELECT * FROM parquet.path `. The documentation used for this lesson doesn't describe that syntax, so this lesson doesn't make claims about it. Check the Spark SQL data sources guide for the details.

    SQL equivalent of a partitioned Parquet writesql
    -- Write partitioned Parquet files by year and month
    CREATE TABLE bookings_parquet_partitioned
    USING PARQUET
    PARTITIONED BY (year, month)
    AS SELECT *, year(check_in) AS year, month(check_in) AS month
    FROM samples.wanderbricks.bookings;

    Checkpoint 7 of 7· Check yourself

    An analyst wants to inspect Parquet files in a volume with SQL, without creating any table. Which statement fits?

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A plain df.write.parquet(path) to an existing location silently overwrites it.Why is that wrong?

      The default mode is error (errorifexists), so the write throws an exception. To replace the data you must ask for overwrite explicitly.

      Covered in Save modes: what happens when the target already exists

    2. 2.mode("ignore") raises an error when data already exists, just like errorifexists.Why is that wrong?

      ignore does nothing when data exists and raises no exception, so the job succeeds even though it wrote nothing.

      Covered in Save modes: what happens when the target already exists

    3. 3.save() without a format always writes Parquet.Why is that wrong?

      Parquet is the default only in Apache Spark. On Databricks, Delta Lake is the default format for reading and writing data and tables.

      Covered in One interface for every file format

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Interface used to load a DataFrame from external storage systems (e.g. file systems, key-value stores, etc).”
      ↩︎ One interface for every file format
      “Loads text files and returns a DataFrame whose schema starts with a string column named "value".”
      ↩︎ Checkpoint
    2. 2.
      “Specify a schema when reading Parquet files to avoid the overhead of schema inference.”
      ↩︎ One interface for every file format
      “Write partitioned Parquet files for optimized query performance on large datasets.”
      ↩︎ partitionBy: one directory per column value
      “Use read_files to query Parquet files directly from cloud storage using SQL without creating a table.”
      ↩︎ Reading and partitioning files from SQL
    3. 3.
      “Databricks uses Delta Lake as the default protocol for reading and writing data and tables, whereas Apache Spark uses Parquet.”
      ↩︎ One interface for every file format
      “Databricks uses Delta Lake as the default protocol for reading and writing data and tables, whereas Apache Spark uses Parquet.”
      ↩︎ Exam trap 3
    4. 4.
      “overwrite: Overwrite existing data.”
      ↩︎ One interface for every file format
      “error or errorifexists: Throw an exception if data already exists (default).”
      ↩︎ Save modes: what happens when the target already exists
      “Interface used to write a DataFrame to external storage systems (e.g. file systems, key-value stores, etc).”
      ↩︎ Key concept
      “error or errorifexists: Throw an exception if data already exists (default).”
      ↩︎ Exam trap 1
      “ignore: Silently ignore this operation if data already exists.”
      ↩︎ Exam trap 2
    5. 5.
      “Use append to add rows to an existing table or overwrite to replace its contents.”
      ↩︎ One interface for every file format
    6. 6.
      “Specifies the behavior when data or table already exists.”
      ↩︎ Save modes: what happens when the target already exists
      “'error' or 'errorifexists' (throw an exception if data exists), and 'ignore' (silently skip if data exists)”
      ↩︎ Checkpoint
    7. 7.
      “If specified, the output is laid out on the file system similar to Hive's partitioning scheme.”
      ↩︎ partitionBy: one directory per column value

    Continue to page 2 of 2

    Spark JDBC, saveAsTable and Temp Views

    Spotted a mistake, or was something unclear? Tell us.