What you will be able to do
- Explain how Unity Catalog volumes and external locations govern access to source data in Amazon S3
- Describe what Auto Loader's cloudFiles source does, including checkpointing, exactly-once processing and file detection modes
- Choose between Auto Loader and COPY INTO based on file volume, schema change and reprocessing needs
- Describe the limits of creating a table by uploading local files through the UI
Key concept
Auto Loader (the cloudFiles source) — Auto Loader is Databricks' way of ingesting files from cloud storage incrementally. You point it at a directory and it picks up each new file as it lands, processing it exactly once, so you don't have to track which files have already been loaded.
1.Ingesting from Amazon S3: volumes and external locations
Much of an organisation's raw data already lives in an S3 bucket. To bring it into Databricks, you first need governed access to that bucket. The Databricks guide to incremental ingestion from S3 reaches the source data through a Unity Catalog volume (the recommended option) or a Unity Catalog external location. Each gives you a different kind of path. A volume path looks like /Volumes/<catalog>/<schema>/<volume>/<path>/<folder>. An external location path looks like s3://<bucket>/<folder>/.
Non-admins need several things from an admin first: a Unity Catalog-enabled workspace, READ VOLUME or READ FILES on the source location, the path to the source data, and USE SCHEMA plus CREATE TABLE on the destination schema. If the source path is a volume path, the cluster must run Databricks Runtime 13.3 LTS or above. Before building a pipeline, the guide has you explore the data in a notebook. First list the directory to confirm you can reach it. Then sample a few records with read_files:
SELECT * from read_files('<path-to-source-data>', format => '<file-format>') LIMIT 10Next, write the ingestion code itself as a streaming table definition. You can't run it as an ordinary notebook cell. If you do, Databricks only checks that the syntax is valid. To actually load data, you create an ETL pipeline from the notebook under Jobs & Pipelines, set Pipeline mode to Triggered, choose Unity Catalog as the destination, and optionally add a schedule.
CREATE OR REFRESH STREAMING TABLE <table-name>AS SELECT *FROM STREAM read_files( '<path-to-source-data>', format => '<file-format>' )Checkpoint 1 of 5· Check yourself
You paste a CREATE OR REFRESH STREAMING TABLE statement into a notebook cell and run it. What happens?
Pipeline syntax in a notebook cell is only checked for validity. The logic runs once you create an ETL pipeline from the notebook.
“Lakeflow pipelines aren't designed to run interactively in notebook cells.”Source: docs.databricks.com
Sources1
2.Auto Loader: the cloudFiles source
In Python, the ingestion step from the previous section uses Auto Loader. You give Auto Loader a directory, and it processes new files as they arrive. It can also process the files that are already there. It reads from Amazon S3 (s3://), ADLS (abfss://), Google Cloud Storage (gs://), Unity Catalog volumes (/Volumes/) and Azure Blob Storage. It handles JSON, CSV, XML, PARQUET, AVRO, ORC, TEXT and BINARYFILE files, including pre-compressed ones.
Checkpoint 2 of 5· Fill the gap
Which format name selects Auto Loader in this Python ingestion code?
spark.readStream.format(' ? ') .option('cloudFiles.format', '<file-format>') .load(f'{<path-to-source-data>}')Auto Loader is used through the cloudFiles Structured Streaming source. The actual file format, such as JSON or CSV, is set separately with the cloudFiles.format option.
Source: docs.databricks.comAuto Loader's reliability comes from its checkpoint. As it discovers files, it records their metadata in a RocksDB key-value store inside the checkpoint location. That record is what guarantees each file is processed exactly once. If the stream fails, it resumes from the checkpoint, so you don't manage any state yourself. Databricks recommends running Auto Loader in Lakeflow pipelines, which manage the schema and checkpoint location for you. Auto Loader can also detect schema drift, tell you when the schema changes, and rescue data that would otherwise be lost.
Auto Loader finds new files in one of two modes. Directory listing is the default. File notification mode is what Databricks recommends for most workloads, because it skips directory listing altogether and can cut cloud costs. In either mode, Auto Loader does not guarantee the order in which files are discovered or processed. Design downstream logic to cope with late arrivals, for example by using soft deletes or comparing timestamps before an upsert.
Checkpoint 3 of 5· Exam question
A retail analytics team receives a continuous stream of CSV order files landing in an S3 bucket at unpredictable intervals throughout the day. They want new files picked up automatically as they arrive, without re-scanning files that were already loaded, and they want schema drift in the source files to be detected rather than silently ignored. Which approach best satisfies these requirements?
Correct answer: A — Configure Auto Loader with the `cloudFiles` format pointed at the S3 path, so it incrementally discovers new files and tracks schema drift using its checkpoint and schema location.
- A. This is correct because Auto Loader's `cloudFiles` source performs incremental file discovery and maintains a checkpoint of processed files, and its schema inference tracks schema drift over time so new or changed columns are surfaced instead of dropped.
- B. A nightly `COPY INTO` job only checks the bucket on a fixed schedule, so files arriving between runs sit unprocessed for hours, and relying on timestamps for dedup is far less robust than Auto Loader's built-in checkpointing.
- C. Manually uploading each file through Catalog Explorer does not scale for an unpredictable, continuous stream of arrivals and defeats the goal of automatic pickup of new files.
- D. Delta Sharing is a protocol for exposing existing Unity Catalog tables and volumes to recipients, not a mechanism for discovering and loading raw files sitting in an S3 bucket.
Sources2
3.Auto Loader or COPY INTO?
Auto Loader isn't the only incremental option for cloud object storage. COPY INTO lets SQL users load data idempotently and incrementally into Delta tables from Databricks SQL, notebooks or Lakeflow Jobs. Which one you choose depends on volume, how often the schema changes, and how often you need to reprocess files.
| Consideration | COPY INTO | Auto Loader |
|---|---|---|
| Expected file count over time | Thousands of files | Millions of files or more; fewer operations to discover files |
| Frequently evolving schema | Less suited | Better primitives for schema inference and evolution |
| Reprocessing a subset of re-uploaded files | Easier to manage | Harder; COPY INTO can reload the subset while the stream keeps running |
Checkpoint 4 of 5· Check yourself
A team expects millions of JSON files to land in S3 over the next year, and the schema changes often. Which approach fits best?
At millions of files, Auto Loader needs fewer discovery operations and can split work into batches. It also handles schema inference and evolution better.
“If you are expecting files in the order of millions or more over time, use Auto Loader.”Source: docs.databricks.com
Sources3
4.Uploading small files through the UI
Not every dataset sits in a bucket. To load a small file from your laptop, use the Create or modify a table using file upload page. It creates or overwrites a managed Delta table from CSV, TSV, JSON, Avro, Parquet or text files. You can upload up to 10 files at a time, and their total size must be under 2 gigabytes. Compressed files such as zip and tar aren't supported. Uploaded files go to a secure internal location that is garbage collected daily. You need a running compute resource to preview and configure the table. Workspace admins can turn the page off.
Checkpoint 5 of 5· Put it in order
Put the steps for starting a file-upload table in order
- 1.Click Create or modify a table.
- 2.Click New > Add or upload data.
- 3.Click browse or drag and drop files directly on the drop zone.
You open the add-data entry point, choose the table upload option, and then add the files.
“Click New > Add or upload data.”Source: docs.databricks.com
The total size of uploaded files must be under 2 GB, and compressed archives such as zip and tar aren't supported. A file this large belongs in cloud storage, loaded with Auto Loader or COPY INTO.
Sources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Auto Loader uses file notification mode by default.Why is that wrong?
Directory listing is the default. Databricks recommends file notification mode for most workloads, but you have to configure it.
Covered in Auto Loader: the cloudFiles source
2.COPY INTO is the better choice at any scale because it is idempotent.Why is that wrong?
COPY INTO suits thousands of files. At millions of files Auto Loader is cheaper and more efficient, because it needs fewer discovery operations.
Covered in Auto Loader or COPY INTO?
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“a Unity Catalog volume (recommended) or a Unity Catalog external location”
↩︎ Ingesting from Amazon S3: volumes and external locations“The USE SCHEMA and CREATE TABLE privileges on the schema you want to load data into.”
↩︎ Ingesting from Amazon S3: volumes and external locations“For Pipeline mode, select Triggered.”
↩︎ Ingesting from Amazon S3: volumes and external locations“The READ VOLUME permission on the Unity Catalog external volume or the READ FILES permission on the Unity Catalog external location”
↩︎ Prediction“Lakeflow pipelines aren't designed to run interactively in notebook cells.”
↩︎ Checkpoint - 2.
“automatically processes new files as they arrive, with the option of also processing existing files in that directory”
↩︎ Auto Loader: the cloudFiles source“Auto Loader can ingest JSON, CSV, XML, PARQUET, AVRO, ORC, TEXT, and BINARYFILE file formats.”
↩︎ Auto Loader: the cloudFiles source“This key-value store ensures that data is processed exactly once.”
↩︎ Auto Loader: the cloudFiles source“You do not need to provide a schema or checkpoint location because Lakeflow pipelines automatically manage these settings for your pipelines.”
↩︎ Auto Loader: the cloudFiles source“However, Databricks recommends file notification mode using file events for most workloads.”
↩︎ Auto Loader: the cloudFiles source“Auto Loader does not guarantee the order in which files are discovered or processed,”
↩︎ Auto Loader: the cloudFiles source“Auto Loader can detect schema drifts, notify you when schema changes happen, and rescue data that would have been otherwise ignored or lost.”
↩︎ Auto Loader: the cloudFiles source“It provides a Structured Streaming source called cloudFiles.”
↩︎ Key concept“By default, Auto Loader uses directory listing mode.”
↩︎ Exam trap 1 - 3.
“With COPY INTO, SQL users can idempotently and incrementally ingest data from cloud object storage into Delta tables.”
↩︎ Auto Loader or COPY INTO?“you can use COPY INTO to reload the subset of files while an Auto Loader stream is running simultaneously.”
↩︎ Auto Loader or COPY INTO?“If you're going to ingest files in the order of thousands over time, you can use COPY INTO.”
↩︎ Exam trap 2“If you are expecting files in the order of millions or more over time, use Auto Loader.”
↩︎ Checkpoint - 4.
“The total size of uploaded files must be under 2 gigabytes.”
↩︎ Uploading small files through the UI“Compressed files such as zip and tar files are not supported.”
↩︎ Uploading small files through the UI“supports uploading up to 10 files at a time”
↩︎ Uploading small files through the UI“Click New > Add or upload data.”
↩︎ Checkpoint