What you will be able to do
- Open the Create or modify a table using file upload page from the workspace UI and get from a local file to a managed Delta table
- State the file formats, extensions, file count and size limits the upload page accepts
- Set the CSV and JSON format options and know how several files are combined
- Edit column names and types in the preview without causing NULLs or casting errors
Key concept
Create or modify a table using file upload — A workspace UI page that takes a few small files from your local machine and turns them straight into a managed Delta table, either a new one or one that replaces an existing table. You pick the catalog, schema and table name, and you check the format options and column types in a live preview before the table is created.
1.What the upload page accepts, and what it needs
Sometimes an analyst has a spreadsheet export or a JSON extract on a laptop and just needs it as a table. The quickest way to do that in Databricks is the Create or modify a table using file upload page. It reads the files you give it and writes a managed Delta table. Workspaces with Unity Catalog can put that table in Unity Catalog or in the legacy Hive metastore.
The page is meant for *small* files. You can upload at most 10 files at a time, and their combined size must be under 2 GB. Each file needs one of these extensions: .csv, .tsv (or .tab), .json, .avro, .parquet or .txt. Compressed archives such as zip and tar files are rejected, so unzip them before you upload.
Uploaded files first land in a secure internal staging location in your account, and that location is garbage collected every day. You can upload to staging without any compute attached. To preview and configure the table, though, you have to select an active compute resource. The page supports SQL warehouses, serverless compute and dedicated compute. It does not support group clusters.
You also need permission to create tables in the target schema. On top of that, a workspace admin may have turned the feature off. Admins manage the Upload data using the UI toggle on the workspace settings page, under the Security tab. That setting enables or prevents uploading data to a Delta table or DBFS from the homepage, the Data tab or the notebook File menu, and it also controls uploads to a Genie Agent, so it is broader than this one page. It does not block programmatic access to DBFS, such as the DBFS command-line interface.
Checkpoint 1 of 6· Check yourself
A colleague wants to load sales_2024.zip (300 MB, containing three CSVs) through the file upload page. What should they do?
The upload page does not accept compressed archives. Three plain CSV files totalling well under 2 GB are within its limits.
“Compressed files such as zip and tar files are not supported.”Source: docs.databricks.com
2.From the New menu to a created table
You start from the workspace sidebar. Click New > Add or upload data, then choose Create or modify a table. Click browse to pick the files, or drag and drop them onto the drop zone. After you attach compute, a preview of the data appears. It shows up to 50 rows, and you can switch between grid and list views.
Next, choose where the table goes. In a Unity Catalog-enabled workspace you first pick a catalog, or the legacy hive_metastore. Then you pick a schema and, if you want, change the suggested table name. The data files are stored in the locations configured for that schema, and you need permission to create a table in that schema. A dropdown sets the write mode to Create new table or Overwrite existing table. If you choose Create new table and the name is already taken, the page shows an error instead of quietly replacing the existing table. When the options and columns look right, click Create at the bottom of the page.
Checkpoint 2 of 6· Put it in order
Put the steps for creating a table from a local CSV in order.
- 1.Browse for the file or drag it onto the drop zone
- 2.Select a catalog and schema, and optionally edit the table name
- 3.Click New > Add or upload data
- 4.Click Create or modify a table
- 5.Click Create at the bottom of the page
You open the add-data menu, pick the table-creating option, supply the files, set the destination and finally create the table.
“Click New > Add or upload data.”Source: docs.databricks.com
You get an error because the name conflicts with the existing table. To replace the table, set the dropdown to Overwrite existing table.
Checkpoint 3 of 6· Exam question
A data analyst working in a Unity Catalog-enabled workspace wants to use the Workspace UI to upload a local CSV file and create a new table from it in the `sales.raw` schema inside the `bronze` catalog. The analyst already has `SELECT` on other tables in that schema but has never created objects there. Which combination of privileges lets the analyst complete the entire upload-to-table workflow through the UI?
Correct answer: A — `USE CATALOG` on `bronze`, `USE SCHEMA` on `sales.raw`, `CREATE TABLE` on `sales.raw`, and `WRITE VOLUME` on the target volume that stores the uploaded file.
- A. This is correct: uploading a file and creating a table from it through the UI requires the full chain of Unity Catalog privileges — catalog and schema usage to traverse the namespace, table creation rights on the target schema, and write access to the volume that stages the uploaded bytes.
- B. There is no privilege model where read access on a volume plus catalog-level `MODIFY` substitutes for the required chain; Unity Catalog does not infer write or create rights from unrelated read grants.
- C. Unity Catalog privileges do not cascade automatically from a catalog down to its schemas and volumes; each level still needs its own explicit grant, so a catalog-level create privilege alone is not sufficient.
- D. Cluster-level workspace permissions control who can attach to or manage compute resources, not who can create objects or write files in the Unity Catalog metastore, so this does not enable the upload workflow.
Sources1
3.Format options and multi-file uploads
The format options you see depend on the type of file you uploaded. The common ones are in the header bar, and the rarer ones are in the Advanced attributes dialog. The preview refreshes each time you change an option, so you can see the effect right away. The table below compares the defaults for the two text formats you are most likely to upload.
| Option | CSV | JSON |
|---|---|---|
| First row contains the header | Enabled by default | Not offered |
| Column delimiter | Comma by default for CSV files; a single character only, backslash is not supported | Not offered |
| Automatically detect column types | Enabled by default (off = every column STRING) | Enabled by default (off = every column STRING) |
| Rows span multiple lines | Disabled by default | Enabled by default |
| Merge the schema across multiple files | Offered; if disabled, the schema from one file is used | Not offered |
| Allow comments / Allow single quotes | Not offered | Both enabled by default |
| Infer timestamp | Not offered | Enabled by default |
Uploading several files at once adds two rules. First, the header setting applies to every file. If some files have a header row and others don't, you can lose data, so make them consistent before you upload. Second, the files are combined by appending their rows to the target table. The upload cannot join or merge records across files. If the files need to be combined on a key, load them as separate tables and join them afterwards.
Checkpoint 4 of 6· Check yourself
You upload orders_jan.csv and orders_feb.csv together into one new table. How are they combined?
A multi-file upload appends all the rows. It cannot join or merge records.
“Uploaded files combine by appending all data as rows in the target table.”Source: docs.databricks.com
Checkpoint 5 of 6· Exam question
A business analyst has an 8 GB point-of-sale export file and wants to load it into a bronze table using the Databricks Workspace UI's file upload flow. When they drag the file into the 'Create or modify table' upload dialog, the upload does not complete. What is the most likely cause, and what should the analyst do instead?
Correct answer: A — The browser-based upload UI enforces a 5 GB file size limit, so the analyst should use the Databricks SDK for Python or another programmatic path to load a file this large into a volume.
- A. The Workspace UI upload path caps browser-based uploads at 5 GB, so an 8 GB file exceeds that ceiling; the Databricks SDK or another programmatic ingestion route is the documented way to move a file that large into a volume.
- B. The UI's limit is measured per file, not as a cumulative 500 MB total, and splitting into many small chunks is not the documented workaround for the actual 5 GB per-file constraint.
- C. Delta Sharing configuration on a catalog has no bearing on the browser upload dialog's file size handling, so disabling sharing would not resolve an 8 GB file failing to upload.
- D. The upload dialog's failure at this size is driven by the file size limit, not character encoding, so re-encoding the file would not address why an 8 GB upload stalls.
Sources1
4.Editing column names and types
For CSV and JSON files, column types are inferred by default. If you want every column read as STRING, turn off Advanced attributes > Automatically detect column types. To change a type, click the type icon on a column. Nested STRUCT and ARRAY types can't be edited here. To rename a column, click the input box at the top of it. Column names cannot contain commas, backslashes or unicode characters such as emojis. Names with other special characters are supported because the page uses Column Mapping.
| Data type | Description |
|---|---|
| BIGINT | 8-byte signed integer numbers |
| BOOLEAN | Boolean (true, false) values |
| DATE | Year, month and day, without a time zone |
| DOUBLE | 8-byte double-precision floating point numbers |
| STRING | Character string values |
| TIMESTAMP | Year through second, with the session local timezone |
| STRUCT | Structure described by a sequence of fields |
| ARRAY | Sequence of elements of type elementType |
| DECIMAL(P,S) | Maximum precision P and fixed scale S |
Schema inference makes a best-effort guess, and overriding it can backfire. If you change a column to a type that some values can't be cast to, those values become NULL. Casting BIGINT to DATE or TIMESTAMP is not supported, and the docs warn that casting a BIGINT to a type like DATE, for example dates in the format of 'yyyy', may trigger errors. Databricks recommends creating the table first and then transforming those columns with SQL functions. Column comments can't be added during the upload either. Create the table, then open it in Catalog Explorer and add the comments there.
After the table exists. Inferred as BIGINT, the year values can't be cast to DATE in the upload page, and 'yyyy' style values may trigger errors. Databricks recommends creating the table first and then transforming the column with SQL functions.
Checkpoint 6 of 6· Check yourself
In the preview, an inferred BIGINT column holds year values. What is the recommended way to get a DATE column?
Casting BIGINT to DATE is not supported in the upload page. The documented approach is to create the table and then transform the column with SQL.
“Casting BIGINT to DATE or TIMESTAMP columns is not supported.”Source: docs.databricks.com
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The file upload page can take a zipped or tarred bundle of CSVs and unpack it.Why is that wrong?
Compressed archives are rejected. Upload the extracted files instead: up to 10 at a time, under 2 GB in total.
2.You need a running compute resource before you can upload anything.Why is that wrong?
Files can be uploaded to staging without compute. Compute is needed only to preview and configure the table, and group clusters are not supported for that step.
3.Uploading several related files together merges or joins their records into one table.Why is that wrong?
A multi-file upload only appends rows. Join or merge the data later with SQL.
Covered in Format options and multi-file uploads
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The total size of uploaded files must be under 2 gigabytes.”
↩︎ What the upload page accepts, and what it needs“Imported files are uploaded to a secure internal location within your account which is garbage collected daily.”
↩︎ What the upload page accepts, and what it needs“Group clusters are not supported.”
↩︎ What the upload page accepts, and what it needs“(For Unity Catalog-enabled workspaces only) You can select a catalog or the legacy hive_metastore.”
↩︎ From the New menu to a created table“You need proper permissions to create a table in a schema.”
↩︎ From the New menu to a created table“Operations that attempt to create new tables with name conflicts display an error message.”
↩︎ From the New menu to a created table“You can preview 50 rows of your data when you configure the options for the uploaded table.”
↩︎ From the New menu to a created table“Only a single character is allowed, and backslash is not supported.”
↩︎ Format options and multi-file uploads“Header settings apply to all files.”
↩︎ Format options and multi-file uploads“If this is set to false, all column types are inferred as STRING.”
↩︎ Format options and multi-file uploads“Changing column types can lead to some values being cast to NULL if the value cannot be cast correctly to the target data type.”
↩︎ Editing column names and types“such as dates in the format of 'yyyy', may trigger errors.”
↩︎ Editing column names and types“Column names do not support commas, backslashes, or unicode characters (such as emojis).”
↩︎ Editing column names and types“To add comments to columns, create the table and navigate to Catalog Explorer where you can add comments.”
↩︎ Editing column names and types“allows you to upload CSV, TSV, or JSON, Avro, Parquet, or text files to create or overwrite a managed Delta Lake table.”
↩︎ Key concept“Compressed files such as zip and tar files are not supported.”
↩︎ Exam trap 1“you must select an active compute resource to preview and configure your table”
↩︎ Exam trap 2“Joining or merging records during file upload is not supported.”
↩︎ Exam trap 3“you must select an active compute resource to preview and configure your table”
↩︎ Prediction“Compressed files such as zip and tar files are not supported.”
↩︎ Checkpoint“Click New > Add or upload data.”
↩︎ Checkpoint“Uploaded files combine by appending all data as rows in the target table.”
↩︎ Checkpoint“Casting BIGINT to DATE or TIMESTAMP columns is not supported.”
↩︎ Checkpoint - 2.
“It also controls whether users can upload files to a Genie Agent.”
↩︎ What the upload page accepts, and what it needs“This setting does not control programmatic access to the Databricks File System”
↩︎ What the upload page accepts, and what it needs