What you will be able to do
- Place the chunk-table write correctly within the Databricks unstructured data pipeline for RAG
- Define a chunk table schema with a unique key, the text to embed, and source and metadata columns
- Create the table as a Unity Catalog managed Delta table under a three-level name with the required privileges
Key concept
Chunk table as the AI Search source table — Chunked text is written to a Unity Catalog managed Delta table with one row per chunk. Each row has a unique chunk ID, the text that will be embedded, and columns that link the chunk back to its source document. Every write decision in this lesson exists to keep that table in this shape, because it is the table a vector index syncs from.
1.Where the chunk write sits in the pipeline
Databricks describes the data pipeline behind a RAG application as a chain of components. It starts with corpus composition and ingestion. Next comes data preprocessing, which covers parsing, enrichment, metadata extraction, deduplication and filtering. After that come chunking, embedding, and finally indexing and storage. This lesson covers one hand-off in that chain: the point where chunked text stops being an in-memory DataFrame and becomes a governed Delta Lake table in Unity Catalog. Embedding and indexing both read from that table, so its shape, and how you write to it, decide how smoothly the rest of the pipeline runs.
The chunk table is not the first table in the pipeline. Databricks recommends ingesting data in a scalable and incremental way, and it states a best practice for what happens next: raw source data should be ingested and stored in a target table. The guidance says this approach ensures data preservation, traceability, and auditing. In practice that gives you two tables. First a document table holding raw (or parsed) content, then a chunk table derived from it. The pipeline guidance expects you to experiment and tune chunking, indexing and parsing. When you change a chunking parameter, you rebuild the chunk table from the preserved source table instead of going back to the original systems.
Checkpoint 1 of 5· Put it in order
Put these pipeline stages in the order Databricks lists them for a RAG data pipeline.
- 1.Ingest raw source data and store it in a target table
- 2.Break the parsed data into smaller, manageable chunks
- 3.Convert the chunked text into vector representations
- 4.Create vector indices for optimized search
- 5.Parse the raw data to extract the relevant information
Chunking works on parsed data, embedding works on chunks, and indexing works on embeddings. That is why the chunk table sits between parsing and embedding.
“Chunking: Break down the parsed data into smaller, manageable chunks for efficient retrieval.”Source: docs.databricks.com
Sources1
2.What a chunk row contains
The ai_prep_search SQL function documentation gives a concrete example of a chunk table. It reads files from a Unity Catalog volume, parses them with ai_parse_document, prepares them with ai_prep_search, and then flattens the output into individual chunk rows and writes them to a Delta table. The key operation is the explode (variant_explode): one input document becomes many rows, one for each chunk. The final projection of that query is shown below.
SELECT
chunk.value:chunk_id::STRING AS chunk_id,
chunk.value:chunk_position::INT AS chunk_position,
chunk.value:chunk_to_retrieve::STRING AS chunk_to_retrieve,
chunk.value:chunk_to_embed::STRING AS chunk_to_embed,
chunk.value:metadata AS metadata,
prepped_documents.path AS source_uri
FROM
prepped_documents,
LATERAL variant_explode(prepped_documents.result:document.contents) AS chunk;| Column name | Type |
|---|---|
| chunk_id | STRING |
| chunk_position | INT |
| chunk_to_retrieve | STRING |
| chunk_to_embed | STRING |
| metadata | VARIANT |
| source_uri | STRING |
The same page says what two of these columns are for. The table can be an AI Search index source, with chunk_to_embed as the embedding column and chunk_id as the primary key. The source_uri value repeats on every chunk from the same file, so it links a chunk back to its document but cannot identify a single row. Only chunk_id can do that.
The metadata column matters too. Databricks says that storing metadata alongside chunked documents or their embeddings is essential for optimal performance, because it lets retrieval narrow down results. Combining vector similarity search with keyword-based filtering on that metadata is what the guidance calls a hybrid search pipeline. Chunks produced by other libraries should land in the same shape: a unique key, the text to embed, and source and metadata columns.
Checkpoint 2 of 5· Check yourself
A team writes chunks to a Delta table with only two columns, source_uri and text. What is the main design problem for using this table as an AI Search source?
Many chunks share one source_uri, so the table has no per-row key. The documented pattern adds a chunk_id column and uses it as the primary key.
“using chunk_to_embed as the embedding column and chunk_id as the primary key”Source: docs.databricks.com
3.Creating the managed table under a three-level name
Databricks calls Unity Catalog managed tables the default and recommended table type for Delta Lake. Unity Catalog handles reads, writes, storage and optimization for these tables. Managed tables also bring addressing rules that show up on the exam. All reads and writes to a managed table must use its full catalog_name.schema_name.table_name. Path-based access is not supported (except in Compatibility Mode), because it bypasses Unity Catalog access controls and can cause data corruption or loss. The Delta tutorial says the same thing from the other direction: path-based access works only for volumes and external tables. Source files can live in a volume path, but the chunk table itself is always referenced by name.
To create the table you need three privileges: USE CATALOG on the parent catalog, USE SCHEMA on the parent schema, and CREATE TABLE on the parent schema.
-- Create a managed Delta table
CREATE TABLE <catalog-name>.<schema-name>.<table-name>
(
<column-specification>
);In Python, you can create the same empty table by writing an empty DataFrame with a schema through saveAsTable(), or you can use the DeltaTableBuilder API. The tutorial points out that, compared with the DataFrame writers, the builder makes it easier to specify column comments, table properties and generated columns. That is handy if you want to document what chunk_to_embed holds, or set a table property up front.
Checkpoint 3 of 5· Fill the gap
Which DeltaTableBuilder method creates this table only if it is not already there, leaving an existing table in place?
from delta.tables import DeltaTable
(
DeltaTable. ? (spark)
.tableName("workspace.default.people_10k_2")
.addColumn("id", "INT")
.addColumn("firstName", "STRING")
.addColumn("lastName", "STRING", comment="surname")
.addColumn("gender", "STRING")
.addColumn("age", "INT")
.execute()
)createIfNotExists builds the table only when it is missing. createOrReplace would replace an existing table, and forName and merge work on tables that already exist rather than defining new ones.
Source: docs.databricks.comCheckpoint 4 of 5· Check yourself
A data engineer holds USE CATALOG on the catalog and CREATE TABLE on the target schema, but cannot create the chunk table. Which privilege is missing?
Creating a managed table requires USE CATALOG, USE SCHEMA and CREATE TABLE together. Without USE SCHEMA, the engineer cannot use the schema the table would live in.
“USE SCHEMA on the table's parent schema.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
A generative AI engineer has produced a Spark DataFrame of chunked text records, each with `document_id`, `chunk_index`, `chunk_text`, and `source_uri` columns. The chunks need to land in Unity Catalog in a form that a Delta Sync Index can later consume. Which operation correctly persists the data for this purpose?
Correct answer: A — Write the DataFrame using `df.write.format("delta").mode("append").saveAsTable("catalog.schema.chunks")`, ensuring the table has a unique primary-key column, then enable `delta.enableChangeDataFeed` on the table so a Delta Sync Index can track incremental updates.
- A. A managed Delta table written to a three-level `catalog.schema.table` name with a primary-key column and Change Data Feed enabled is exactly what a Delta Sync Index needs to incrementally track and embed new rows. This is the standard sequence for landing chunks in Unity Catalog ahead of indexing.
- B. AI Search indexes sync from Delta tables, not directly from raw Parquet files in a volume; volumes store unstructured files, not queryable table data that Delta Sync can track for incremental changes. Writing to a volume skips the transaction log that the sync mechanism relies on.
- C. Unity Catalog table references require an explicit three-level namespace; an unqualified table name resolves relative to the current session catalog and schema rather than being auto-resolved by the index, which can point the index at the wrong table or fail outright.
- D. Overwriting the table on every run destroys the table's history and defeats the purpose of Change Data Feed, which is designed to let the index sync only the incremental rows added since the last sync rather than requiring a full table replacement.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.You can write chunks to a Unity Catalog managed table by pointing the writer at its storage path.Why is that wrong?
Managed tables must be read and written by their catalog.schema.table name. Path-based access is supported only for volumes and external tables.
Covered in Creating the managed table under a three-level name
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/quality-data-pipeline-ragOfficial docs
“As a best practice, raw source data should be ingested and stored in a target table.”
↩︎ Where the chunk write sits in the pipeline“This approach ensures data preservation, traceability, and auditing.”
↩︎ Where the chunk write sits in the pipeline“Storing metadata alongside chunked documents or their corresponding embeddings is essential for optimal performance.”
↩︎ What a chunk row contains“Chunking: Break down the parsed data into smaller, manageable chunks for efficient retrieval.”
↩︎ Checkpoint - 2.
“The following example flattens the output into individual chunk rows and writes them to a Delta table.”
↩︎ What a chunk row contains“using chunk_to_embed as the embedding column and chunk_id as the primary key”
↩︎ Key concept“using chunk_to_embed as the embedding column and chunk_id as the primary key”
↩︎ Prediction - 3.https://docs.databricks.com/aws/en/tables/managedOfficial docs
“All reads and writes to managed tables must use table names and catalog and schema names where they exist.”
↩︎ Creating the managed table under a three-level name“USE SCHEMA on the table's parent schema.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/delta/tutorialOfficial docs
“Compared to DataFrameWriter and DataFrameWriterV2, the DeltaTableBuilder API makes it easier to specify additional information like column comments, table properties, and generated columns.”
↩︎ Creating the managed table under a three-level name“Path-based access is only supported for volumes and external tables, not for managed tables.”
↩︎ Exam trap 1