CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 2 · Lesson 10/56

    Designing a Unity Catalog Delta Table for Chunked Text

    Define operations and sequence to write given chunked text into Delta Lake tables in Unity Catalog

    9 min read
    1.79% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Place the chunk-table write correctly within the Databricks unstructured data pipeline for RAG
    • Define a chunk table schema with a unique key, the text to embed, and source and metadata columns
    • Create the table as a Unity Catalog managed Delta table under a three-level name with the required privileges

    Key concept

    Chunk table as the AI Search source table — Chunked text is written to a Unity Catalog managed Delta table with one row per chunk. Each row has a unique chunk ID, the text that will be embedded, and columns that link the chunk back to its source document. Every write decision in this lesson exists to keep that table in this shape, because it is the table a vector index syncs from.

    1.Where the chunk write sits in the pipeline

    Databricks describes the data pipeline behind a RAG application as a chain of components. It starts with corpus composition and ingestion. Next comes data preprocessing, which covers parsing, enrichment, metadata extraction, deduplication and filtering. After that come chunking, embedding, and finally indexing and storage. This lesson covers one hand-off in that chain: the point where chunked text stops being an in-memory DataFrame and becomes a governed Delta Lake table in Unity Catalog. Embedding and indexing both read from that table, so its shape, and how you write to it, decide how smoothly the rest of the pipeline runs.

    The chunk table is not the first table in the pipeline. Databricks recommends ingesting data in a scalable and incremental way, and it states a best practice for what happens next: raw source data should be ingested and stored in a target table. The guidance says this approach ensures data preservation, traceability, and auditing. In practice that gives you two tables. First a document table holding raw (or parsed) content, then a chunk table derived from it. The pipeline guidance expects you to experiment and tune chunking, indexing and parsing. When you change a chunking parameter, you rebuild the chunk table from the preserved source table instead of going back to the original systems.

    Checkpoint 1 of 5· Put it in order

    Put these pipeline stages in the order Databricks lists them for a RAG data pipeline.

    1. 1.Ingest raw source data and store it in a target table
    2. 2.Break the parsed data into smaller, manageable chunks
    3. 3.Convert the chunked text into vector representations
    4. 4.Create vector indices for optimized search
    5. 5.Parse the raw data to extract the relevant information

    Sources1

    2.What a chunk row contains

    The ai_prep_search SQL function documentation gives a concrete example of a chunk table. It reads files from a Unity Catalog volume, parses them with ai_parse_document, prepares them with ai_prep_search, and then flattens the output into individual chunk rows and writes them to a Delta table. The key operation is the explode (variant_explode): one input document becomes many rows, one for each chunk. The final projection of that query is shown below.

    Flattening prepared documents into one row per chunk, ready to be written to a Delta tablesql
    SELECT
      chunk.value:chunk_id::STRING AS chunk_id,
      chunk.value:chunk_position::INT AS chunk_position,
      chunk.value:chunk_to_retrieve::STRING AS chunk_to_retrieve,
      chunk.value:chunk_to_embed::STRING AS chunk_to_embed,
      chunk.value:metadata AS metadata,
      prepped_documents.path AS source_uri
    FROM
      prepped_documents,
      LATERAL variant_explode(prepped_documents.result:document.contents) AS chunk;
    Schema of the resulting chunk rows, as documented for ai_prep_search
    Column nameType
    chunk_idSTRING
    chunk_positionINT
    chunk_to_retrieveSTRING
    chunk_to_embedSTRING
    metadataVARIANT
    source_uriSTRING

    The same page says what two of these columns are for. The table can be an AI Search index source, with chunk_to_embed as the embedding column and chunk_id as the primary key. The source_uri value repeats on every chunk from the same file, so it links a chunk back to its document but cannot identify a single row. Only chunk_id can do that.

    The metadata column matters too. Databricks says that storing metadata alongside chunked documents or their embeddings is essential for optimal performance, because it lets retrieval narrow down results. Combining vector similarity search with keyword-based filtering on that metadata is what the guidance calls a hybrid search pipeline. Chunks produced by other libraries should land in the same shape: a unique key, the text to embed, and source and metadata columns.

    Checkpoint 2 of 5· Check yourself

    A team writes chunks to a Delta table with only two columns, source_uri and text. What is the main design problem for using this table as an AI Search source?

    Sources21

    3.Creating the managed table under a three-level name

    Databricks calls Unity Catalog managed tables the default and recommended table type for Delta Lake. Unity Catalog handles reads, writes, storage and optimization for these tables. Managed tables also bring addressing rules that show up on the exam. All reads and writes to a managed table must use its full catalog_name.schema_name.table_name. Path-based access is not supported (except in Compatibility Mode), because it bypasses Unity Catalog access controls and can cause data corruption or loss. The Delta tutorial says the same thing from the other direction: path-based access works only for volumes and external tables. Source files can live in a volume path, but the chunk table itself is always referenced by name.

    To create the table you need three privileges: USE CATALOG on the parent catalog, USE SCHEMA on the parent schema, and CREATE TABLE on the parent schema.

    SQL syntax for creating an empty managed Delta table under a three-level namesql
    -- Create a managed Delta table
    CREATE TABLE <catalog-name>.<schema-name>.<table-name>
    (
      <column-specification>
    );

    In Python, you can create the same empty table by writing an empty DataFrame with a schema through saveAsTable(), or you can use the DeltaTableBuilder API. The tutorial points out that, compared with the DataFrame writers, the builder makes it easier to specify column comments, table properties and generated columns. That is handy if you want to document what chunk_to_embed holds, or set a table property up front.

    Checkpoint 3 of 5· Fill the gap

    Which DeltaTableBuilder method creates this table only if it is not already there, leaving an existing table in place?

    from delta.tables import DeltaTable
    
    (
      DeltaTable. ? (spark)
        .tableName("workspace.default.people_10k_2")
        .addColumn("id", "INT")
        .addColumn("firstName", "STRING")
        .addColumn("lastName", "STRING", comment="surname")
        .addColumn("gender", "STRING")
        .addColumn("age", "INT")
        .execute()
    )

    Checkpoint 4 of 5· Check yourself

    A data engineer holds USE CATALOG on the catalog and CREATE TABLE on the target schema, but cannot create the chunk table. Which privilege is missing?

    Checkpoint 5 of 5· Exam question

    A generative AI engineer has produced a Spark DataFrame of chunked text records, each with `document_id`, `chunk_index`, `chunk_text`, and `source_uri` columns. The chunks need to land in Unity Catalog in a form that a Delta Sync Index can later consume. Which operation correctly persists the data for this purpose?

    Sources34

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.You can write chunks to a Unity Catalog managed table by pointing the writer at its storage path.Why is that wrong?

      Managed tables must be read and written by their catalog.schema.table name. Path-based access is supported only for volumes and external tables.

      Covered in Creating the managed table under a three-level name

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “As a best practice, raw source data should be ingested and stored in a target table.”
      ↩︎ Where the chunk write sits in the pipeline
      “This approach ensures data preservation, traceability, and auditing.”
      ↩︎ Where the chunk write sits in the pipeline
      “Storing metadata alongside chunked documents or their corresponding embeddings is essential for optimal performance.”
      ↩︎ What a chunk row contains
      “Chunking: Break down the parsed data into smaller, manageable chunks for efficient retrieval.”
      ↩︎ Checkpoint
    2. 2.
      “The following example flattens the output into individual chunk rows and writes them to a Delta table.”
      ↩︎ What a chunk row contains
      “using chunk_to_embed as the embedding column and chunk_id as the primary key”
      ↩︎ Key concept
      “using chunk_to_embed as the embedding column and chunk_id as the primary key”
      ↩︎ Prediction
    3. 3.
      “All reads and writes to managed tables must use table names and catalog and schema names where they exist.”
      ↩︎ Creating the managed table under a three-level name
      “USE SCHEMA on the table's parent schema.”
      ↩︎ Checkpoint
    4. 4.
      “Compared to DataFrameWriter and DataFrameWriterV2, the DeltaTableBuilder API makes it easier to specify additional information like column comments, table properties, and generated columns.”
      ↩︎ Creating the managed table under a three-level name
      “Path-based access is only supported for volumes and external tables, not for managed tables.”
      ↩︎ Exam trap 1

    Continue to page 2 of 2

    Writing and Updating Chunked Text in Delta Lake Tables

    Spotted a mistake, or was something unclear? Tell us.