CertSafari
    Snowflake SnowPro Specialty: Gen AI (GES-C02)· Lessons

    Domain 2 · Lesson 6/15

    Extracting, Enriching and Augmenting Data with Cortex AI

    Apply Snowflake Cortex functions in data pipelines.

    9 min read
    7.6% of exam
    5 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Choose between the LAYOUT and OCR modes of AI_PARSE_DOCUMENT, and use page_split and page_filter
    • Write AI_EXTRACT responseFormat values for entity, list and table extraction, and work around the 4096-token table limit
    • Enrich rows with AI_CLASSIFY labels on text and on staged files, within the limits that apply to document input
    • Augment extracted data with confidence scores so that a pipeline can route low-scoring results to a person

    1.Data extraction: AI_PARSE_DOCUMENT and AI_EXTRACT

    Document pipelines usually begin with files on a stage. Two functions turn those files into data. Both process files directly from object storage, so you don't need to move the data first.

    AI_PARSE_DOCUMENT turns a document into text with its reading order preserved. It keeps structural elements such as tables and headers, and returns Markdown inside a JSON object with page metadata. You choose a mode, and optionally page options:

    AI_PARSE_DOCUMENT modes and options
    OptionBehaviour
    'mode': 'LAYOUT'Preferred for most use cases and for complex documents. Extracts tables, headers and layout relationships, and is required for image extraction
    'mode': 'OCR'Fast, high-quality text extraction from scanned or text-heavy documents such as contracts and manuals
    'page_split': trueSplits a multi-page document into separate pages in the response
    page_filterProcesses only the specified pages. Implies page_split
    'extract_images': trueLAYOUT mode only. Extracts embedded images at no additional cost
    Parsing a staged PDF page by page with LAYOUT mode; TO_FILE supplies the file referencesql
    SELECT AI_PARSE_DOCUMENT (
        TO_FILE('@docs.doc_stage','research-paper-example.pdf'),
        {'mode': 'LAYOUT' , 'page_split': true}) AS research_paper_example;

    AI_EXTRACT goes a step further. Instead of returning the whole text, it returns only the fields you ask for, from either a string or a file. The responseFormat argument defines what you want:

    - Entities: a simple object that maps each label to a question, such as {'name': 'What is the last name of the employee?'}. An array of questions or a JSON schema with 'type': 'string' also works. - Lists: a JSON schema property with 'type': 'array'. - Tables: a JSON schema property with 'type': 'object', a column_ordering, and one array-typed property per column.

    The JSON schema format can't be mixed with the other formats in one call. String is the only supported scalar type.

    For tables, two limits shape how you design a pipeline. First, the table model stops after 4096 tokens of output. The documented workaround is to split a multi-page table into one-page documents and join the results afterwards. Second, scanned input can be poor. For large pages, small text or OCR typos, the config object's scale_factor (1.0 to 4.0) upscales pages before extraction.

    Checkpoint 1 of 4· Check yourself

    An AI_EXTRACT table extraction over a 12-page statement returns only the first part of the table. What does the documentation recommend?

    Sources123

    2.Data enrichment: adding labels and signals to rows

    Enrichment means adding new columns to existing data: a category, a sentiment, extracted entities. Snowflake lists "Extracting entities to enrich metadata and streamline validation" as a core AI Functions use case. The AI_PARSE_DOCUMENT documentation also names data enrichment explicitly: extracted text and images add "visual and contextual signals". You can chain the outputs, for example by passing extracted images to AI_EXTRACT or AI_COMPLETE for automatic tagging.

    AI_CLASSIFY is the workhorse for enrichment. It takes an input and an array of at least two categories, and returns an object with a labels array. Read that array with :labels to project it as a column. The input can be a text column or files listed through a stage's directory table, as in the following example.

    Enriching every image on a stage with a category labelsql
    WITH food_pictures AS (
      SELECT
          TO_FILE(file_url) AS img
      FROM DIRECTORY(@file_stage)
    )
    SELECT
    *,
    AI_CLASSIFY(img, ['dessert', 'drink', 'main dish', 'side dish']):labels AS classification
    FROM food_pictures;

    The optional config object lets you tune the call. task_description adds context in 50 words or fewer. output_mode: 'multi' allows multiple labels per row, and examples supplies few-shot examples. Each of these adds input tokens on every row, so in a large pipeline they increase cost. Document input has tighter rules: output_mode must be 'single' and examples are not supported. Accuracy may also drop beyond about twenty categories.

    Checkpoint 2 of 4· Fill the gap

    Complete the call so that a review can receive more than one topic label

    SELECT AI_CLASSIFY(
      'One day I will see the world and learn to cook my favorite dishes',
      ['travel', 'cooking', 'reading', 'driving'],
      {'output_mode': ' ? '}
    );

    Sources425

    3.Data augmentation: confidence scores and RAG-ready content

    The sources available for this lesson don't define "data augmentation" as a separate Cortex feature. They do document two ways in which pipelines add information beyond the raw extracted value, and these two are covered here.

    The first is extraction scores. Call AI_EXTRACT with the named-argument syntax and scores => TRUE, and the result gains a scoring object alongside response. It holds a score between 0 and 1 for each field, where a higher score means the value is more likely to be correct. Requesting scores costs nothing extra. A pipeline can use the scores to apply thresholds, set up fallbacks, or send low-scoring extractions to human review. For lists and tables, you get one aggregate score per list or table, not per cell.

    Checkpoint 3 of 4· Fill the gap

    Complete the call so that each extracted field comes back with a confidence value

    SELECT AI_EXTRACT(
      file => TO_FILE('@db.schema.files', 'document.pdf'),
      responseFormat => {'name': 'What is the last name of the employee?', 'date': 'What is the inspection date?'},
       ?  => TRUE
    );

    The second is preparing content for retrieval-augmented generation. AI_PARSE_DOCUMENT is documented for multimodal RAG: it combines extracted text and images to improve retrieval quality and to build richer knowledge bases. The retrieval and embedding mechanics belong to another lesson. For this objective, the point is that parsing in LAYOUT mode with images is the pipeline step that produces RAG-ready content.

    Checkpoint 4 of 4· Check yourself

    A pipeline extracts a list of account holders with scores => TRUE. Reviewers want to flag only the individual names the model was unsure about. What do they receive?

    Sources32

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.AI_CLASSIFY with output_mode 'multi' and few-shot examples works the same way on PDF documents as on text.Why is that wrong?

      Document input requires single-label output and doesn't accept examples.

      Covered in Data enrichment: adding labels and signals to rows

    2. 2.Turning on AI_EXTRACT scores gives a confidence value for every table cell and adds to the cost.Why is that wrong?

      Tables and lists get one aggregate score, and requesting scores is free.

      Covered in Data augmentation: confidence scores and RAG-ready content

    3. 3.OCR mode is the right AI_PARSE_DOCUMENT mode when you also need embedded images from the document.Why is that wrong?

      Image extraction requires LAYOUT mode, which is also the preferred mode for complex documents.

      Covered in Data extraction: AI_PARSE_DOCUMENT and AI_EXTRACT

    Practise it for real

    Build a small staged-document pipeline: parse a PDF, extract scored fields from it, and label images on a stage

    1. 1.Run AI_PARSE_DOCUMENT on a staged PDF with TO_FILE and {'mode': 'LAYOUT' , 'page_split': true}

      Why: LAYOUT mode keeps tables and headers, and page_split makes each page a separate unit

      You should see: A JSON object with metadata.pageCount and a pages array, each page holding Markdown content and an index

    2. 2.Run AI_EXTRACT on the same file with an entity responseFormat of two questions and scores => TRUE

      Why: Scores let you route low-confidence values to review

      You should see: A JSON object with response, scoring.scores (one score between 0 and 1 per field) and error set to null

    3. 3.Run AI_CLASSIFY over TO_FILE(file_url) FROM DIRECTORY(@your_stage) with four category labels, and project :labels

      Why: Enriches every staged image with a category column in a single set-based query

      You should see: One row per file, with a classification column holding an array of exactly one label

    Stuck? Get a nudge

    If an AI_EXTRACT table comes back cut off, check whether it spans several pages before you change anything else.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “extract structured information, such as entities, lists, and tables, from text or document files, by asking questions in natural language”
      ↩︎ Data extraction: AI_PARSE_DOCUMENT and AI_EXTRACT
      “The model for table extraction returns answers that are up to 4096 tokens long.”
      ↩︎ Data extraction: AI_PARSE_DOCUMENT and AI_EXTRACT
      “If the table covers several pages, split the document into multiple one-page documents, and join the results in postprocessing.”
      ↩︎ Checkpoint
    2. 2.
      “If using page_filter, page_split is implied, and you do not need to set it explicitly.”
      ↩︎ Data extraction: AI_PARSE_DOCUMENT and AI_EXTRACT
      “Data enrichment: Extract text and images to add visual and contextual signals for deeper insights”
      ↩︎ Data enrichment: adding labels and signals to rows
      “Multimodal RAG: Combine text and images to improve retrieval-augmented generation (RAG) quality”
      ↩︎ Data augmentation: confidence scores and RAG-ready content
      “LAYOUT mode is the preferred choice for most use cases, especially for complex documents.”
      ↩︎ Exam trap 3
    3. 3.
      “String is currently the only supported scalar type.”
      ↩︎ Data extraction: AI_PARSE_DOCUMENT and AI_EXTRACT
      “You can use these scores to set thresholds for business logic, such as flagging low-scoring extractions for human review.”
      ↩︎ Data augmentation: confidence scores and RAG-ready content
      “Requesting scores does not incur additional cost.”
      ↩︎ Exam trap 2
      “Per-element scores for individual list items and table cells are not available.”
      ↩︎ Checkpoint
    4. 4.
      “Extracting entities to enrich metadata and streamline validation”
      ↩︎ Data enrichment: adding labels and signals to rows
    5. 5.
      “Each label, description, and example increases the number of input tokens for every AI_CLASSIFY call, which affects cost.”
      ↩︎ Data enrichment: adding labels and signals to rows
      “When the input is a document, output_mode must be 'single' and examples is not supported.”
      ↩︎ Exam trap 1

    Ready to test yourself?

    Practise the 27 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.