CertSafari
    Snowflake SnowPro Specialty: Gen AI (GES-C02)· Lessons

    Domain 4 · Lesson 12/15

    AI_EXTRACT: response formats and prompting for document extraction

    Use document parsing functions.

    10 min read
    3.75% of exam
    2 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Call AI_EXTRACT on a staged file or a string using positional or named arguments
    • Choose the right responseFormat shape for entities, lists and tables, and follow the JSON schema rules
    • Read the JSON that AI_EXTRACT returns, including the scoring object added by scores => TRUE
    • Write extraction questions and table definitions that follow Snowflake's prompting guidance

    1.What AI_EXTRACT does and how to call it

    AI_PARSE_DOCUMENT returns the whole document as text. AI_EXTRACT returns only the values you ask for. It pulls structured information (entities, lists and tables) from text or document files, using questions in natural language or descriptions of the information you want. It runs on arctic-extract, a proprietary vision-based LLM. Because the model reads the page visually, it can also extract handwriting such as signatures, checkmarks and content inside tables.

    AI_EXTRACT accepts either a string or a FILE. Supported files include PDF, PNG, PPTX/PPT, EML, DOC/DOCX, JPEG/JPG, HTM/HTML, TEXT/TXT, TIF/TIFF, BMP, GIF, WEBP and MD, and each file must be under 100 MB. Calls can be positional, as in AI_EXTRACT(<file>, <responseFormat>), or use named arguments. Only the named form accepts the optional config and scores arguments:

    Named-argument form of AI_EXTRACT for filessql
    AI_EXTRACT( file => <file>,
                responseFormat => <responseFormat>,
                [ config => <config_object> ],
                [ scores => TRUE | FALSE ] )

    config currently has one key, scale_factor, which takes a value from 1.0 to 4.0. It scales pages up before the model processes them, which can improve OCR quality. The reference suggests it for pages larger than A4, small text or dense layouts, and output with character-level OCR errors. If you omit it, the value is 1.0.

    Checkpoint 1 of 6· Check yourself

    Extracted values from a dense, small-print engineering drawing contain character-level typos. Which documented setting addresses this?

    Sources12

    2.responseFormat: entities, lists and tables

    The responseFormat argument defines what to extract. It accepts several shapes, and which shape you use depends on whether you want single values, a list or a table. The first three shapes below are shortcuts that only extract single values. Lists and tables need a JSON schema.

    responseFormat shapes and what each extracts
    ShapeExampleExtracts
    Simple object (label → question){'name': 'What is the last name of the employee?'}Single values (entities)
    Array of strings['What is the last name of the employee?']Single values (entities)
    Array of [label, question] pairs[['name', 'What is the last name of the employee?']]Single values (entities)
    JSON schema, sub-object 'type': 'string''title': {'description': ..., 'type': 'string'}Single values (entities)
    JSON schema, sub-object 'type': 'array''employees': {'description': ..., 'type': 'array'}List of values
    JSON schema, sub-object 'type': 'object' with column_ordering'income_table': {..., 'column_ordering': ['month', 'income']}Table, returned as column arrays
    Table extraction schema: each column is an array property, and column_ordering fixes their orderjson
    { 'schema': { 'type': 'object', 'properties': { 'income_table': { 'description': 'Income for FY2026Q2', 'type': 'object', 'column_ordering': ['month', 'income'], 'properties': { 'month': { 'description': 'Month', 'type': 'array' }, 'income': { 'description': 'Income', 'type': 'array' } } } } } }

    The schema form has strict rules: - The top-level type must be an object. Each sub-object inside it is extracted independently and must be a string, a list of strings or a table. - String is the only supported scalar type, so there are no numbers or booleans. - column_ordering is case-sensitive, must match the names under properties, and should follow the column order in the document. - Once responseFormat contains a 'schema' key, every question has to be defined inside that schema. You cannot mix it with the simple object or array shapes.

    The description field is how you give the model context, for example to point it at the right table.

    Checkpoint 2 of 6· Check yourself

    A developer passes {'schema': {...list of employees...}, 'invoice_no': 'What is the invoice number?'} as responseFormat. What is wrong?

    Checkpoint 3 of 6· Exam question

    When AI_PARSE_DOCUMENT is called on a multi-page document without setting page_split (or with it left FALSE), which structure does the function return?

    Sources2

    3.Reading the response, with and without scores

    AI_EXTRACT returns a JSON object with an "error" field and a "response" field, and the response keys are your labels. An entity comes back as a string, a list as an array, and a table as an object of column arrays. If you combine all three kinds in one call, they all appear in the same response:

    Entity, list and table extracted in a single calljson
    {
      "error": null,
      "response": {
        "employees": [
          "Smith",
          "Johnson",
          "Doe"
        ],
        "income_table": {
          "income": ["$120 678","$130 123","$150 998"],
          "month": ["February", "March", "April"]
        },
        "title": "Financial report"
      }
    }

    Setting scores => TRUE adds a "scoring" object next to "response". It contains a score between 0 and 1 for each field, and a higher score means the value is more likely to be correct. Lists and tables get one aggregate score each. Individual list items and table cells are not scored. Scores cost nothing extra. A common use is a threshold that sends low-scoring extractions to human review.

    Checkpoint 4 of 6· Fill the gap

    Which argument makes this call return a scoring object?

    SELECT AI_EXTRACT(
      file => TO_FILE('@db.schema.files', 'document.pdf'),
      responseFormat => {'name': 'What is the last name of the employee?', 'date': 'What is the inspection date?'},
       ?  => TRUE
    );

    Sources2

    4.Writing questions and table definitions that extract well

    Most of the prompting work happens in the question strings and in the schema descriptions. Snowflake's guidelines for questions: - Use plain English. - Know what answer you expect from each question. - Be specific. If a document has both an issuing date and a signature date, "What is the date?" is ambiguous. - Ask for one value per question. - Do not expect the model to guess your intent or to bring specialist domain knowledge.

    For tables, the guidance is about how you name and order columns. Copy column names exactly as they appear in the document, for example Product Code rather than product_code, and keep the document's casing. List columns in the order they appear, usually left to right, and put columns whose value repeats across rows, such as Invoice Number and Invoice Date, first. Give columns meaningful names; avoid col1 or val1. For hierarchical headers, join the parent names as a prefix. If a table is divided into named sections, add a Section column.

    The description field is optional. Use it when a document contains several similar tables and the model reads the wrong one, by giving the table's title or number. Table answers are capped at 4096 tokens. If a table spans several pages, split the document into one-page documents and join the results afterwards. If a single page is still too dense, split the table's columns into two separate definitions.

    Checkpoint 5 of 6· Check yourself

    Which question follows Snowflake's AI_EXTRACT prompting guidance for a purchase agreement that contains several dates?

    Checkpoint 6 of 6· Check yourself

    A 12-page statement has one transaction table that runs across every page, and the extracted rows stop part-way through. What does the guidance recommend?

    Sources1

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.You can add a few simple label: question pairs alongside a JSON schema in the same responseFormat.Why is that wrong?

      The JSON schema format cannot be combined with other response formats. Every question has to be defined inside the schema.

      Covered in responseFormat: entities, lists and tables

    2. 2.Turning on scores adds cost, and it returns a score for every table cell.Why is that wrong?

      Scores cost nothing extra. Lists and tables get one aggregate score, and individual cells are not scored.

      Covered in Reading the response, with and without scores

    3. 3.Table columns should be given clean, database-style names such as product_code or REPORT_DATE.Why is that wrong?

      Copy the column names exactly as they appear in the document, for example Product Code.

      Covered in Writing questions and table definitions that extract well

    Practise it for real

    Extract fields with confidence scores from a staged PDF and see how the response changes

    1. 1.Run SELECT AI_EXTRACT(file => TO_FILE('@db.schema.files', 'document.pdf'), responseFormat => {'name': 'What is the last name of the employee?', 'date': 'What is the inspection date?'}, scores => TRUE); using your own stage and file.

      Why: The named-argument form is the only way to request scores.

      You should see: A JSON object with response.name, response.date, a scoring.scores entry for each field, and error set to null.

    2. 2.Run the same call with scores => FALSE, or leave the argument out.

      Why: This confirms that scores are optional and off by default.

      You should see: The same response object without a scoring object.

    3. 3.Replace responseFormat with a 'schema' that defines a table, using 'type': 'object', column_ordering and one 'type': 'array' property per column copied from a table in your document.

      Why: This exercises the table-extraction format and the rules for naming and ordering columns.

      You should see: response.<table_label> holds one array per column. With scores => TRUE, the table gets a single aggregate score.

    Stuck? Get a nudge

    If the table values come from the wrong table, add the table's title to the description field.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “AI_EXTRACT uses arctic-extract, a proprietary vision-based large language model (LLM) that delivers high extraction accuracy.”
      ↩︎ What AI_EXTRACT does and how to call it
      “Do not expect AI_EXTRACT to guess your intentions or have extended knowledge in a specific domain.”
      ↩︎ Writing questions and table definitions that extract well
      “The model for table extraction returns answers that are up to 4096 tokens long.”
      ↩︎ Writing questions and table definitions that extract well
      “To improve accuracy, define the columns in the same order as they appear in the document”
      ↩︎ Writing questions and table definitions that extract well
      “You can copy the column names from the document so that they’re exactly the same.”
      ↩︎ Exam trap 3
      “Ask for a single value in each question.”
      ↩︎ Checkpoint
      “If the table covers several pages, split the document into multiple one-page documents, and join the results in postprocessing.”
      ↩︎ Checkpoint
    2. 2.
      “The files must be less than 100 MB in size.”
      ↩︎ What AI_EXTRACT does and how to call it
      “Top level type must always be an object, which contains independently extracted sub-objects.”
      ↩︎ responseFormat: entities, lists and tables
      “String is currently the only supported scalar type.”
      ↩︎ responseFormat: entities, lists and tables
      “The column_ordering field is case-sensitive and must match the column names defined in the properties field.”
      ↩︎ responseFormat: entities, lists and tables
      “Each field in scoring.scores corresponds to a field in response and contains a score value between 0 and 1.”
      ↩︎ Reading the response, with and without scores
      “Requesting scores does not incur additional cost.”
      ↩︎ Reading the response, with and without scores
      “You can’t combine the JSON schema format with other response formats.”
      ↩︎ Exam trap 1
      “Per-element scores for individual list items and table cells are not available.”
      ↩︎ Exam trap 2
      “Scales pages of an input file before they are processed by the underlying model, which can enhance OCR quality and improve extraction results.”
      ↩︎ Checkpoint
      “If responseFormat contains the schema key, you must define all questions within the JSON schema.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 13 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.