What you will be able to do
- Call AI_EXTRACT on a staged file or a string using positional or named arguments
- Choose the right responseFormat shape for entities, lists and tables, and follow the JSON schema rules
- Read the JSON that AI_EXTRACT returns, including the scoring object added by scores => TRUE
- Write extraction questions and table definitions that follow Snowflake's prompting guidance
1.What AI_EXTRACT does and how to call it
AI_PARSE_DOCUMENT returns the whole document as text. AI_EXTRACT returns only the values you ask for. It pulls structured information (entities, lists and tables) from text or document files, using questions in natural language or descriptions of the information you want. It runs on arctic-extract, a proprietary vision-based LLM. Because the model reads the page visually, it can also extract handwriting such as signatures, checkmarks and content inside tables.
AI_EXTRACT accepts either a string or a FILE. Supported files include PDF, PNG, PPTX/PPT, EML, DOC/DOCX, JPEG/JPG, HTM/HTML, TEXT/TXT, TIF/TIFF, BMP, GIF, WEBP and MD, and each file must be under 100 MB. Calls can be positional, as in AI_EXTRACT(<file>, <responseFormat>), or use named arguments. Only the named form accepts the optional config and scores arguments:
AI_EXTRACT( file => <file>,
responseFormat => <responseFormat>,
[ config => <config_object> ],
[ scores => TRUE | FALSE ] )config currently has one key, scale_factor, which takes a value from 1.0 to 4.0. It scales pages up before the model processes them, which can improve OCR quality. The reference suggests it for pages larger than A4, small text or dense layouts, and output with character-level OCR errors. If you omit it, the value is 1.0.
Checkpoint 1 of 6· Check yourself
Extracted values from a dense, small-print engineering drawing contain character-level typos. Which documented setting addresses this?
scale_factor scales pages up before the model processes them, which targets small text and OCR errors. 'mode' is an AI_PARSE_DOCUMENT option, and scores only report confidence.
“Scales pages of an input file before they are processed by the underlying model, which can enhance OCR quality and improve extraction results.”Source: docs.snowflake.com
2.responseFormat: entities, lists and tables
The responseFormat argument defines what to extract. It accepts several shapes, and which shape you use depends on whether you want single values, a list or a table. The first three shapes below are shortcuts that only extract single values. Lists and tables need a JSON schema.
| Shape | Example | Extracts |
|---|---|---|
| Simple object (label → question) | {'name': 'What is the last name of the employee?'} | Single values (entities) |
| Array of strings | ['What is the last name of the employee?'] | Single values (entities) |
| Array of [label, question] pairs | [['name', 'What is the last name of the employee?']] | Single values (entities) |
| JSON schema, sub-object 'type': 'string' | 'title': {'description': ..., 'type': 'string'} | Single values (entities) |
| JSON schema, sub-object 'type': 'array' | 'employees': {'description': ..., 'type': 'array'} | List of values |
| JSON schema, sub-object 'type': 'object' with column_ordering | 'income_table': {..., 'column_ordering': ['month', 'income']} | Table, returned as column arrays |
{ 'schema': { 'type': 'object', 'properties': { 'income_table': { 'description': 'Income for FY2026Q2', 'type': 'object', 'column_ordering': ['month', 'income'], 'properties': { 'month': { 'description': 'Month', 'type': 'array' }, 'income': { 'description': 'Income', 'type': 'array' } } } } } }The schema form has strict rules: - The top-level type must be an object. Each sub-object inside it is extracted independently and must be a string, a list of strings or a table. - String is the only supported scalar type, so there are no numbers or booleans. - column_ordering is case-sensitive, must match the names under properties, and should follow the column order in the document. - Once responseFormat contains a 'schema' key, every question has to be defined inside that schema. You cannot mix it with the simple object or array shapes.
The description field is how you give the model context, for example to point it at the right table.
Checkpoint 2 of 6· Check yourself
A developer passes {'schema': {...list of employees...}, 'invoice_no': 'What is the invoice number?'} as responseFormat. What is wrong?
When responseFormat contains a schema key, every question must be defined in that schema and no other keys are allowed. Number would also be invalid, because string is the only scalar type.
“If responseFormat contains the schema key, you must define all questions within the JSON schema.”Source: docs.snowflake.com
Checkpoint 3 of 6· Exam question
When AI_PARSE_DOCUMENT is called on a multi-page document without setting page_split (or with it left FALSE), which structure does the function return?
Correct answer: A — A single object with a content field holding the document's extracted text or markdown as one combined string
- A. Without page_split enabled, AI_PARSE_DOCUMENT returns a single content field containing the whole document's extracted text or markdown as one combined value.
- B. A pages array with per-page content and index fields is the response shape produced when page_split is set to TRUE, not the default behavior.
- C. PageCount only appears inside the metadata object when return_error_details is TRUE, and it is unrelated to whether page_split was configured.
- D. AI_PARSE_DOCUMENT does not return per-paragraph confidence scores in its response; that granularity and field are not part of the documented output.
Sources2
3.Reading the response, with and without scores
AI_EXTRACT returns a JSON object with an "error" field and a "response" field, and the response keys are your labels. An entity comes back as a string, a list as an array, and a table as an object of column arrays. If you combine all three kinds in one call, they all appear in the same response:
{
"error": null,
"response": {
"employees": [
"Smith",
"Johnson",
"Doe"
],
"income_table": {
"income": ["$120 678","$130 123","$150 998"],
"month": ["February", "March", "April"]
},
"title": "Financial report"
}
}Setting scores => TRUE adds a "scoring" object next to "response". It contains a score between 0 and 1 for each field, and a higher score means the value is more likely to be correct. Lists and tables get one aggregate score each. Individual list items and table cells are not scored. Scores cost nothing extra. A common use is a threshold that sends low-scoring extractions to human review.
Checkpoint 4 of 6· Fill the gap
Which argument makes this call return a scoring object?
SELECT AI_EXTRACT(
file => TO_FILE('@db.schema.files', 'document.pdf'),
responseFormat => {'name': 'What is the last name of the employee?', 'date': 'What is the inspection date?'},
? => TRUE
);scores => TRUE adds the scoring object and is supported only in named-argument syntax. return_error_details belongs to AI_PARSE_DOCUMENT, and scale_factor goes inside config.
Source: docs.snowflake.comSources2
4.Writing questions and table definitions that extract well
Most of the prompting work happens in the question strings and in the schema descriptions. Snowflake's guidelines for questions: - Use plain English. - Know what answer you expect from each question. - Be specific. If a document has both an issuing date and a signature date, "What is the date?" is ambiguous. - Ask for one value per question. - Do not expect the model to guess your intent or to bring specialist domain knowledge.
For tables, the guidance is about how you name and order columns. Copy column names exactly as they appear in the document, for example Product Code rather than product_code, and keep the document's casing. List columns in the order they appear, usually left to right, and put columns whose value repeats across rows, such as Invoice Number and Invoice Date, first. Give columns meaningful names; avoid col1 or val1. For hierarchical headers, join the parent names as a prefix. If a table is divided into named sections, add a Section column.
The description field is optional. Use it when a document contains several similar tables and the model reads the wrong one, by giving the table's title or number. Table answers are capped at 4096 tokens. If a table spans several pages, split the document into one-page documents and join the results afterwards. If a single page is still too dense, split the table's columns into two separate definitions.
Checkpoint 5 of 6· Check yourself
Which question follows Snowflake's AI_EXTRACT prompting guidance for a purchase agreement that contains several dates?
It is specific about which date, and it asks for a single value. The other options are ambiguous, ask for several values at once, or rely on domain knowledge the guidance says not to expect.
“Ask for a single value in each question.”Source: docs.snowflake.com
Checkpoint 6 of 6· Check yourself
A 12-page statement has one transaction table that runs across every page, and the extracted rows stop part-way through. What does the guidance recommend?
Table extraction stops at a 4096-token answer, so a long table has to be split across separate calls. Generic column names and an unordered schema both make accuracy worse.
“If the table covers several pages, split the document into multiple one-page documents, and join the results in postprocessing.”Source: docs.snowflake.com
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.You can add a few simple label: question pairs alongside a JSON schema in the same responseFormat.Why is that wrong?
The JSON schema format cannot be combined with other response formats. Every question has to be defined inside the schema.
Covered in responseFormat: entities, lists and tables
2.Turning on scores adds cost, and it returns a score for every table cell.Why is that wrong?
Scores cost nothing extra. Lists and tables get one aggregate score, and individual cells are not scored.
3.Table columns should be given clean, database-style names such as product_code or REPORT_DATE.Why is that wrong?
Copy the column names exactly as they appear in the document, for example Product Code.
Covered in Writing questions and table definitions that extract well
Practise it for real
Extract fields with confidence scores from a staged PDF and see how the response changes
1.Run SELECT AI_EXTRACT(file => TO_FILE('@db.schema.files', 'document.pdf'), responseFormat => {'name': 'What is the last name of the employee?', 'date': 'What is the inspection date?'}, scores => TRUE); using your own stage and file.
Why: The named-argument form is the only way to request scores.
You should see: A JSON object with response.name, response.date, a scoring.scores entry for each field, and error set to null.
2.Run the same call with scores => FALSE, or leave the argument out.
Why: This confirms that scores are optional and off by default.
You should see: The same response object without a scoring object.
3.Replace responseFormat with a 'schema' that defines a table, using 'type': 'object', column_ordering and one 'type': 'array' property per column copied from a table in your document.
Why: This exercises the table-extraction format and the rules for naming and ordering columns.
You should see: response.<table_label> holds one array per column. With scores => TRUE, the table gets a single aggregate score.
Stuck? Get a nudge
If the table values come from the wrong table, add the table's title to the description field.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“AI_EXTRACT uses arctic-extract, a proprietary vision-based large language model (LLM) that delivers high extraction accuracy.”
↩︎ What AI_EXTRACT does and how to call it“Do not expect AI_EXTRACT to guess your intentions or have extended knowledge in a specific domain.”
↩︎ Writing questions and table definitions that extract well“The model for table extraction returns answers that are up to 4096 tokens long.”
↩︎ Writing questions and table definitions that extract well“To improve accuracy, define the columns in the same order as they appear in the document”
↩︎ Writing questions and table definitions that extract well“You can copy the column names from the document so that they’re exactly the same.”
↩︎ Exam trap 3“Ask for a single value in each question.”
↩︎ Checkpoint“If the table covers several pages, split the document into multiple one-page documents, and join the results in postprocessing.”
↩︎ Checkpoint - 2.
“The files must be less than 100 MB in size.”
↩︎ What AI_EXTRACT does and how to call it“Top level type must always be an object, which contains independently extracted sub-objects.”
↩︎ responseFormat: entities, lists and tables“String is currently the only supported scalar type.”
↩︎ responseFormat: entities, lists and tables“The column_ordering field is case-sensitive and must match the column names defined in the properties field.”
↩︎ responseFormat: entities, lists and tables“Each field in scoring.scores corresponds to a field in response and contains a score value between 0 and 1.”
↩︎ Reading the response, with and without scores“Requesting scores does not incur additional cost.”
↩︎ Reading the response, with and without scores“You can’t combine the JSON schema format with other response formats.”
↩︎ Exam trap 1“Per-element scores for individual list items and table cells are not available.”
↩︎ Exam trap 2“Scales pages of an input file before they are processed by the underlying model, which can enhance OCR quality and improve extraction results.”
↩︎ Checkpoint“If responseFormat contains the schema key, you must define all questions within the JSON schema.”
↩︎ Checkpoint