What you will be able to do
- Read the error field of an AI_EXTRACT result and map each error message to its cause and fix
- Use GET_PRESIGNED_URL to open a staged document, knowing its expiration limits, required stage privileges and encryption requirement
- List the role, stage and model privileges needed to run AI_EXTRACT, fine-tune arctic-extract and run inference on a tuned model
- Estimate AI_EXTRACT cost from page tokens and scale_factor, and apply the warehouse-sizing and table-extraction best practices
- Prepare a valid fine-tuning Dataset, start a FINETUNE job for arctic-extract within its limits, and call the tuned model from AI_EXTRACT
Key concept
The AI_EXTRACT result envelope — Every AI_EXTRACT call returns a JSON object with an error field next to the response field. Troubleshooting starts there: a null error with a weak answer is a quality problem you fix with prompts, scale_factor or fine-tuning, while a non-null error is an access, format or limit problem you fix before you spend any more credits.
1.Extracting query errors and opening the document with GET_PRESIGNED_URL
When you run AI_EXTRACT over a whole directory of files, one bad document does not need to stop you from reading the others. Each result is a JSON object, and the error key sits beside response. On a successful call error is null. When you troubleshoot a batch, you separate the rows whose error is not null from the rows that came back with answers, and then you read the message.
{
"error": null,
"response": {
"title": "Financial report"
}
}Most of the documented messages point to a single fix. Some are about the file: it is missing, you are not allowed to read it, or it is in the wrong format. Some are about the request: it is empty, it is not valid JSON, it reuses a key, or it asks too many questions. The rest are about size: page count, page dimensions or file size. "Not found" and "cannot be accessed" look alike but mean different things. The first is a path problem. The second means the current user lacks privileges on the file.
| Message | Explanation | What to change |
|---|---|---|
| Provided file cannot be found. | The file was not found. | The stage path or relative path |
| Provided file cannot be accessed. | The current user does not have sufficient privileges to access the file. | The role's privileges |
| The provided file isn't in the expected format or is corrupted. | The document is corrupted or isn’t in a supported format. | The source file |
| Invalid response format. | The response format is not valid JSON. | The responseFormat argument |
| Duplicate feature name found: {feature_name}. | The response format contains one or more duplicate feature names. | The keys in responseFormat |
| Too many questions: {number} complex and {number} simple = {number} total, complex question weight {number}. | The number of questions exceeds the allowed limit. | Split the questions across calls |
| Maximum number of 125 pages exceeded. The document has {actual_pages} pages. | The document exceeds the 125-page limit. | Split the document |
| Maximum file size of 104857600 bytes exceeded. The file size is {actual_size} bytes. | The document is larger than 100 MB. | Reduce or split the file |
| Internal error. | A system error occurred. Wait and try again. If the error persists, contact Snowflake support. | Retry, then contact support |
After you have isolated a failing row, or a successful row whose answer looks wrong, the next step is to look at the document yourself. GET_PRESIGNED_URL is the file function that makes this possible. You give it a stage name and a relative path, and it returns a URL that you can open in a web browser, click in a Snowsight results table, or send to the REST API for file support. The documentation describes the two functions separately. Pairing them in a review workflow is a practical pattern, not a documented feature.
GET_PRESIGNED_URL( @<stage_name> , '<relative_file_path>' , [ <expiration_time> ] )SELECT GET_PRESIGNED_URL(@images_stage, 'us/yosemite/half_dome.jpg', 3600);The expiration_time is in seconds and defaults to 3600. The ceiling depends on the stage. It is 3600 when the stage connects to S3 through an AWS IAM role (AWS_ROLE) and for Microsoft Fabric OneLake stages, and 604800 (7 days) otherwise. On OneLake, asking for more than the limit returns an error. On an Azure external stage, the function only works when the container is accessed through a storage integration. Querying it fails if the container uses a SAS token you generated. If a stage name contains spaces or special characters, wrap it in single quotes, for example '@"my stage"'.
One behaviour catches people out during troubleshooting. The function does not check that the file exists. It will return a well-formed URL for a misspelled path, and you only find out when the browser shows a NoSuchKey error in XML.
Checkpoint 1 of 8· Check yourself
An AI_EXTRACT row returned "Provided file cannot be found." You call GET_PRESIGNED_URL with the same relative path and get a URL back. What can you conclude?
GET_PRESIGNED_URL does not check that the file exists. Getting a URL back proves nothing until you open it, and a missing file shows as NoSuchKey in the browser.
“This SQL function generates a pre-signed URL for the file path that you specify, even if the file does not exist on the stage.”Source: docs.snowflake.com
Checkpoint 2 of 8· Exam question
A pipeline step calls GET_PRESIGNED_URL against a file that a separate process deletes from an internal stage moments earlier, then hands the returned URL to a document-extraction step. The GET_PRESIGNED_URL call itself succeeds and returns a URL, but the subsequent fetch of that URL fails. What is the most likely explanation?
Correct answer: A — GET_PRESIGNED_URL generates a URL without checking whether the target file currently exists, so the fetch failed only once the deleted file was actually accessed
- A. GET_PRESIGNED_URL builds the URL from the stage and path arguments without verifying the object is present, so a URL is returned even for a file that no longer exists. The failure only surfaces later, when something actually tries to fetch content from that URL and gets a not-found style error.
- B. Missing server-side encryption would prevent the function from generating a URL at all, not cause a delayed fetch failure after a valid URL was already returned. Since the call to GET_PRESIGNED_URL succeeded, encryption was already correctly configured on the stage.
- C. An expiration value that exceeds the stage's maximum lifetime is capped or rejected at the time the URL is generated, not after a successful call. Because the URL was returned successfully, the expiration value was already within bounds.
- D. Presigned URLs can be generated for files in internal stages, and internal stages are commonly used with document-extraction workflows. This option describes a blanket restriction that does not reflect how internal stages are actually supported.
2.Requirements and privileges
Some errors, such as "cannot be accessed", come from missing privileges rather than from the document. The privileges build up in layers. To call AI_EXTRACT at all, your role needs the SNOWFLAKE.CORTEX_USER database role. Reading the file comes next, and GET_PRESIGNED_URL is a good example of how stage privileges work: on an internal stage the role needs READ, and on an external stage it needs USAGE.
Encryption is a separate requirement, and the three features you use while troubleshooting do not agree on it:
- AI_EXTRACT reads documents on stages with client-side or server-side encryption, including in accounts that restrict public network access to stages with PrivateLink or network policies.
- GET_PRESIGNED_URL requires server-side encryption on the stage. If files downloaded from an internal stage are corrupted, check that the stage was created with ENCRYPTION = (TYPE = 'SNOWFLAKE_SSE').
- Fine-tuning arctic-extract does not support client-side encrypted stages, and it does not currently work with custom network policies.
So a stage that AI_EXTRACT reads without trouble can still fail when you try to fine-tune from it or create presigned URLs for it.
Fine-tuning needs the most privileges. The ACCOUNTADMIN role must grant SNOWFLAKE.CORTEX_USER to the calling user. The job then needs access to the database and schema that hold the Dataset, to the stage that holds the documents, and to the schema where the tuned model will be created. After training, anyone who runs inference with the tuned model needs OWNERSHIP, USAGE and READ on the model object.
| Privilege | Object | Notes |
|---|---|---|
| USAGE or OWNERSHIP | DATABASE | The database that the Dataset object is stored in. |
| USAGE or OWNERSHIP | SCHEMA | The schema that the Dataset object is stored in. |
| READ or OWNERSHIP | STAGE | The internal or external named stage that stores the document files. |
| USAGE or OWNERSHIP | SCHEMA | The schema that the fine-tuned model is stored in. |
| CREATE MODEL | SCHEMA | The schema that the fine-tuned model is stored in. |
Checkpoint 3 of 8· Match them up
Match each requirement to the feature it applies to
Tap a term, then the definition that fits it.
Each feature sets its own stage requirements. AI_EXTRACT accepts both kinds of encryption, presigned URLs need SSE, and fine-tuning rejects client-side encryption and also needs CREATE MODEL where the model will live.
“Client-side encrypted stages are not supported.”Source: docs.snowflake.com
3.Cost and best-practice considerations
AI_EXTRACT bills compute for three things: pages, input prompt tokens and output tokens. For paged formats (PDF, DOCX, TIF, TIFF), each page counts as 970 tokens. Image files (JPEG, JPG, PNG) are billed as one page each, also at 970 tokens. Because of this, sending an extra 40-page appendix through the function costs money even if none of your questions are about it. Warehouse size does not help: Snowflake recommends MEDIUM or smaller because larger warehouses don’t increase performance. Requesting scores => TRUE costs nothing extra.
The config argument's scale_factor (1.0 to 4.0) enlarges each page before the model reads it. It is the setting to try when answers have character-level OCR typos, when text is small or layouts are dense, or when pages are larger than A4. It has two costs. Input tokens grow in proportion to the factor, and the maximum number of pages per document shrinks by the same factor.
| scale_factor value | Token count per page | Max. number of pages per document |
|---|---|---|
| 1.0 (default) | 970 | 125 |
| 2 | 970 * 2 = 1940 tokens | 125/2 = 62.5 (rounded down to 62) |
| 2.5 | 970 * 2.5 = 2425 tokens | 125/2.5 = 50 |
| 4 | 970 * 4 = 3880 tokens | 125/4 = 31.25 (rounded down to 31) |
Each call also has a question budget: at most 100 entity questions and at most 10 table questions. A table question weighs as much as 10 entity questions, so 4 tables plus 60 entities fill the budget exactly. Going over it triggers the "Too many questions" error from the first section. Entity answers are capped at 512 output tokens per question. Table answers are capped at 4096 tokens, and the model simply stops extracting when it reaches that limit. If a table spans several pages, split the document into one-page documents and join the results afterwards. If a single page is still too dense, split the table by columns, for example two definitions of five columns each instead of one with ten.
Most table-extraction accuracy problems can be fixed in how you write the response format:
- Keep each workload to documents of the same type.
- Copy column names exactly as they appear in the document (Report Date, not REPORT_DATE). Avoid meaningless names such as col1.
- Define columns in document order. Put columns whose value repeats across rows, such as Invoice Number and Invoice Date, first.
- Add a Section column (and a Subsection column if needed) when the table is divided into named sections.
- Flatten hierarchical headers by prefixing each column with its parent names.
- Fill in the optional description field only when similar tables confuse the model, and give the table title or number.
For entity questions, use plain English, ask for one value per question, and be specific. "What is the date?" is ambiguous on a form that has both an issue date and a signature date.
Because scores are free, you can use them to flag low-scoring extractions for human review instead of re-running whole batches. Each list or table gets a single aggregate score. Individual items and cells are not scored.
Checkpoint 4 of 8· Check yourself
A scanned 100-page PDF returns answers with character-level OCR typos. You set scale_factor to 2.0. What happens?
scale_factor multiplies the input tokens per page and divides the page ceiling by the same factor. At 2.0, a 100-page document is over the 62-page limit.
“The maximum number of pages per document that can be processed by AI_EXTRACT decreases by scale_factor.”Source: docs.snowflake.com
Checkpoint 5 of 8· Exam question
A pipeline calls GET_PRESIGNED_URL against files sitting in two different stages: one internal stage and one external stage on cloud storage. Which privilege must the calling role hold on each stage type?
Correct answer: A — READ on the internal stage and USAGE on the external stage
- A. Internal stages require the READ privilege for GET_PRESIGNED_URL, while external stages require the USAGE privilege. This pairing is correct and matches the documented privilege model for the function.
- B. This reverses the actual requirement: USAGE is what external stages need, and READ is what internal stages need, not the other way around. Applying this pairing would leave the role under-privileged on the internal stage.
- C. OWNERSHIP is a much broader grant than the function requires and is not the documented minimum privilege for either stage type. Requiring OWNERSHIP everywhere would over-provision access well beyond what GET_PRESIGNED_URL needs.
- D. SELECT is a table-level privilege and is not meaningful on a stage object, so it cannot satisfy the internal-stage requirement. The function's documented privilege model does not reference SELECT at all.
4.Fine-tuning arctic-extract models
If better prompts, table definitions and scale_factor still leave accuracy too low for a document type, the last option is to fine-tune. You train arctic-extract with the Cortex FINETUNE function on a Snowflake Dataset, then call the resulting model from AI_EXTRACT. FINETUNE has four modes: CREATE, DESCRIBE, SHOW and CANCEL. CREATE takes a model name, the base model 'arctic-extract', a training Dataset, and optionally a validation Dataset and an options JSON.
SELECT SNOWFLAKE.CORTEX.FINETUNE(
'CREATE',
'@database.schema.model_name',
'arctic-extract',
'snow://dataset/training_ds/versions/2',
'snow://dataset/validation_ds/versions/4',
'{"max_epochs": 3}'
);max_epochs takes an integer from 2 to 10. If you leave out options, the system chooses the number of epochs. The Dataset must contain three columns. File is a stage path such as @db.schema.stage/file.pdf. Prompt holds the questions, in any format that AI_EXTRACT's responseFormat accepts. Response holds the answers, matched to the questions by key. Column names are case-insensitive and can appear in any order. Extra columns are ignored, but missing any of the three makes the Dataset invalid. Documents do not all need the same questions, and the model's schema is the unique set of questions across the Dataset. Set the response to None when a document has no answer. Also add rows even where the base model already answers correctly, because that confirms the answer is right. The limits below are hard requirements, and the per-question limits differ from the per-call AI_EXTRACT budget only in that here they count unique questions across the whole Dataset.
| Constraint | Limit |
|---|---|
| Recommended minimum documents | At least 20 |
| Unique questions | 100 for entity extraction, 10 for table extraction |
| Unique document files | 1,000 (the same file may be referenced multiple times) |
| Questions × total pages | Equal or less than 50,000 |
| Pages per document | 64 in AWS US West 2 and AWS Europe Central 1; 125 in AWS US East 1 and Azure East US 2 |
| File formats | PDF, PNG, JPG/JPEG, TIFF/TIF |
Checkpoint 6 of 8· Put it in order
Put the documented steps for training an arctic-extract model from scratch in order
- 1.ALTER DATASET ... ADD VERSION from the table, building paths with FL_GET_STAGE and FL_GET_RELATIVE_PATH
- 2.CREATE OR REPLACE DATASET
- 3.Insert rows using TO_FILE, a prompt JSON and a response JSON
- 4.Create a table with FILE, prompt and response columns
- 5.Call SNOWFLAKE.CORTEX.FINETUNE('CREATE', ...) on the Dataset version
FINETUNE takes a versioned Dataset as input, so the data table is filled first, then the Dataset is created, then a version is added, and only then is the job started.
“Create a new version of the Dataset that adds the training data”Source: docs.snowflake.com
FINETUNE('DESCRIBE') reports on a job: its status, its progress (1.0 when finished), the number of trained tokens, and a training_result with training_loss and validation_loss. You can only see a validation loss if you supplied a validation Dataset. When the status is SUCCESS, pass the model to AI_EXTRACT with the model argument. You can override the questions it was trained on with responseFormat, combine it with config => {'scale_factor': 2.0}, and request scores, which are supported for fine-tuned models. A tuned model can also be copied between databases, schemas and accounts.
Checkpoint 7 of 8· Fill the gap
Which named argument tells AI_EXTRACT to use your fine-tuned model?
SELECT AI_EXTRACT(
? => 'db.schema.my_tuned_model',
file => TO_FILE('@db.schema.files','document.pdf')
);The tuned model is passed with the model argument. No responseFormat is required, because the model already carries the questions it was trained on.
Source: docs.snowflake.comCheckpoint 8 of 8· Exam question
A document team wants to fine-tune an arctic-extract model so it better recognizes their company's invoice layout, but they currently have only 8 sample invoices available. Following Snowflake's guidance, what should they do before starting the fine-tuning job?
Correct answer: A — Gather additional invoices until reaching Snowflake's recommended minimum of roughly 20 documents, since fine-tuning quality depends on enough representative training examples
- A. Snowflake recommends using at least around 20 documents when fine-tuning arctic-extract models, since too few examples limits how well the model generalizes to layout variation. Expanding the sample set before training is the documented, supported path.
- B. Snowflake does publish a recommended minimum document count for fine-tuning, so this claim of no minimum requirement is incorrect. Training on far fewer documents than recommended risks a poorly adapted model.
- C. Duplicating the same handful of invoices adds rows without adding new layout variation, so it does not address the underlying lack of representative training data. The model would still only have seen 8 distinct documents.
- D. Combining files into one PDF reduces rather than increases the number of distinct training documents available to the model. Question count does not substitute for having a sufficiently diverse set of source documents.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If GET_PRESIGNED_URL returns a URL, the file exists on the stage.Why is that wrong?
The function returns a URL for any path you give it. You only find out the file is missing when you open the URL and get a NoSuchKey error.
Covered in Extracting query errors and opening the document with GET_PRESIGNED_URL
2.If AI_EXTRACT can read documents on a stage, fine-tuning can use that stage too.Why is that wrong?
AI_EXTRACT accepts client-side encrypted stages and restricted networks, but fine-tuning does not. It rejects client-side encrypted stages and does not work with custom network policies.
Covered in Requirements and privileges
3.A larger warehouse makes AI_EXTRACT run faster.Why is that wrong?
Snowflake recommends a warehouse no larger than MEDIUM, because larger warehouses do not improve AI_EXTRACT performance.
Covered in Cost and best-practice considerations
4.A table question uses one slot of the 100-question budget, just like an entity question.Why is that wrong?
A table question counts as 10 entity questions, so a few tables quickly use up the per-call budget and trigger the Too many questions error.
Covered in Cost and best-practice considerations
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The current user does not have sufficient privileges to access the file.”
↩︎ Extracting query errors and opening the document with GET_PRESIGNED_URL“Users must use a role that has been granted the SNOWFLAKE.CORTEX_USER database role.”
↩︎ Requirements and privileges“The Cortex AI_EXTRACT function incurs compute cost based on the number of pages per document, input prompt tokens, and output tokens processed.”
↩︎ Cost and best-practice considerations“The number of input tokens consumed increases proportionally with scale_factor.”
↩︎ Cost and best-practice considerations“Requesting scores does not incur additional cost.”
↩︎ Cost and best-practice considerations“Scores are supported for fine-tuned models.”
↩︎ Fine-tuning arctic-extract models“AI_EXTRACT can produce the following error messages:”
↩︎ Key concept“Larger warehouses don’t increase performance.”
↩︎ Exam trap 3“A table extraction question is equal to 10 entity extraction questions.”
↩︎ Exam trap 4“Snowflake recommends executing queries that call the Cortex AI_EXTRACT function in a smaller warehouse (no larger than MEDIUM).”
↩︎ Prediction“The maximum number of pages per document that can be processed by AI_EXTRACT decreases by scale_factor.”
↩︎ Checkpoint - 2.
“Navigate to the pre-signed URL directly in a web browser.”
↩︎ Extracting query errors and opening the document with GET_PRESIGNED_URL“Otherwise, the maximum expiration time is 604800 (7 days).”
↩︎ Extracting query errors and opening the document with GET_PRESIGNED_URL“Server-side encryption is required on the internal or external stage.”
↩︎ Requirements and privileges“If the file does not exist, the browser returns a NoSuchKey error in XML format.”
↩︎ Exam trap 1“This SQL function generates a pre-signed URL for the file path that you specify, even if the file does not exist on the stage.”
↩︎ Checkpoint - 3.
“AI_EXTRACT supports documents on stages that use client-side or server-side encryption”
↩︎ Requirements and privileges“If the table covers several pages, split the document into multiple one-page documents, and join the results in postprocessing.”
↩︎ Cost and best-practice considerations“If the answers come from a different source table than expected or the model can’t find the table, try using the description field.”
↩︎ Cost and best-practice considerations - 4.
“the ACCOUNTADMIN role must grant the SNOWFLAKE.CORTEX_USER database role to the user who will call the function”
↩︎ Requirements and privileges“To use the fine-tuned arctic-extract model for inference, ensure you have the following privileges on the model object:”
↩︎ Requirements and privileges“You can specify max_epochs with an integer from 2 through 10 (inclusive) to control how many epochs the job runs.”
↩︎ Fine-tuning arctic-extract models“all required columns (File, Prompt, and Response) must be present for the Dataset to be valid”
↩︎ Fine-tuning arctic-extract models“Number of questions multiplied by total number of pages in all document files in the Dataset must be equal or less than 50,000.”
↩︎ Fine-tuning arctic-extract models“Fine-tuning arctic-extract models is currently incompatible with custom network policies.”
↩︎ Exam trap 2“Client-side encrypted stages are not supported.”
↩︎ Checkpoint“Create a new version of the Dataset that adds the training data”
↩︎ Checkpoint