CertSafari
    Snowflake SnowPro Specialty: Gen AI (GES-C02)· Lessons

    Domain 4 · Lesson 13/15

    Staging and Uploading Documents for Snowflake AI Document Functions

    Prepare and manage documents and implement extracting workflows.

    17 min read
    3.75% of exam
    8 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Explain why document functions read files from a stage through a FILE object instead of from a table
    • Create a stage whose encryption setting works with AI_PARSE_DOCUMENT, AI_EXTRACT and AI_COMPLETE
    • Upload documents with PUT, Snowsight, cloud-provider tools or the Document Processing Playground, and confirm the upload with LIST
    • Check a document against the format, size, page and resolution limits before you run a function on it
    • Grant the roles and privileges needed to run a document function against staged files

    Key concept

    Stage-resident document (FILE object) — Snowflake's document functions do not take document contents loaded into a table. They read a file where it sits on an internal or external stage, and you pass that file as a FILE object. Before any parsing or extraction can run, you have to get the document onto a stage the function can read, in a format and size the function accepts.

    1.Why document workflows start with a stage

    In most Snowflake work, you load data into tables before you query it. Document processing works differently. AI_PARSE_DOCUMENT, AI_EXTRACT and AI_COMPLETE all read a document directly from a stage. The parse-document guide says the function "processes documents stored on internal or external stages" and keeps the reading order, tables and headers. The AI_EXTRACT guide has the same design note: documents are processed straight from object storage so that data doesn't have to be moved.

    These functions take a FILE object, not a path string or a VARCHAR column. The AI_PARSE_DOCUMENT reference defines its first argument, file_object, as a FILE object that points to a document on a Snowflake stage, and it refers you to TO_FILE for building one. TO_FILE takes two arguments: the stage reference and the file's relative path on that stage. So when you prepare documents, you decide which stage holds them and what path each file has on it. The extraction query then names that stage and path.

    A FILE object built with TO_FILE from a named stage (@docs.doc_stage) and a relative file name, then passed to AI_PARSE_DOCUMENTsql
    SELECT AI_PARSE_DOCUMENT (
        TO_FILE('@docs.doc_stage','research-paper-example.pdf'),
        {'mode': 'LAYOUT' , 'page_split': true}) AS research_paper_example;

    The rest of this lesson follows from this design. Because the function reads the file in place, three things decide whether the call succeeds: how the stage is encrypted, how the file reached the stage, and whether the file itself is within the function's limits. The sections below cover each one in turn. How to choose a parsing mode is covered in a separate lesson.

    Checkpoint 1 of 9· Check yourself

    A team has 5,000 PDF contracts and wants to run AI_PARSE_DOCUMENT on them. What do they have to do first?

    Sources12

    2.Creating a stage the functions can read

    You can create a named stage in Snowsight. For an internal stage, choose Create » Stage » Snowflake Managed. For an external stage, choose Create » Stage » External Stage, then pick Amazon S3, Microsoft Azure or Google Cloud Platform. Enter the storage URL, and enable Authentication if the storage isn't public. Two settings in the internal-stage dialog matter for document work.

    The first is the directory table. It is selected by default, and you can deselect it. A directory table lets you browse the files on the stage, but it needs a warehouse to refresh, so it costs credits. You can turn it on later if you skip it now. On an external stage you can also enable auto-refresh, so the directory table updates through event notifications when files are added or removed.

    The second setting is encryption, and it is the one that causes the most problems later. Snowsight's guidance is: "To enable data access, use server-side encryption." If you leave the default, staged files are client-side encrypted. You can't change the setting later, because the encryption type cannot be changed after the stage is created. Fixing a wrong choice means creating a new stage and uploading the files again.

    How much the setting matters depends on which function will read the files. AI_PARSE_DOCUMENT and AI_EXTRACT both accept client-side or server-side encrypted stages, including in accounts that use PrivateLink or network policies that block public access to stages. AI_COMPLETE is stricter. It lists only server-side encryption, and it doesn't currently work with custom network policies. If your workflow might later send the same documents to AI_COMPLETE, server-side encryption keeps that option open.

    Stage requirements by document function
    FunctionStage encryption acceptedNetwork-policy note
    AI_PARSE_DOCUMENTClient-side or server-sideWorks with PrivateLink or network policies that restrict public stage access
    AI_EXTRACTClient-side or server-sideWorks with PrivateLink or network policies that restrict public stage access
    AI_COMPLETEServer-side onlyStage processing incompatible with custom network policies

    Checkpoint 2 of 9· Exam question

    A team needs to run AI_PARSE_DOCUMENT against a scanned engineering manual that is 120 MB. The call fails immediately with a size-related error. What is the most direct fix consistent with the function's document requirements?

    Checkpoint 3 of 9· Check yourself

    You find that a named internal stage was created with client-side encryption, and AI_COMPLETE needs to read its files. What is the correct fix?

    Sources345

    3.Uploading documents: PUT, Snowsight, cloud tools and the Playground

    Once a stage exists, how you upload depends on the stage type and the tool you use. For internal stages, the main method is the PUT command. You run it from the Snowflake CLI, SnowSQL or a driver. The prefix after PUT tells Snowflake which kind of internal stage you mean: @~ is your user stage, @%mytable is the stage attached to a table, and @ alone is a named stage.

    Checkpoint 4 of 9· Fill the gap

    This command uploads a local file to the table stage for mytable. Which command completes it?

     ?  file:///data/data.csv @%mytable;

    Snowsight is the option without SQL, and it has more restrictions than PUT. It uploads only to named internal stages. You cannot use it to upload to user stages or table stages. The maximum file size for a Snowsight upload is 250 MB. Your role needs USAGE on the database and schema and WRITE on the stage. Only one upload runs at a time: if another upload is in progress, it has to finish before you can upload more files to the stage.

    Checkpoint 5 of 9· Put it in order

    Put the Snowsight steps for uploading files to a named internal stage in order.

    1. 1.Select Upload
    2. 2.Select the files to upload, then select the database, schema and stage
    3. 3.Select Load files into a Stage
    4. 4.In the navigation menu, select Ingestion » Add Data
    5. 5.Sign in to Snowsight

    External stages don't support either method. To upload files to an external stage, use the tools from your cloud provider: Amazon S3, Microsoft Azure or Google Cloud Storage. Snowflake reads the files where they are stored. If an external stage looks empty in Snowsight, check that it has a directory table enabled and refreshed, and that the URL is correct. If the URL contains a subpath, it needs a trailing slash.

    For any upload method, check the result before running any extraction. The LIST command shows what is on a user stage (LIST @~;), a table stage (LIST @%mytable;) or a named stage (LIST @my_stage;). The same check is available in the Python API:

    Listing the files on a named stage with the Snowflake Python APIpython
    stage_files = root.databases["<database>"].schemas["<schema>"].stages["my_stage"].list_files()
    for stage_file in stage_files:
      print(stage_file)

    For trying documents out before you build a workflow, the Document Processing Playground (AI & ML » AI Studio) offers a fourth way in. It has its own limits. You can add up to 10 documents, and each file can be at most 50 MB. You can upload from your local machine, which requires a personal database, or select Add from stage to pick documents from a database, schema and stage. The Playground shows AI_EXTRACT answers, the LAYOUT (Markdown) and OCR (Text) output, and generates SQL or Python snippets you can copy. If you added the files from a stage, you can open the snippets directly in Workspaces.

    Checkpoint 6 of 9· Exam question

    Before a Snowpark pipeline can call AI_PARSE_DOCUMENT or AI_EXTRACT on newly arrived contracts, what must the engineering team do to make the files available to the function?

    Sources637

    4.Document requirements: formats, size, pages and resolution

    A file that uploads successfully may still be rejected by the function. The upload limits above (250 MB in Snowsight, 50 MB in the Playground) only control what reaches the stage. Each function then applies its own input limits. For AI_PARSE_DOCUMENT, a file over 100 MB or a document over 2,000 pages fails with an error. A 200 MB PDF uploaded through Snowsight is therefore on the stage, but it can't be parsed.

    AI_PARSE_DOCUMENT input requirements
    RequirementLimit or rule
    Maximum file size100 MB
    Maximum pages per document2,000
    Maximum page resolution10000 x 10000 pixels
    Supported file typesPDF, PPTX, DOCX, JPEG, JPG, PNG, TIFF, TIF, HTML, TXT
    Stage encryptionClient-side or server-side encryption
    Font size8 points or larger for best results
    Embedded imagesUp to 50 images per document
    page_splitPDF, PowerPoint (.pptx) and Word (.docx) only; other formats return an error

    The page_split row affects how you prepare files. The reference recommends setting page_split to TRUE to process long documents that exceed the function's token limit. However, the option works only for PDF, PPTX and DOCX files. On any other format it returns an error, so a long document in another format can't be handled this way. Accuracy has its own requirements. Text should be 8 points or larger. Decorative or script fonts may be harder to recognize, although handwriting is recognized. Page orientation is detected automatically. The function is trained on 15 languages, including English, Chinese, Hindi and Ukrainian.

    AI_EXTRACT has its own list of formats. It accepts more file types than AI_PARSE_DOCUMENT, but has the same file-size limit: the reference says "The files must be less than 100 MB in size." AI_COMPLETE's limits depend on the model. All models accept .txt, .md and .pdf. Only Claude models also take Word, Excel, CSV and .xhtml files, and their per-file limit is lower.

    Accepted formats and file-size limits across document functions
    FunctionAccepted file formatsFile-size limit
    AI_PARSE_DOCUMENTPDF, PPTX, DOCX, JPEG, JPG, PNG, TIFF, TIF, HTML, TXT100 MB
    AI_EXTRACTPDF, PNG, PPTX, PPT, EML, DOC, DOCX, JPEG, JPG, HTM, HTML, TEXT, TXT, TIF, TIFF, BMP, GIF, WEBP, MDLess than 100 MB
    AI_COMPLETE (Claude models).txt, .md, .pdf, .doc, .docx, .xls, .xlsx, .csv, .xhtml22MB
    AI_COMPLETE (gemini-3.1-pro).pdf, .txt, .md37.5MB

    Checkpoint 7 of 9· Match them up

    Match each limit to what it applies to.

    Tap a term, then the definition that fits it.

    Checkpoint 8 of 9· Exam question

    An operations team submits a single PDF containing 3,000 scanned pages to AI_PARSE_DOCUMENT and the call fails. Which change addresses the root cause?

    Sources128

    5.Access and compute for running functions on staged files

    After the files are staged and within limits, two more conditions decide whether a query runs: permissions and the warehouse. For permissions, a user with ACCOUNTADMIN must grant the SNOWFLAKE.CORTEX_USER database role to anyone who calls AI_PARSE_DOCUMENT or AI_COMPLETE. The Playground also requires a role that has CORTEX_USER. The AI_COMPLETE guide adds that users "must also have READ access to the stage and file being processed." For stage operations, Snowsight lists these privileges: WRITE to upload, READ to view files, and OWNERSHIP to edit, clone or drop the stage.

    For compute, Snowflake recommends running AI_PARSE_DOCUMENT and AI_EXTRACT queries on a warehouse no larger than MEDIUM. A larger warehouse does not make them faster. Both functions are horizontally scalable and process many documents in parallel, so more documents don't call for a bigger warehouse.

    Privileges by task
    TaskRequired privilege or role
    Upload files to an internal stage in SnowsightUSAGE on database and schema, WRITE on the stage
    View staged files in SnowsightUSAGE on database and schema, READ on the stage
    Edit, clone or drop a stage in SnowsightUSAGE on database and schema, OWNERSHIP on the stage
    Call AI_PARSE_DOCUMENT or AI_COMPLETESNOWFLAKE.CORTEX_USER database role, granted by ACCOUNTADMIN

    Checkpoint 9 of 9· Check yourself

    An analyst's AI_PARSE_DOCUMENT queries are slow on a SMALL warehouse. They have the CORTEX_USER role. Which change does Snowflake's guidance support?

    Sources158

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Snowsight can upload a document to any internal stage, including your user stage or a table stage.Why is that wrong?

      Snowsight uploads only to named internal stages. For user stages (@~) and table stages (@%table), use PUT from the CLI, SnowSQL or a driver.

      Covered in Uploading documents: PUT, Snowsight, cloud tools and the Playground

    2. 2.If Snowsight accepted a 200 MB PDF (under its 250 MB upload limit), AI_PARSE_DOCUMENT can parse it.Why is that wrong?

      The upload limit and the function's limit are separate. AI_PARSE_DOCUMENT rejects any document larger than 100 MB.

      Covered in Document requirements: formats, size, pages and resolution

    3. 3.AI_PARSE_DOCUMENT can read a stage with default (client-side) encryption, so every document function can read it.Why is that wrong?

      AI_PARSE_DOCUMENT and AI_EXTRACT accept either encryption type. AI_COMPLETE requires server-side encryption, and the encryption type can't be changed after the stage is created.

      Covered in Creating a stage the functions can read

    4. 4.page_split can split any long document AI_PARSE_DOCUMENT accepts, including HTML, TXT and images.Why is that wrong?

      page_split supports only PDF, PPTX and DOCX files. Other formats return an error.

      Covered in Document requirements: formats, size, pages and resolution

    Practise it for real

    Stage a local PDF on a server-side-encrypted named internal stage, confirm it is there, and parse it.

    1. 1.In Snowsight, select Create » Stage » Snowflake Managed. Name the stage, choose server-side encryption, and select Create.

      Why: Server-side encryption lets every document function read the files, including AI_COMPLETE, and the setting can't be changed later.

      You should see: The new stage appears under Catalog » Explorer » Stages in your schema.

    2. 2.From SnowSQL or the Snowflake CLI, run PUT file:///<path>/<your>.pdf @<your_stage>; with a PDF under 100 MB and 2,000 pages.

      Why: PUT uploads local files to an internal stage, and keeping the file within the limits avoids size and page-count errors.

      You should see: PUT reports the file as uploaded.

    3. 3.Run LIST @<your_stage>;

      Why: Confirms the file name and path before you build the FILE object.

      You should see: Your PDF is listed with its size.

    4. 4.Using a role with SNOWFLAKE.CORTEX_USER and a warehouse no larger than MEDIUM, run SELECT AI_PARSE_DOCUMENT(TO_FILE('@<your_stage>','<your>.pdf'));

      Why: TO_FILE turns the stage reference and path into the FILE object the function requires.

      You should see: A JSON result containing the document's extracted text.

    Stuck? Get a nudge

    If you get 'Provided file cannot be accessed', check that your role has the privileges it needs on the stage.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “The function processes documents stored on internal or external stages, preserving reading order and structural elements such as tables and headers.”
      ↩︎ Why document workflows start with a stage
      “The document exceeds the 2,000-page limit.”
      ↩︎ Document requirements: formats, size, pages and resolution
      “a user with the ACCOUNTADMIN role must grant the SNOWFLAKE.CORTEX_USER database role to the user who will call the function.”
      ↩︎ Access and compute for running functions on staged files
      “Documents can be processed directly from object storage to avoid unnecessary data movement.”
      ↩︎ Key concept
      “The document is larger than 100 MB.”
      ↩︎ Exam trap 2
      “Maximum file size of 104857600 bytes exceeded. The file size is {actual_size} bytes.”
      ↩︎ Checkpoint
      “Snowflake recommends executing queries that call the Cortex AI_PARSE_DOCUMENT function in a smaller warehouse (no larger than MEDIUM).”
      ↩︎ Checkpoint
    2. 2.
      “A FILE object that specifies the document to parse, stored in a Snowflake stage.”
      ↩︎ Why document workflows start with a stage
      “This feature supports only PDF, PowerPoint (.pptx), and Word (.docx) documents.”
      ↩︎ Document requirements: formats, size, pages and resolution
      “Documents in other formats return an error.”
      ↩︎ Exam trap 4
    3. 3.
      “To enable data access, use server-side encryption.”
      ↩︎ Creating a stage the functions can read
      “Directory tables let you see files on the stage, but require a warehouse and thus incur a cost.”
      ↩︎ Creating a stage the functions can read
      “use the tools provided by your external cloud service”
      ↩︎ Uploading documents: PUT, Snowsight, cloud tools and the Playground
      “If another upload is in progress, it must complete before you can upload additional files onto the stage.”
      ↩︎ Uploading documents: PUT, Snowsight, cloud tools and the Playground
      “You can’t upload files onto user stages or table stages using Snowsight.”
      ↩︎ Exam trap 1
      “Otherwise, staged files are client-side encrypted by default and unreadable when downloaded.”
      ↩︎ Prediction
      “You can’t change the encryption type after you create the stage.”
      ↩︎ Checkpoint
      “In the navigation menu, select Ingestion » Add Data.”
      ↩︎ Checkpoint
    4. 4.
      “AI_EXTRACT supports documents on stages that use client-side or server-side encryption”
      ↩︎ Creating a stage the functions can read
    5. 5.
      “Processing files from stages with AI_COMPLETE is currently incompatible with custom network policies.”
      ↩︎ Creating a stage the functions can read
      “Users must also have READ access to the stage and file being processed.”
      ↩︎ Access and compute for running functions on staged files
      “Ensure that the stage uses server-side encryption.”
      ↩︎ Exam trap 3
    6. 6.
      “Note that the @~ character combination identifies a user stage.”
      ↩︎ Uploading documents: PUT, Snowsight, cloud tools and the Playground
      “To see files that have been uploaded to a Snowflake stage, use the LIST command:”
      ↩︎ Uploading documents: PUT, Snowsight, cloud tools and the Playground

    Spotted a mistake, or was something unclear? Tell us.