CertSafari
    Snowflake SnowPro Advanced: MLOps Engineer (MLA-B01)· Lessons

    Domain 1 · Lesson 1/17

    Storing and Loading ML Training and Inference Data: Tables, Stages and Data Shares

    Construct distributed feature engineering pipelines.

    9 min read
    4% of exam
    7 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Load structured table data into ML code with DataConnector and write results back with DataSink
    • Choose between internal and external stages for file-based training data, and keep staged-file storage costs under control
    • Explain what a data share costs the consumer and what governance limits it imposes
    • Pick a storage method for a training or inference dataset based on performance, cost and governance requirements

    1.Tables: load with DataConnector, write with DataSink

    When training or inference data sits in Snowflake tables, Snowflake ML provides one API to read it and one to write it. Snowflake Notebooks and Snowflake ML Jobs both run on Container Runtime, which uses Ray to process data across multiple compute nodes. In that environment, DataConnector loads structured data from tables and Snowflake Datasets, and DataSink writes structured data back to tables.

    DataConnector speeds up loading by parallelizing reads across the compute nodes. The docs say it outperforms to_pandas on large datasets. The loaded data can become a pandas DataFrame, a PyTorch dataset or a TensorFlow dataset. For the best performance, you can instead pass the connector directly to Snowflake's distributed training APIs.

    Which Snowflake ML API loads and writes each data source
    Data typeData sourceAPI for loadingAPI for writing
    StructuredSnowflake TablesDataConnectorDataSink
    StructuredSnowflake DatasetsDataConnectorDataSink
    StructuredCSV Files (Stage)DataSource APIDataSink
    StructuredParquet Files (Stage)DataSource APIDataSink
    UnstructuredOther Staged FilesDataSource APIN/A

    A connector can be built from either source, and the docs pair each with a stage of the lifecycle. A Snowpark DataFrame gives direct access to the table and is best used during development. A Snowflake Dataset is a versioned, schema-level object and is best used for production workflows.

    Checkpoint 1 of 6· Fill the gap

    Which factory method builds a DataConnector from a Snowpark DataFrame over a table?

    from snowflake.ml.data.data_connector import DataConnector
    from snowflake.snowpark.context import get_active_session
    
    session = get_active_session()
    
    # Create DataConnector from a Snowflake table
    data_connector = DataConnector. ? (session.table("example-table-name"))

    Sources12

    2.Internal and external stages for file-based data

    Data that arrives as files, such as CSV, Parquet or images, lives in a stage. Snowflake has two kinds. An internal stage stores the files inside Snowflake. An external stage points to files outside Snowflake, in Amazon S3 buckets, Google Cloud Storage buckets or Microsoft Azure containers, and the location can be private/protected or public. External stages have one catch: data in archival storage classes that need restoring before retrieval, such as S3 Glacier Deep Archive, can't be read through them.

    You can read staged files in two ways. In Snowpark, session.read plus a format method turns a staged file into a DataFrame, which you can then feed into the preprocessors and transformations above. In Container Runtime, the DataSource APIs load CSV, Parquet, images and other formats from stages. Two optimizations help when reading staged files. Stage sharding consolidates many small files into larger shards to cut per-file I/O. A disk cache, enabled by default, serves repeated reads from node-local disk rather than fetching again from Snowflake.

    Reading a staged JSON file into a Snowpark DataFramepython
    df_json = session.read.json("@my_stage2/data1.json")

    Internal stages also have a cost angle. Files in an internal stage incur standard data storage costs but none of the extra Time Travel and Fail-safe costs. Snowflake recommends removing staged files once they are loaded and no longer needed, either during COPY INTO or afterwards with REMOVE. Regular purging can also improve data loading performance.

    Checkpoint 2 of 6· Check yourself

    A team keeps raw training files in an internal stage after loading them into tables. What does Snowflake say about the cost of those files?

    Sources314

    3.Data shares for training data owned by another account

    When training data belongs to another Snowflake account, Secure Data Sharing gives you access without moving the data. A provider creates a share containing selected objects from one of their databases, and the consumer imports it. Shareable objects include tables, dynamic tables, external tables, Iceberg tables, regular and secure views, secure materialized views, and UDFs.

    No data is copied between accounts, because sharing runs through Snowflake's services layer and metadata store. Two consequences follow for ML data. The shared data takes up no storage in the consumer account, so the consumer pays only for the compute that queries it. And access is near-instant once the share is set up. The governance trade-off is control: every shared object is read-only for the consumer. You can query a shared table to build features, but you can't modify it or add data to it.

    Checkpoint 3 of 6· Check yourself

    A consumer team wants to add engineered feature columns to a table imported through a data share. What happens?

    Sources5

    4.Matching storage to performance, cost and governance

    For a given training or inference dataset, choose the storage method by the requirement that matters most. The table below collects the deciding facts for each option. It covers only what the sources here state. In particular, they say little about governance controls on ordinary tables, so that cell describes how the data is accessed rather than which policies apply.

    Storage options for ML data, compared on the factors these sources document
    Storage methodPerformance for ML loadingCostGovernance and control
    Snowflake tableDataConnector parallelizes reads across nodes; DataSink writes backQueried with your own warehouse computeRead through a Snowpark DataFrame (best during development) or a Dataset (best for production)
    Internal stageDataSource APIs; stage sharding and default disk cacheStandard storage, no Time Travel or Fail-safe costs; purge after loadingFiles stored internally within Snowflake; a named stage is a database object, so access control privileges govern who can create, modify, use or drop it
    External stageDataSource APIs; archival storage classes cannot be readFiles stay in your S3, GCS or Azure storageLocation can be private/protected or public; a storage integration is recommended over supplying credentials
    Data shareNear-instant access with no data copyNo consumer storage; consumer pays only for query computeRead-only for the consumer; consumer access to the imported database uses standard role-based access control; the provider can revoke access at any time

    Checkpoint 4 of 6· Match them up

    Match each requirement to the storage method that fits it best

    Tap a term, then the definition that fits it.

    Checkpoint 5 of 6· Check yourself

    A data provider shares training data with a partner account, and later the agreement ends. The provider wants the partner to lose access without copying or deleting anything on the partner's side. What can the provider do?

    Checkpoint 6 of 6· Exam question

    Which Python module contains the scikit-learn-style distributed preprocessors such as `MinMaxScaler`, `OneHotEncoder`, and `OrdinalEncoder` that execute inside Snowflake?

    Sources3567

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Consuming a data share copies the provider's tables into your account, so you pay storage for them.Why is that wrong?

      Sharing copies nothing. Shared data takes up no storage in the consumer account, and the consumer pays only for the warehouses that query it.

      Covered in Data shares for training data owned by another account

    2. 2.Files left in an internal stage carry Time Travel and Fail-safe costs, just like table data.Why is that wrong?

      Internally staged files incur only standard storage costs. Snowflake still recommends removing them after loading to control cost and improve loading performance.

      Covered in Internal and external stages for file-based data

    3. 3.A DataConnector over a live Snowpark DataFrame is the recommended input for production training.Why is that wrong?

      The docs pair Snowpark DataFrames with development. For production workflows they recommend Snowflake Datasets, which are versioned.

      Covered in Tables: load with DataConnector, write with DataSink

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “The DataConnector accelerates data loading by parallelizing the reads across multiple compute nodes.”
      ↩︎ Tables: load with DataConnector, write with DataSink
      “Snowpark DataFrames: Provide direct access to the data in your Snowflake tables. Best used during development.”
      ↩︎ Tables: load with DataConnector, write with DataSink
      “Stage sharding: Cut per-file I/O when reading many small files by consolidating them into larger shards.”
      ↩︎ Internal and external stages for file-based data
      “Snowflake Datasets: Versioned schema-level objects. Best used for production workflows.”
      ↩︎ Exam trap 3
    2. 2.
      “it provides improved performance over to_pandas for loading large datasets.”
      ↩︎ Tables: load with DataConnector, write with DataSink
    3. 3.
      “References data files stored in a location outside of Snowflake.”
      ↩︎ Internal and external stages for file-based data
      “You cannot access data held in archival cloud storage classes that requires restoration before it can be retrieved.”
      ↩︎ Internal and external stages for file-based data
      “The storage location can be either private/protected or public.”
      ↩︎ Matching storage to performance, cost and governance
    4. 4.
      “Periodic purging of staged files can have other benefits, such as improved data loading performance.”
      ↩︎ Internal and external stages for file-based data
      “remove them from the stages once the data has been loaded and the files are no longer needed”
      ↩︎ Exam trap 2
      “internal stages are not subject to the additional costs associated with Time Travel and Fail-safe, but they do incur standard data storage costs”
      ↩︎ Checkpoint
    5. 5.
      “Secure Data Sharing lets you share selected objects in a database in your account with other Snowflake accounts.”
      ↩︎ Data shares for training data owned by another account
      “access to the imported data is near-instantaneous for consumers”
      ↩︎ Data shares for training data owned by another account
      “The only charges to consumers are for the compute resources (i.e. virtual warehouses) used to query the imported data.”
      ↩︎ Matching storage to performance, cost and governance
      “Access to a share (or any of the objects in a share) can be revoked at any time.”
      ↩︎ Matching storage to performance, cost and governance
      “Access to this database is configurable using the same, standard role-based access control that Snowflake provides for all objects in the system.”
      ↩︎ Matching storage to performance, cost and governance
      “With Secure Data Sharing, no actual data is copied or transferred between accounts.”
      ↩︎ Exam trap 1
      “The only charges to consumers are for the compute resources (i.e. virtual warehouses) used to query the imported data.”
      ↩︎ Prediction
      “All database objects shared between accounts are read-only”
      ↩︎ Checkpoint
      “Shared data does not take up any storage in a consumer account”
      ↩︎ Checkpoint
    6. 6.
      “the ability to create, modify, use, or drop them can be controlled using security access control privileges”
      ↩︎ Matching storage to performance, cost and governance

    Ready to test yourself?

    Practise the 15 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.