What you will be able to do
- Load structured table data into ML code with DataConnector and write results back with DataSink
- Choose between internal and external stages for file-based training data, and keep staged-file storage costs under control
- Explain what a data share costs the consumer and what governance limits it imposes
- Pick a storage method for a training or inference dataset based on performance, cost and governance requirements
1.Tables: load with DataConnector, write with DataSink
When training or inference data sits in Snowflake tables, Snowflake ML provides one API to read it and one to write it. Snowflake Notebooks and Snowflake ML Jobs both run on Container Runtime, which uses Ray to process data across multiple compute nodes. In that environment, DataConnector loads structured data from tables and Snowflake Datasets, and DataSink writes structured data back to tables.
DataConnector speeds up loading by parallelizing reads across the compute nodes. The docs say it outperforms to_pandas on large datasets. The loaded data can become a pandas DataFrame, a PyTorch dataset or a TensorFlow dataset. For the best performance, you can instead pass the connector directly to Snowflake's distributed training APIs.
| Data type | Data source | API for loading | API for writing |
|---|---|---|---|
| Structured | Snowflake Tables | DataConnector | DataSink |
| Structured | Snowflake Datasets | DataConnector | DataSink |
| Structured | CSV Files (Stage) | DataSource API | DataSink |
| Structured | Parquet Files (Stage) | DataSource API | DataSink |
| Unstructured | Other Staged Files | DataSource API | N/A |
A connector can be built from either source, and the docs pair each with a stage of the lifecycle. A Snowpark DataFrame gives direct access to the table and is best used during development. A Snowflake Dataset is a versioned, schema-level object and is best used for production workflows.
Checkpoint 1 of 6· Fill the gap
Which factory method builds a DataConnector from a Snowpark DataFrame over a table?
from snowflake.ml.data.data_connector import DataConnector
from snowflake.snowpark.context import get_active_session
session = get_active_session()
# Create DataConnector from a Snowflake table
data_connector = DataConnector. ? (session.table("example-table-name"))session.table() returns a Snowpark DataFrame, so the factory is from_dataframe. from_dataset takes a Snowflake Dataset, and to_pandas converts an existing connector.
Source: docs.snowflake.com2.Internal and external stages for file-based data
Data that arrives as files, such as CSV, Parquet or images, lives in a stage. Snowflake has two kinds. An internal stage stores the files inside Snowflake. An external stage points to files outside Snowflake, in Amazon S3 buckets, Google Cloud Storage buckets or Microsoft Azure containers, and the location can be private/protected or public. External stages have one catch: data in archival storage classes that need restoring before retrieval, such as S3 Glacier Deep Archive, can't be read through them.
You can read staged files in two ways. In Snowpark, session.read plus a format method turns a staged file into a DataFrame, which you can then feed into the preprocessors and transformations above. In Container Runtime, the DataSource APIs load CSV, Parquet, images and other formats from stages. Two optimizations help when reading staged files. Stage sharding consolidates many small files into larger shards to cut per-file I/O. A disk cache, enabled by default, serves repeated reads from node-local disk rather than fetching again from Snowflake.
df_json = session.read.json("@my_stage2/data1.json")Internal stages also have a cost angle. Files in an internal stage incur standard data storage costs but none of the extra Time Travel and Fail-safe costs. Snowflake recommends removing staged files once they are loaded and no longer needed, either during COPY INTO or afterwards with REMOVE. Regular purging can also improve data loading performance.
Checkpoint 2 of 6· Check yourself
A team keeps raw training files in an internal stage after loading them into tables. What does Snowflake say about the cost of those files?
Internally staged files are billed as standard storage and are exempt from Time Travel and Fail-safe. That is why Snowflake recommends purging them once they are loaded.
“internal stages are not subject to the additional costs associated with Time Travel and Fail-safe, but they do incur standard data storage costs”Source: docs.snowflake.com
3.Data shares for training data owned by another account
When training data belongs to another Snowflake account, Secure Data Sharing gives you access without moving the data. A provider creates a share containing selected objects from one of their databases, and the consumer imports it. Shareable objects include tables, dynamic tables, external tables, Iceberg tables, regular and secure views, secure materialized views, and UDFs.
No data is copied between accounts, because sharing runs through Snowflake's services layer and metadata store. Two consequences follow for ML data. The shared data takes up no storage in the consumer account, so the consumer pays only for the compute that queries it. And access is near-instant once the share is set up. The governance trade-off is control: every shared object is read-only for the consumer. You can query a shared table to build features, but you can't modify it or add data to it.
Checkpoint 3 of 6· Check yourself
A consumer team wants to add engineered feature columns to a table imported through a data share. What happens?
Shared database objects are read-only for the consumer and no local copy exists. Engineered features have to be written to an object the consumer owns.
“All database objects shared between accounts are read-only”Source: docs.snowflake.com
Sources5
4.Matching storage to performance, cost and governance
For a given training or inference dataset, choose the storage method by the requirement that matters most. The table below collects the deciding facts for each option. It covers only what the sources here state. In particular, they say little about governance controls on ordinary tables, so that cell describes how the data is accessed rather than which policies apply.
| Storage method | Performance for ML loading | Cost | Governance and control |
|---|---|---|---|
| Snowflake table | DataConnector parallelizes reads across nodes; DataSink writes back | Queried with your own warehouse compute | Read through a Snowpark DataFrame (best during development) or a Dataset (best for production) |
| Internal stage | DataSource APIs; stage sharding and default disk cache | Standard storage, no Time Travel or Fail-safe costs; purge after loading | Files stored internally within Snowflake; a named stage is a database object, so access control privileges govern who can create, modify, use or drop it |
| External stage | DataSource APIs; archival storage classes cannot be read | Files stay in your S3, GCS or Azure storage | Location can be private/protected or public; a storage integration is recommended over supplying credentials |
| Data share | Near-instant access with no data copy | No consumer storage; consumer pays only for query compute | Read-only for the consumer; consumer access to the imported database uses standard role-based access control; the provider can revoke access at any time |
Checkpoint 4 of 6· Match them up
Match each requirement to the storage method that fits it best
Tap a term, then the definition that fits it.
Shares avoid consumer storage. External stages reference cloud storage outside Snowflake. Internal stages hold files within Snowflake. Tables pair with DataConnector and DataSink.
“Shared data does not take up any storage in a consumer account”Source: docs.snowflake.com
Checkpoint 5 of 6· Check yourself
A data provider shares training data with a partner account, and later the agreement ends. The provider wants the partner to lose access without copying or deleting anything on the partner's side. What can the provider do?
Shares are controlled completely by the provider account, and access to a share or any object in it can be revoked at any time. No data copy exists in the consumer account.
“Access to a share (or any of the objects in a share) can be revoked at any time.”Source: docs.snowflake.com
Checkpoint 6 of 6· Exam question
Which Python module contains the scikit-learn-style distributed preprocessors such as `MinMaxScaler`, `OneHotEncoder`, and `OrdinalEncoder` that execute inside Snowflake?
Correct answer: A — snowflake.ml.modeling.preprocessing
- A. Correct: the Snowflake ML preprocessing classes live in snowflake.ml.modeling.preprocessing and mirror a subset of scikit-learn's preprocessors.
- B. Incorrect: snowflake.snowpark.functions holds SQL-function wrappers for DataFrame columns and has no preprocessing submodule of transformers.
- C. Incorrect: the Feature Store package manages entities and feature views, not the scaler and encoder classes.
- D. Incorrect: snowflake.ml.data covers data connectors and ingestion helpers rather than fit/transform preprocessors.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Consuming a data share copies the provider's tables into your account, so you pay storage for them.Why is that wrong?
Sharing copies nothing. Shared data takes up no storage in the consumer account, and the consumer pays only for the warehouses that query it.
Covered in Data shares for training data owned by another account
2.Files left in an internal stage carry Time Travel and Fail-safe costs, just like table data.Why is that wrong?
Internally staged files incur only standard storage costs. Snowflake still recommends removing them after loading to control cost and improve loading performance.
3.A DataConnector over a live Snowpark DataFrame is the recommended input for production training.Why is that wrong?
The docs pair Snowpark DataFrames with development. For production workflows they recommend Snowflake Datasets, which are versioned.
Covered in Tables: load with DataConnector, write with DataSink
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The DataConnector accelerates data loading by parallelizing the reads across multiple compute nodes.”
↩︎ Tables: load with DataConnector, write with DataSink“Snowpark DataFrames: Provide direct access to the data in your Snowflake tables. Best used during development.”
↩︎ Tables: load with DataConnector, write with DataSink“Stage sharding: Cut per-file I/O when reading many small files by consolidating them into larger shards.”
↩︎ Internal and external stages for file-based data“Snowflake Datasets: Versioned schema-level objects. Best used for production workflows.”
↩︎ Exam trap 3 - 2.
“it provides improved performance over to_pandas for loading large datasets.”
↩︎ Tables: load with DataConnector, write with DataSink - 3.
“References data files stored in a location outside of Snowflake.”
↩︎ Internal and external stages for file-based data“You cannot access data held in archival cloud storage classes that requires restoration before it can be retrieved.”
↩︎ Internal and external stages for file-based data“The storage location can be either private/protected or public.”
↩︎ Matching storage to performance, cost and governance - 4.
“Periodic purging of staged files can have other benefits, such as improved data loading performance.”
↩︎ Internal and external stages for file-based data“remove them from the stages once the data has been loaded and the files are no longer needed”
↩︎ Exam trap 2“internal stages are not subject to the additional costs associated with Time Travel and Fail-safe, but they do incur standard data storage costs”
↩︎ Checkpoint - 5.
“Secure Data Sharing lets you share selected objects in a database in your account with other Snowflake accounts.”
↩︎ Data shares for training data owned by another account“access to the imported data is near-instantaneous for consumers”
↩︎ Data shares for training data owned by another account“The only charges to consumers are for the compute resources (i.e. virtual warehouses) used to query the imported data.”
↩︎ Matching storage to performance, cost and governance“Access to a share (or any of the objects in a share) can be revoked at any time.”
↩︎ Matching storage to performance, cost and governance“Access to this database is configurable using the same, standard role-based access control that Snowflake provides for all objects in the system.”
↩︎ Matching storage to performance, cost and governance“With Secure Data Sharing, no actual data is copied or transferred between accounts.”
↩︎ Exam trap 1“The only charges to consumers are for the compute resources (i.e. virtual warehouses) used to query the imported data.”
↩︎ Prediction“All database objects shared between accounts are read-only”
↩︎ Checkpoint“Shared data does not take up any storage in a consumer account”
↩︎ Checkpoint - 6.
“the ability to create, modify, use, or drop them can be controlled using security access control privileges”
↩︎ Matching storage to performance, cost and governance - 7.
“We highly recommend the use of storage integrations.”
↩︎ Matching storage to performance, cost and governance