What you will be able to do
- Explain what a pipe object is and the two ways Snowpipe learns about new files
- Contrast Snowpipe with bulk COPY loading on compute, billing, load history and transactions
- Choose between Elastic Channels and Named Channels in Snowpipe Streaming
- Decide whether files-in-cloud-storage or rows-from-applications calls for Snowpipe or Snowpipe Streaming, and place Openflow in that picture
Key concept
Pipe — A pipe is a named Snowflake object that wraps a COPY statement: it names the source and the target table. Snowpipe runs that statement whenever new files arrive, and Snowpipe Streaming uses a pipe as the route that rows take into a table.
1.Snowpipe: loading files as soon as they land
Bulk loading means someone runs COPY statements on a schedule and loads large batches. Snowpipe works differently. It loads each file once it is available in a stage, in micro-batches, so the data can be queried within minutes. The pipe object holds the COPY statement that Snowpipe runs. That statement names the stage holding the files and the target table, and every data type is supported, including semi-structured formats such as JSON and Avro.
Snowpipe can find out about new files in two ways:
- Automated loading with cloud messaging. Cloud storage sends event notifications to a queue. Snowpipe polls that queue and uses the metadata in it to load new files serverlessly. Snowflake recommends turning on cloud event filtering to cut cost, event noise and latency. - Calling the Snowpipe REST endpoints. Your client application sends Snowflake a pipe name and a list of file names. Any matching files in the pipe's stage are queued for loading. REST calls require key pair authentication with a JSON Web Token (JWT).
Each pipe has one queue. That doesn't mean files load in order: several processes pull from the queue, so older files usually load first but there is no guarantee. Snowpipe also keeps file-loading metadata for each pipe, recording the path and name of every file it has loaded. A file with the same name is not loaded again, even if its contents have changed since.
Access control is pipe-specific. Creating a pipe requires CREATE PIPE on the schema, USAGE on an external stage or READ on an internal stage, and SELECT and INSERT on the target table. A role other than the owner needs the OPERATE privilege on the pipe to pause or resume it. Pipes are managed with CREATE PIPE, ALTER PIPE, DROP PIPE, DESCRIBE PIPE and SHOW PIPES.
Checkpoint 1 of 4· Check yourself
A pipe has already loaded sales_0412.csv. An upstream job overwrites that file in the stage with corrected rows, keeping the same name. What does Snowpipe do?
Snowpipe recognises files by path and name, not by content. A file that was modified but kept its name is still treated as already loaded.
“prevents loading files with the same name even if they were later modified (i.e. have a different eTag).”Source: docs.snowflake.com
Sources1
2.How Snowpipe differs from bulk COPY
Snowpipe and bulk loading both run a COPY statement, but they differ in compute, billing, transactions and where load history is kept. Exam questions usually test one of these rows.
| Aspect | Bulk data load (COPY) | Snowpipe |
|---|---|---|
| Compute | Requires a user-specified warehouse to execute COPY statements | Uses Snowflake-supplied compute resources |
| Billing | Billed for the amount of time each virtual warehouse is active | Billed according to the compute resources used in the Snowpipe warehouse while loading the files |
| Load history | Stored in the metadata of the target table for 64 days | Stored in the metadata of the pipe for 14 days; requested via REST endpoint, SQL table function, or ACCOUNT_USAGE view |
| Transactions | Always a single transaction | Combined or split into one or more transactions based on the number and size of rows per file |
| Authentication | Security options supported by the client | REST endpoints require key pair authentication with JWT |
Two recommendations follow from this. First, don't load the same set of files with both bulk COPY and Snowpipe. Each one tracks loaded files in its own metadata (the table for bulk, the pipe for Snowpipe), so mixing them can load files twice. Second, for the best balance of cost and latency, follow the file-sizing guidance and stage files about once per minute. Cost here means the resources spent on Snowpipe queue management as well as the actual load. Latency depends on file format, file size and how complex the COPY transformation is, so Snowflake can't promise a figure. Measure it with a typical set of loads.
Checkpoint 2 of 4· Match them up
Match each property to the loading method it describes
Tap a term, then the definition that fits it.
Bulk COPY runs on your warehouse and keeps 64 days of history on the table. Snowpipe is serverless and keeps 14 days of history on the pipe.
“Stored in the metadata of the pipe for 14 days.”Source: docs.snowflake.com
Sources1
3.Snowpipe Streaming: rows, not files
Snowpipe still depends on files. Snowpipe Streaming takes that step away: applications, devices and services send rows directly into Snowflake tables or Snowflake-managed Apache Iceberg tables. Nothing has to be written to a staging file or to intermediate object storage. In either mode, rows travel through a channel, which is a logical path through a pipe to a target table.
There are two channel modes:
- Elastic Channels are the recommended starting point for new applications. Producers write without creating or coordinating channels, and Snowflake scales ingestion as traffic changes. Snowflake sends an acknowledgement once the append is durably buffered, and the producer can then drop its copy. Delivery is at-least-once, with no ordering guarantee. - Named Channels use offset tokens to give ordered, exactly-once ingestion within each channel. Use them when the source needs strict ordering, such as Kafka partitions or CDC feeds.
The service is serverless. Compute scales with ingestion load, and billing is by throughput, in credits per uncompressed GB ingested. Snowflake quotes up to 20 GB/s per table and ingest-to-queryable latency as low as 5 seconds, though actual results depend on the workload. Because the pipe holds COPY-style syntax, you can reorder columns, cast types and apply expressions before data is committed. Schema evolution can also add new columns that it detects in the stream.
You can connect through the Java, Python or Node.js SDKs (which share a Rust-based core and handle batching and compression for you), the REST API (for lightweight, IoT or edge workloads, sending NDJSON), or the Snowflake Connector for Kafka. Snowflake advises using an SDK over REST where possible.
Checkpoint 3 of 4· Check yourself
A team streams events from Kafka partitions and must not lose ordering or get duplicates when a producer restarts. Which Snowpipe Streaming option fits?
Elastic Channels are at-least-once and don't guarantee order. Named Channels use offset tokens to give ordered, exactly-once ingestion within each channel.
“Use Named Channels when reading from a source that requires strict ordering semantics, such as Kafka partitions or Change Data Capture (CDC).”Source: docs.snowflake.com
Sources2
4.Choosing a path: Snowpipe, Snowpipe Streaming, or Openflow
Snowflake presents Snowpipe and Snowpipe Streaming as complementary, and the shape of your data decides between them. If an upstream system already writes files to cloud storage and minutes of latency is fine, use Snowpipe. If data arrives as rows from applications, devices or services and needs to be available within seconds, use Snowpipe Streaming.
Openflow is the third ingestion option on the exam guide. Snowflake describes it as an integration service that connects data sources to destinations through processors, and it is built on Apache NiFi. The exam guide states that Openflow will not be tested until it is globally GA. The sources for this lesson describe only what it is, so this lesson stops there.
| Situation | Use |
|---|---|
| Pipeline already produces files in cloud storage; batch-oriented, higher-latency loading is acceptable | Snowpipe |
| Rows arrive from applications, devices, or services and need low-latency availability | Snowpipe Streaming |
| Kafka topic ingestion | Snowpipe Streaming via the Snowflake Connector for Kafka |
Checkpoint 4 of 4· Check yourself
A partner drops hourly Parquet extracts into an S3 bucket, and analysts are fine seeing the data a few minutes later. Which service is the documented fit?
The data already arrives as files in cloud storage and higher latency is acceptable, which is exactly the case Snowflake describes for Snowpipe.
“Use Snowpipe when your pipeline already produces files in cloud storage and batch-oriented, higher-latency loading is acceptable.”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Because each pipe has a single queue, Snowpipe loads files in exactly the order they were staged.Why is that wrong?
Several processes pull files from the queue. Older files usually load first, but the order is not guaranteed.
Covered in Snowpipe: loading files as soon as they land
2.Snowpipe load history sits on the target table for 64 days, just like bulk COPY history.Why is that wrong?
The 64 days on the table applies to bulk COPY. Snowpipe history is kept on the pipe for 14 days and has to be requested.
Covered in How Snowpipe differs from bulk COPY
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Snowpipe polls the event notifications from a queue.”
↩︎ Snowpipe: loading files as soon as they land“Your client application calls a public REST endpoint with the name of a pipe object and a list of data filenames.”
↩︎ Snowpipe: loading files as soon as they land“a role that has the following minimum permissions can pause or resume the pipe”
↩︎ Snowpipe: loading files as soon as they land“we recommend loading data from a specific set of files using either bulk data loading or Snowpipe but not both.”
↩︎ How Snowpipe differs from bulk COPY“Stored in the metadata of the target table for 64 days.”
↩︎ How Snowpipe differs from bulk COPY“Billed according to the compute resources used in the Snowpipe warehouse while loading the files.”
↩︎ How Snowpipe differs from bulk COPY“resources spent on Snowpipe queue management and the actual load”
↩︎ How Snowpipe differs from bulk COPY“A pipe is a named, first-class Snowflake object that contains a COPY statement used by Snowpipe.”
↩︎ Key concept“there is no guarantee that files are loaded in the same order they are staged.”
↩︎ Exam trap 1“Stored in the metadata of the pipe for 14 days.”
↩︎ Exam trap 2“while Snowpipe generally loads older files first, there is no guarantee that files are loaded in the same order they are staged.”
↩︎ Prediction“prevents loading files with the same name even if they were later modified (i.e. have a different eTag).”
↩︎ Checkpoint“Stored in the metadata of the pipe for 14 days.”
↩︎ Checkpoint - 2.https://docs.snowflake.com/en/user-guide/snowpipe-streaming/data-load-snowpipe-streaming-overviewOfficial docs
“a channel is a logical path that carries rows through a pipe to a target table”
↩︎ Snowpipe Streaming: rows, not files“Elastic Channels provide at-least-once delivery without an ordering guarantee.”
↩︎ Snowpipe Streaming: rows, not files“Throughput-based billing calculated by credits per uncompressed GB of data ingested.”
↩︎ Snowpipe Streaming: rows, not files“Snowpipe Streaming and Snowpipe complement each other.”
↩︎ Choosing a path: Snowpipe, Snowpipe Streaming, or Openflow“Use Named Channels when reading from a source that requires strict ordering semantics, such as Kafka partitions or Change Data Capture (CDC).”
↩︎ Checkpoint“Use Snowpipe when your pipeline already produces files in cloud storage and batch-oriented, higher-latency loading is acceptable.”
↩︎ Checkpoint - 3.
“Built on Apache NiFi, Openflow lets you run a fully managed service in your own cloud for complete control.”
↩︎ Choosing a path: Snowpipe, Snowpipe Streaming, or Openflow