What you will be able to do
- Choose the right Spark connector version and data transfer mode
- Read and write Snowflake data from Spark DataFrames and explain query pushdown limits
- Install the Snowflake Connector for Python with the right extras and use its PEP-249 and pandas APIs
- Describe what a native connector is and how the Native App Framework hosts it
1.The Spark connector: a plugin inside your Spark cluster
The Snowflake Connector for Spark makes Snowflake look like any other Spark data source, comparable to PostgreSQL, HDFS or S3. It runs as a Spark plugin and is distributed as the spark-snowflake package. The Spark cluster can be self-hosted or provided by a service such as AWS EMR, Databricks or Qubole. Databricks and Qubole have both integrated the connector into their platforms. Internally, the connector uses Scala 2.12.x or 2.13.x and talks to Snowflake through the Snowflake JDBC driver.
Two version lines matter. Connector 2.x supports Spark 3.2–3.4 and ships a separate build for each Spark version, so you must match the build to your Spark version. Connector 3.x supports Spark 3.2 through 4.1, and each package works with most of those versions. To enforce row access and masking policies on Iceberg tables queried through Horizon Catalog, you need version 3.1.6 or later.
The connector isn't strictly required, because third-party JDBC drivers can also connect Spark to Snowflake. Snowflake still recommends it because it is optimized for large transfers and supports query pushdown. If you want all of the processing to happen inside Snowflake, the documentation suggests Snowpark or Snowpark Connect for Spark as alternatives.
In code, you name the data source class net.snowflake.spark.snowflake (usually through the SNOWFLAKE_SOURCE_NAME constant) and pass connection options. To read, you supply either dbtable (the whole table) or query (an exact SELECT). DataFrame reads support SELECT only, not SHOW, DESC or DML. To write, you use dbtable plus a Spark SaveMode:
df.write .format(SNOWFLAKE_SOURCE_NAME) .options(sfOptions) .option("dbtable", "t2") .mode(SaveMode.Overwrite) .save()Checkpoint 1 of 9· Check yourself
A team runs Spark 3.3 and wants to install connector 2.x. What does the documentation require?
Connector 2.x has a separate build for each Spark version, so the build has to match. Single packages that cover most Spark versions are a 3.x feature.
“Use the correct version of the connector for your version of Spark.”Source: docs.snowflake.com
2.Spark transfer modes and query pushdown
Data moving between Spark and Snowflake is staged, and the connector offers two ways to do that. With internal transfer, the connector creates a Snowflake internal stage when the session starts, uses it during the session, and drops it at the end, which removes the temporary data. With external transfer, you provide the storage location yourself (an S3 bucket or an Azure Blob container) as part of installation. The connector does not delete the files it writes there. You can delete them manually, set the connector's purge parameter, or use a storage lifecycle rule.
| Aspect | Internal transfer | External transfer |
|---|---|---|
| Storage | Snowflake internal stage, created and dropped by the connector | User-created location (S3 bucket or Azure Blob container) |
| Cleanup | Automatic at end of session | Manual, purge parameter, or storage lifecycle policy |
| Column mapping (columnmapping) | Supported | Not supported |
| When to choose | Default recommendation | Connector 2.1.x or lower, or transfers of 36 hours or more |
Internal transfer also needs a minimum connector version that depends on the cloud platform of your Snowflake account: 2.2.0 or higher on AWS, 2.4.0 or higher on Azure, and 2.7.0 or higher on GCP. External transfer on Azure is likewise supported only from version 2.4.0. For external transfer, the storage location must be created and configured as part of the Spark connector installation, and the parameters that specify it are among the connector's configuration options.
Checkpoint 2 of 9· Check yourself
A nightly Spark job writing to Snowflake has grown to about 40 hours. Which transfer mode should it use?
Internal transfer relies on temporary credentials that expire after 36 hours, so a longer transfer needs a user-managed external location.
“Your transfer is likely to take 36 hours or more (internal transfer uses temporary credentials that expire after 36 hours).”Source: docs.snowflake.com
Query pushdown, available from version 2.1.0 and enabled by default, translates Spark logical plans into SQL that runs in Snowflake. Only operations with a close Snowflake equivalent are pushed down, including aggregates, filters, joins, sorts, unions, window functions and many string and math functions. Anything else runs in Spark. Spark UDFs, for example, are never pushed down. You can turn pushdown off for a session with SnowflakeConnectorUtils.disablePushdownSession, or for one DataFrame with autopushdown set to off. To see the SQL that was actually sent, call Utils.getLastSelect().
Checkpoint 3 of 9· Fill the gap
Which option name makes this read run an exact SELECT statement?
val df: DataFrame = sqlContext.read .format(SNOWFLAKE_SOURCE_NAME) .options(sfOptions) .option(" ? ", "SELECT DEPT, SUM(SALARY) AS SUM_SALARY FROM T1") .load()query runs the exact SELECT you give it. dbtable takes a table name and reads every row and column.
Source: docs.snowflake.com3.The Python connector: a pip package in your application
The Spark connector depends on Spark and JDBC. The Snowflake Connector for Python has neither dependency. It is a pure Python package with no JDBC or ODBC underneath, and it is the Python alternative to writing Java or C/C++ applications on those drivers. SnowSQL is built on it. You install it with pip on Linux, macOS or Windows:
pip install snowflake-connector-python
The connector requires Python 3.9 (marked deprecated) or later. It implements the Python DB API v2 (PEP-249). A Connection object connects to Snowflake, and Cursor objects run DDL, DML and queries. If you don't use Snowflake on AWS, you can leave out the boto3 and botocore dependencies to save disk space and memory:
Checkpoint 4 of 9· Fill the gap
Which environment variable skips the AWS libraries during installation?
? =true pip install snowflake-connector-pythonSetting SNOWFLAKE_NO_BOTO=true at install time excludes boto3 and botocore.
Source: docs.snowflake.comOptional features come as extras in square brackets. Quote the package name so your shell doesn't treat the brackets as a wildcard, and separate multiple extras with commas. The pandas extra installs the matching PyArrow version for you. Snowflake warns that a different PyArrow version should not be reinstalled afterwards. To read, you run a query on a cursor and call fetch_pandas_all() or fetch_pandas_batches(). To write, you call write_pandas(), or use pandas.DataFrame.to_sql() with pd_writer().
Checkpoint 5 of 9· Check yourself
You have run a query on a Cursor and want the result as a pandas DataFrame. Which call fits?
fetch_pandas_all() and fetch_pandas_batches() are the Cursor methods for reading into pandas. write_pandas(), pd_writer() and to_sql() move data the other way, from a DataFrame into Snowflake.
“To read data into a pandas DataFrame, you use a Cursor to retrieve the data and then call one of these Cursor methods”Source: docs.snowflake.com
pip install "snowflake-connector-python[secure-local-storage,pandas]"4.Native connectors: running inside Snowflake
The connectors above all run on infrastructure you operate: a Kafka Connect cluster, a Spark cluster, or a Python process. A native connector is a different category. Snowflake defines a connector as an application that moves data from an external source system into Snowflake. A native connector is one built and deployed with the Snowflake Native App Framework. Through that framework, providers can publish and monetize the app on Snowflake Marketplace.
Connectors follow either a pull-based or a push-based pattern. A pull-based connector uses direct external access to connect to the source application, performs outbound authentication, and fetches data directly into the customer's Snowflake account. It fits when the source data provider does not manage customer data in Snowflake. A push-based connector fits when inbound access through a customer firewall is not feasible. It uses an agent, a standalone application distributed as a Docker image and deployed in the customer environment, which sends initial and incremental loads by reading changes from a source CDC stream. The Snowflake Native SDK for Connectors, a Java library with templates and examples for building these apps, currently supports only the pull-based pattern. The SDK is built on Snowflake features such as the Native App Framework and external network access.
In the push-based design, the Snowflake Native App runs within Snowflake and coordinates the integration. It manages the replication process, controls the agent's state, and creates the required objects, including the target databases. The available sources describe the SDK and the framework. They do not give installation or scheduling steps for any specific Snowflake-provided connector, so check that connector's own documentation for those details.
The SDK documentation does show the general flow, using an example connector that ingests GitHub issues by pulling from the GitHub API. You build the connector, deploy it as an application package with snow app deploy, and install an instance with snow app run. The installed app opens a wizard that guides configuration: prerequisites, connector configuration, and connection to the source system. Configuration grants the app two account-level privileges, CREATE DATABASE (for the destination database) and EXECUTE TASK (to schedule periodic ingestion tasks), and chooses a warehouse that is referenced when scheduling those tasks. For the source connection, OAuth2 is recommended over user/password or plaintext tokens. Afterwards, in daily use, you can view statistics, configure resources for ingestion, and pause or resume the connector. The docs also use a Salesforce connector as an example of the ingestion pattern, where an External Access Integration that you explicitly authorize makes the outbound call.
Checkpoint 6 of 9· Check yourself
What makes a connector a native connector?
Being native has nothing to do with language or ingestion method. It means the connector is a Snowflake Native App, which runs inside Snowflake.
“A native connector is a connector application built and deployed using the Snowflake Native App Framework.”Source: docs.snowflake.com
The Kafka questions below draw on settings of the classic Kafka connector (v3 and earlier), which Snowflake still supports but plans to deprecate in favor of v4. In the classic connector, snowflake.ingestion.method chooses how topic data is loaded: SNOWPIPE (the default) or SNOWPIPE_STREAMING. Snowpipe Streaming calls an API to write rows directly to tables, which gives lower load latencies and lower cost for similar volumes than the Snowpipe path, where buffered records are flushed to an internal stage. The v4 connector uses the Snowpipe Streaming architecture natively, so it has no such choice.
Three classic settings control buffering. buffer.flush.time is the number of seconds between flushes from Kafka's memory buffer (default 120 seconds with Snowpipe, 10 with Snowpipe Streaming). buffer.count.records is the number of records buffered per Kafka partition before ingestion (default 10000). buffer.size.bytes is the cumulative size of buffered records per partition (default 5 MB with Snowpipe, 20 MB with Snowpipe Streaming). A flush happens when a threshold is reached. With Snowpipe Streaming, snowflake.streaming.enable.single.buffer set to true skips the connector's internal buffer, which makes buffer.flush.time and buffer.count.records irrelevant. tasks.max sets the number of tasks, usually the number of CPU cores across the Kafka Connect worker nodes. For best performance Snowflake recommends one task per Kafka partition, but no more than the CPU cores, because many tasks raise memory use and rebalances.
Checkpoint 7 of 9· Exam question
A data engineering team configures a single Kafka Connect sink connector to ingest from three topics — orders, refunds, and shipments — into three separately named Snowflake tables instead of the connector's default per-topic naming. Which configuration approach lets them redirect each topic to a specific target table without deploying three separate connectors?
Correct answer: A — Set the `snowflake.topic2table.map` property to a comma-separated list of topic:table pairs so each topic routes to its intended table.
- A. This is correct because `snowflake.topic2table.map` is the property designed to override the connector's default one-topic-to-one-table naming with an explicit mapping. It accepts comma-separated topic:table pairs, letting one connector fan messages out to differently named tables.
- B. This is incorrect because `buffer.count.records` only controls how many records accumulate before the connector flushes to Snowflake; it has no effect on which table a topic's messages land in.
- C. This is incorrect because the key converter class governs how the connector deserializes a record's key format, not which target table receives the record's rows.
- D. This is incorrect because schematization controls how the connector evolves a table's columns to match the ingested record's schema; it does not choose which table a topic writes into.
Checkpoint 8 of 9· Exam question
A streaming platform team must load Kafka messages into Snowflake with sub-second latency so downstream dashboards reflect events almost immediately, rather than tolerating the file-staging delay of periodic loads. Which ingestion mode should they configure the Kafka connector to use?
Correct answer: A — Configure the connector for Snowpipe Streaming, which writes rows directly through the low-latency streaming API without staging files.
- A. This is correct because Snowpipe Streaming ingests rows through a low-latency API that writes directly to table storage, avoiding the file-staging step and giving the sub-second latency this scenario requires.
- B. This is incorrect because the classic Snowpipe mode stages files and waits for an event notification before loading, which adds the file-staging delay the team explicitly wants to avoid.
- C. This is incorrect because batching into hourly stages with a scheduled Task introduces an hour-scale delay, far short of the sub-second freshness the dashboards need.
- D. This is incorrect because external tables read files already sitting in cloud storage and are not a Kafka connector ingestion mode at all, so this does not address loading connector records.
Checkpoint 9 of 9· Exam question
A connector deployment is flushing to Snowflake every few seconds under low traffic, creating far more small load operations than the team wants, even though the record-count threshold is rarely reached. Which setting should they raise to reduce flush frequency during quiet periods?
Correct answer: A — Increase `buffer.flush.time` so the connector waits longer before flushing when neither the record count nor byte size threshold has been hit.
- A. This is correct because `buffer.flush.time` sets the maximum time the connector waits before flushing buffered records, and raising it reduces how often small, low-traffic flushes fire.
- B. This is incorrect because the topic-to-table map only controls routing of topics to tables; changing its entries does not change how often any individual topic's buffer flushes.
- C. This is incorrect because `tasks.max` controls parallelism across Connect tasks, and adding tasks does not raise the time or size threshold that triggers each task's own flush.
- D. This is incorrect because Kafka topic replication factor is a durability setting for the topic's partitions and has no relationship to the Snowflake connector's buffer flush timing.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.With pushdown enabled, the Spark connector runs every DataFrame operation in Snowflake, including Spark UDFs.Why is that wrong?
Only operations that translate to Snowflake expressions are pushed down. Everything else, including Spark UDFs, runs in Spark.
Covered in Spark transfer modes and query pushdown
2.The Python connector is a wrapper around the Snowflake JDBC or ODBC driver.Why is that wrong?
It is a pure Python package with no JDBC or ODBC dependency.
Covered in The Python connector: a pip package in your application
3.The Snowflake Native SDK for Connectors supports both pull-based and push-based connectors.Why is that wrong?
The templates discuss both patterns, but the SDK itself currently supports only pull-based connectors.
Covered in Native connectors: running inside Snowflake
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The connector runs as a Spark plugin and is provided as a Spark package (spark-snowflake).”
↩︎ The Spark connector: a plugin inside your Spark cluster“Each Spark Connector 3 package supports most versions of Spark.”
↩︎ The Spark connector: a plugin inside your Spark cluster“Use the correct version of the connector for your version of Spark.”
↩︎ Checkpoint - 2.
“The Snowflake Connector for Spark is not strictly required to connect Snowflake and Apache Spark; other 3rd-party JDBC drivers can be used.”
↩︎ The Spark connector: a plugin inside your Spark cluster“At the end of the Snowflake session, the connector drops the stage, thereby removing all the temporary data in the stage.”
↩︎ Spark transfer modes and query pushdown“Column mapping is supported only for internal data transfer.”
↩︎ Spark transfer modes and query pushdown“The internal data transfer mode is supported only in version 2.2.0 (and higher) of the connector.”
↩︎ Spark transfer modes and query pushdown“For example, Spark UDFs cannot be pushed down to Snowflake.”
↩︎ Exam trap 1“Your transfer is likely to take 36 hours or more (internal transfer uses temporary credentials that expire after 36 hours).”
↩︎ Checkpoint - 3.
“When using DataFrames, the Snowflake connector supports SELECT queries only.”
↩︎ The Spark connector: a plugin inside your Spark cluster“When pushdown fails, the connector falls back to a less-optimized execution plan.”
↩︎ Spark transfer modes and query pushdown - 4.https://docs.snowflake.com/en/developer-guide/python-connector/python-connector-installOfficial docs
“To disable these libraries, set the SNOWFLAKE_NO_BOTO environment variable to true during installation:”
↩︎ The Python connector: a pip package in your application“Requires Python version 3.9 (deprecated) or later.”
↩︎ The Python connector: a pip package in your application - 5.
“Call the write_pandas() function.”
↩︎ The Python connector: a pip package in your application“To read data into a pandas DataFrame, you use a Cursor to retrieve the data and then call one of these Cursor methods”
↩︎ Checkpoint - 6.https://docs.snowflake.com/en/developer-guide/native-apps/connector-sdk/about-connector-sdkOfficial docs
“A Snowflake Native App runs within Snowflake and coordinates the integration.”
↩︎ Native connectors: running inside Snowflake“It is primarily responsible for managing the replication process, controlling the agent state and creating required objects, including the target databases.”
↩︎ Native connectors: running inside Snowflake“The Snowflake Native App Framework allows providers to publish and monetize a Snowflake Native App on the Snowflake Marketplace.”
↩︎ Native connectors: running inside Snowflake“An agent is a standalone application, distributed as a Docker image, that is deployed in a customer environment”
↩︎ Native connectors: running inside Snowflake“The Snowflake Native SDK for Connectors currently supports only the pull-based pattern.”
↩︎ Exam trap 3“A native connector is a connector application built and deployed using the Snowflake Native App Framework.”
↩︎ Checkpoint - 7.
“Application requires two account level permissions to operate: CREATE DATABASE and EXECUTE TASK.”
↩︎ Native connectors: running inside Snowflake“To deploy the connector execute the command: snow app deploy --connection=native_sdk_connection.”
↩︎ Native connectors: running inside Snowflake - 8.https://docs.snowflake.com/en/user-guide/snowpipe-streaming/snowpipe-streaming-classic-kafkaOfficial docs
“Specifies whether to use Snowpipe Streaming or standard Snowpipe to load your Kafka topic data.”
↩︎ Native connectors: running inside Snowflake“Note that setting this property to true makes buffer.flush.time and buffer.count.records irrelevant.”
↩︎ Native connectors: running inside Snowflake - 9.
“Number of tasks, usually the same as the number of CPU cores across the worker nodes in the Kafka Connect cluster.”
↩︎ Native connectors: running inside Snowflake
Also cited
“The connector is a native, pure Python package that has no dependencies on JDBC or ODBC.”
↩︎ Exam trap 2