CertSafari
    Snowflake SnowPro Advanced: Data Engineer (DEA-C02)· Lessons

    Domain 1 · Lesson 5/22

    Spark, Python, and Native Connectors for Snowflake

    Install, configure, and use connectors for Snowflake integration.

    14 min read
    4% of exam
    10 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Choose the right Spark connector version and data transfer mode
    • Read and write Snowflake data from Spark DataFrames and explain query pushdown limits
    • Install the Snowflake Connector for Python with the right extras and use its PEP-249 and pandas APIs
    • Describe what a native connector is and how the Native App Framework hosts it

    1.The Spark connector: a plugin inside your Spark cluster

    The Snowflake Connector for Spark makes Snowflake look like any other Spark data source, comparable to PostgreSQL, HDFS or S3. It runs as a Spark plugin and is distributed as the spark-snowflake package. The Spark cluster can be self-hosted or provided by a service such as AWS EMR, Databricks or Qubole. Databricks and Qubole have both integrated the connector into their platforms. Internally, the connector uses Scala 2.12.x or 2.13.x and talks to Snowflake through the Snowflake JDBC driver.

    Two version lines matter. Connector 2.x supports Spark 3.2–3.4 and ships a separate build for each Spark version, so you must match the build to your Spark version. Connector 3.x supports Spark 3.2 through 4.1, and each package works with most of those versions. To enforce row access and masking policies on Iceberg tables queried through Horizon Catalog, you need version 3.1.6 or later.

    The connector isn't strictly required, because third-party JDBC drivers can also connect Spark to Snowflake. Snowflake still recommends it because it is optimized for large transfers and supports query pushdown. If you want all of the processing to happen inside Snowflake, the documentation suggests Snowpark or Snowpark Connect for Spark as alternatives.

    In code, you name the data source class net.snowflake.spark.snowflake (usually through the SNOWFLAKE_SOURCE_NAME constant) and pass connection options. To read, you supply either dbtable (the whole table) or query (an exact SELECT). DataFrame reads support SELECT only, not SHOW, DESC or DML. To write, you use dbtable plus a Spark SaveMode:

    Writing a Spark DataFrame to a Snowflake tablescala
    df.write .format(SNOWFLAKE_SOURCE_NAME) .options(sfOptions) .option("dbtable", "t2") .mode(SaveMode.Overwrite) .save()

    Checkpoint 1 of 9· Check yourself

    A team runs Spark 3.3 and wants to install connector 2.x. What does the documentation require?

    Sources123

    2.Spark transfer modes and query pushdown

    Data moving between Spark and Snowflake is staged, and the connector offers two ways to do that. With internal transfer, the connector creates a Snowflake internal stage when the session starts, uses it during the session, and drops it at the end, which removes the temporary data. With external transfer, you provide the storage location yourself (an S3 bucket or an Azure Blob container) as part of installation. The connector does not delete the files it writes there. You can delete them manually, set the connector's purge parameter, or use a storage lifecycle rule.

    Internal vs external transfer
    AspectInternal transferExternal transfer
    StorageSnowflake internal stage, created and dropped by the connectorUser-created location (S3 bucket or Azure Blob container)
    CleanupAutomatic at end of sessionManual, purge parameter, or storage lifecycle policy
    Column mapping (columnmapping)SupportedNot supported
    When to chooseDefault recommendationConnector 2.1.x or lower, or transfers of 36 hours or more

    Internal transfer also needs a minimum connector version that depends on the cloud platform of your Snowflake account: 2.2.0 or higher on AWS, 2.4.0 or higher on Azure, and 2.7.0 or higher on GCP. External transfer on Azure is likewise supported only from version 2.4.0. For external transfer, the storage location must be created and configured as part of the Spark connector installation, and the parameters that specify it are among the connector's configuration options.

    Checkpoint 2 of 9· Check yourself

    A nightly Spark job writing to Snowflake has grown to about 40 hours. Which transfer mode should it use?

    Query pushdown, available from version 2.1.0 and enabled by default, translates Spark logical plans into SQL that runs in Snowflake. Only operations with a close Snowflake equivalent are pushed down, including aggregates, filters, joins, sorts, unions, window functions and many string and math functions. Anything else runs in Spark. Spark UDFs, for example, are never pushed down. You can turn pushdown off for a session with SnowflakeConnectorUtils.disablePushdownSession, or for one DataFrame with autopushdown set to off. To see the SQL that was actually sent, call Utils.getLastSelect().

    Checkpoint 3 of 9· Fill the gap

    Which option name makes this read run an exact SELECT statement?

    val df: DataFrame = sqlContext.read .format(SNOWFLAKE_SOURCE_NAME) .options(sfOptions) .option(" ? ", "SELECT DEPT, SUM(SALARY) AS SUM_SALARY FROM T1") .load()

    Sources23

    3.The Python connector: a pip package in your application

    The Spark connector depends on Spark and JDBC. The Snowflake Connector for Python has neither dependency. It is a pure Python package with no JDBC or ODBC underneath, and it is the Python alternative to writing Java or C/C++ applications on those drivers. SnowSQL is built on it. You install it with pip on Linux, macOS or Windows:

    pip install snowflake-connector-python

    The connector requires Python 3.9 (marked deprecated) or later. It implements the Python DB API v2 (PEP-249). A Connection object connects to Snowflake, and Cursor objects run DDL, DML and queries. If you don't use Snowflake on AWS, you can leave out the boto3 and botocore dependencies to save disk space and memory:

    Checkpoint 4 of 9· Fill the gap

    Which environment variable skips the AWS libraries during installation?

     ? =true pip install snowflake-connector-python

    Optional features come as extras in square brackets. Quote the package name so your shell doesn't treat the brackets as a wildcard, and separate multiple extras with commas. The pandas extra installs the matching PyArrow version for you. Snowflake warns that a different PyArrow version should not be reinstalled afterwards. To read, you run a query on a cursor and call fetch_pandas_all() or fetch_pandas_batches(). To write, you call write_pandas(), or use pandas.DataFrame.to_sql() with pd_writer().

    Checkpoint 5 of 9· Check yourself

    You have run a query on a Cursor and want the result as a pandas DataFrame. Which call fits?

    Installing the connector with pandas and secure-local-storage extrasbash
    pip install "snowflake-connector-python[secure-local-storage,pandas]"

    Sources45

    4.Native connectors: running inside Snowflake

    The connectors above all run on infrastructure you operate: a Kafka Connect cluster, a Spark cluster, or a Python process. A native connector is a different category. Snowflake defines a connector as an application that moves data from an external source system into Snowflake. A native connector is one built and deployed with the Snowflake Native App Framework. Through that framework, providers can publish and monetize the app on Snowflake Marketplace.

    Connectors follow either a pull-based or a push-based pattern. A pull-based connector uses direct external access to connect to the source application, performs outbound authentication, and fetches data directly into the customer's Snowflake account. It fits when the source data provider does not manage customer data in Snowflake. A push-based connector fits when inbound access through a customer firewall is not feasible. It uses an agent, a standalone application distributed as a Docker image and deployed in the customer environment, which sends initial and incremental loads by reading changes from a source CDC stream. The Snowflake Native SDK for Connectors, a Java library with templates and examples for building these apps, currently supports only the pull-based pattern. The SDK is built on Snowflake features such as the Native App Framework and external network access.

    In the push-based design, the Snowflake Native App runs within Snowflake and coordinates the integration. It manages the replication process, controls the agent's state, and creates the required objects, including the target databases. The available sources describe the SDK and the framework. They do not give installation or scheduling steps for any specific Snowflake-provided connector, so check that connector's own documentation for those details.

    The SDK documentation does show the general flow, using an example connector that ingests GitHub issues by pulling from the GitHub API. You build the connector, deploy it as an application package with snow app deploy, and install an instance with snow app run. The installed app opens a wizard that guides configuration: prerequisites, connector configuration, and connection to the source system. Configuration grants the app two account-level privileges, CREATE DATABASE (for the destination database) and EXECUTE TASK (to schedule periodic ingestion tasks), and chooses a warehouse that is referenced when scheduling those tasks. For the source connection, OAuth2 is recommended over user/password or plaintext tokens. Afterwards, in daily use, you can view statistics, configure resources for ingestion, and pause or resume the connector. The docs also use a Salesforce connector as an example of the ingestion pattern, where an External Access Integration that you explicitly authorize makes the outbound call.

    Checkpoint 6 of 9· Check yourself

    What makes a connector a native connector?

    The Kafka questions below draw on settings of the classic Kafka connector (v3 and earlier), which Snowflake still supports but plans to deprecate in favor of v4. In the classic connector, snowflake.ingestion.method chooses how topic data is loaded: SNOWPIPE (the default) or SNOWPIPE_STREAMING. Snowpipe Streaming calls an API to write rows directly to tables, which gives lower load latencies and lower cost for similar volumes than the Snowpipe path, where buffered records are flushed to an internal stage. The v4 connector uses the Snowpipe Streaming architecture natively, so it has no such choice.

    Three classic settings control buffering. buffer.flush.time is the number of seconds between flushes from Kafka's memory buffer (default 120 seconds with Snowpipe, 10 with Snowpipe Streaming). buffer.count.records is the number of records buffered per Kafka partition before ingestion (default 10000). buffer.size.bytes is the cumulative size of buffered records per partition (default 5 MB with Snowpipe, 20 MB with Snowpipe Streaming). A flush happens when a threshold is reached. With Snowpipe Streaming, snowflake.streaming.enable.single.buffer set to true skips the connector's internal buffer, which makes buffer.flush.time and buffer.count.records irrelevant. tasks.max sets the number of tasks, usually the number of CPU cores across the Kafka Connect worker nodes. For best performance Snowflake recommends one task per Kafka partition, but no more than the CPU cores, because many tasks raise memory use and rebalances.

    Checkpoint 7 of 9· Exam question

    A data engineering team configures a single Kafka Connect sink connector to ingest from three topics — orders, refunds, and shipments — into three separately named Snowflake tables instead of the connector's default per-topic naming. Which configuration approach lets them redirect each topic to a specific target table without deploying three separate connectors?

    Checkpoint 8 of 9· Exam question

    A streaming platform team must load Kafka messages into Snowflake with sub-second latency so downstream dashboards reflect events almost immediately, rather than tolerating the file-staging delay of periodic loads. Which ingestion mode should they configure the Kafka connector to use?

    Checkpoint 9 of 9· Exam question

    A connector deployment is flushing to Snowflake every few seconds under low traffic, creating far more small load operations than the team wants, even though the record-count threshold is rarely reached. Which setting should they raise to reduce flush frequency during quiet periods?

    Sources6789

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.With pushdown enabled, the Spark connector runs every DataFrame operation in Snowflake, including Spark UDFs.Why is that wrong?

      Only operations that translate to Snowflake expressions are pushed down. Everything else, including Spark UDFs, runs in Spark.

      Covered in Spark transfer modes and query pushdown

    2. 2.The Python connector is a wrapper around the Snowflake JDBC or ODBC driver.Why is that wrong?

      It is a pure Python package with no JDBC or ODBC dependency.

      Covered in The Python connector: a pip package in your application

    3. 3.The Snowflake Native SDK for Connectors supports both pull-based and push-based connectors.Why is that wrong?

      The templates discuss both patterns, but the SDK itself currently supports only pull-based connectors.

      Covered in Native connectors: running inside Snowflake

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “The connector runs as a Spark plugin and is provided as a Spark package (spark-snowflake).”
      ↩︎ The Spark connector: a plugin inside your Spark cluster
      “Each Spark Connector 3 package supports most versions of Spark.”
      ↩︎ The Spark connector: a plugin inside your Spark cluster
      “Use the correct version of the connector for your version of Spark.”
      ↩︎ Checkpoint
    2. 2.
      “The Snowflake Connector for Spark is not strictly required to connect Snowflake and Apache Spark; other 3rd-party JDBC drivers can be used.”
      ↩︎ The Spark connector: a plugin inside your Spark cluster
      “At the end of the Snowflake session, the connector drops the stage, thereby removing all the temporary data in the stage.”
      ↩︎ Spark transfer modes and query pushdown
      “Column mapping is supported only for internal data transfer.”
      ↩︎ Spark transfer modes and query pushdown
      “The internal data transfer mode is supported only in version 2.2.0 (and higher) of the connector.”
      ↩︎ Spark transfer modes and query pushdown
      “For example, Spark UDFs cannot be pushed down to Snowflake.”
      ↩︎ Exam trap 1
      “Your transfer is likely to take 36 hours or more (internal transfer uses temporary credentials that expire after 36 hours).”
      ↩︎ Checkpoint
    3. 3.
      “When using DataFrames, the Snowflake connector supports SELECT queries only.”
      ↩︎ The Spark connector: a plugin inside your Spark cluster
      “When pushdown fails, the connector falls back to a less-optimized execution plan.”
      ↩︎ Spark transfer modes and query pushdown
    4. 4.
      “To disable these libraries, set the SNOWFLAKE_NO_BOTO environment variable to true during installation:”
      ↩︎ The Python connector: a pip package in your application
      “Requires Python version 3.9 (deprecated) or later.”
      ↩︎ The Python connector: a pip package in your application
    5. 5.
      “To read data into a pandas DataFrame, you use a Cursor to retrieve the data and then call one of these Cursor methods”
      ↩︎ Checkpoint
    6. 6.
      “A Snowflake Native App runs within Snowflake and coordinates the integration.”
      ↩︎ Native connectors: running inside Snowflake
      “It is primarily responsible for managing the replication process, controlling the agent state and creating required objects, including the target databases.”
      ↩︎ Native connectors: running inside Snowflake
      “The Snowflake Native App Framework allows providers to publish and monetize a Snowflake Native App on the Snowflake Marketplace.”
      ↩︎ Native connectors: running inside Snowflake
      “An agent is a standalone application, distributed as a Docker image, that is deployed in a customer environment”
      ↩︎ Native connectors: running inside Snowflake
      “The Snowflake Native SDK for Connectors currently supports only the pull-based pattern.”
      ↩︎ Exam trap 3
      “A native connector is a connector application built and deployed using the Snowflake Native App Framework.”
      ↩︎ Checkpoint
    7. 7.
      “Application requires two account level permissions to operate: CREATE DATABASE and EXECUTE TASK.”
      ↩︎ Native connectors: running inside Snowflake
      “To deploy the connector execute the command: snow app deploy --connection=native_sdk_connection.”
      ↩︎ Native connectors: running inside Snowflake
    8. 8.
      “Specifies whether to use Snowpipe Streaming or standard Snowpipe to load your Kafka topic data.”
      ↩︎ Native connectors: running inside Snowflake
      “Note that setting this property to true makes buffer.flush.time and buffer.count.records irrelevant.”
      ↩︎ Native connectors: running inside Snowflake
    9. 9.
      “Number of tasks, usually the same as the number of CPU cores across the worker nodes in the Kafka Connect cluster.”
      ↩︎ Native connectors: running inside Snowflake

    Also cited

    Spotted a mistake, or was something unclear? Tell us.