CertSafari
    Snowflake SnowPro Core Certification (COF-C03)· Lessons

    Domain 1 · Lesson 1/19

    Snowflake Architecture: Storage, Compute and Cloud Services Layers

    Describe and use the Snowflake architecture

    8 min read
    5.17% of exam
    4 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Explain how Snowflake combines shared-disk and shared-nothing designs
    • Describe what the database storage layer holds and what Snowflake manages for you
    • Explain why virtual warehouses in the compute layer do not affect each other
    • List the services that run in the cloud services layer

    Key concept

    Separation of storage and compute — Snowflake keeps all persisted data in one central storage layer and runs queries on separate, independent compute clusters. Because the two are decoupled, any number of warehouses can work on the same data without copying it, and each can be sized or stopped on its own.

    1.A hybrid of shared-disk and shared-nothing

    Snowflake is a self-managed service that runs entirely on public cloud infrastructure. You don't select, install or tune hardware, and you can't run Snowflake on-premises or on a private cloud. All three layers of the architecture (storage, compute and cloud services) are deployed on the cloud platform you pick for the account: Amazon Web Services, Google Cloud or Microsoft Azure.

    Snowflake takes something from each traditional design. From shared-disk, it uses a central data repository for persisted data, and every compute node in the platform can reach it. From shared-nothing, it processes queries on massively parallel processing (MPP) compute clusters, where each node in a cluster stores a portion of the data set locally while it works. According to the documentation, the result offers "the data management simplicity of a shared-disk architecture, but with the performance and scale-out benefits of a shared-nothing architecture."

    This is why the design is often called a multi-cluster shared data architecture. One copy of the data sits in the central store, and many independent compute clusters read it at the same time. Two warehouses querying the same table don't need private copies, because both read from the shared repository. The rest of this page goes through the three layers that make this work: database storage, compute and cloud services.

    What Snowflake borrows from each traditional architecture
    Traditional designWhat Snowflake takes from itBenefit
    Shared-diskA central data repository for persisted data, accessible from all compute nodesData management simplicity
    Shared-nothingMPP compute clusters where each node stores a portion of the data set locallyPerformance and scale-out

    Checkpoint 1 of 5· Check yourself

    Which part of Snowflake's architecture comes from the shared-disk tradition?

    Checkpoint 2 of 5· Exam question

    A data engineer runs a query against a 50-billion-row fact table filtered on `order_date`. Snowflake returns results in under two seconds after examining only a small fraction of the table's micro-partitions, based on the min/max values recorded for each partition. Which architectural layer determines which micro-partitions can be skipped for this query?

    Sources12

    2.The database storage layer

    The database storage layer is the central repository from the previous section. When you load data into a Snowflake table, Snowflake doesn't keep it in the shape it arrived in. It reorganizes the data into its own internally optimized, compressed, columnar format and stores that in cloud storage.

    After that, Snowflake manages everything about how the data is stored: the organization, file size, structure, compression, metadata and statistics. You don't choose file layouts or tune compression. All data in Snowflake tables is automatically divided into micro-partitions, which are contiguous units of storage. Micro-partitions have their own lesson, but for architecture purposes the point is that this physical layout is Snowflake's job, not yours.

    The storage layer holds structured data such as rows and columns, semi-structured data such as JSON, and unstructured data through the FILE data type. Apache Iceberg tables are a useful contrast: they keep their data and metadata files in external cloud storage that you manage, and that storage isn't part of Snowflake.

    Checkpoint 3 of 5· Check yourself

    A team loads CSV files into a standard Snowflake table. Which statement describes how that data is stored?

    Sources1

    3.The compute layer: virtual warehouses

    Queries run in the compute layer, on virtual warehouses. A virtual warehouse is a cluster of compute resources. It supplies the CPU, memory and temporary storage needed to run SELECT statements and DML: inserting, updating and deleting rows, loading data with COPY INTO <table>, and unloading data with COPY INTO <location>. Warehouses are required for queries and for all DML operations. Using Snowpark, warehouses can also run code in languages such as Java, Python and Scala.

    The property that matters most for the architecture is isolation. Each virtual warehouse is an independent compute cluster that doesn't share compute resources with any other warehouse, so one warehouse has no effect on the performance of another. Since every warehouse reads from the same central storage layer, a loading job and a reporting dashboard can each have their own warehouse, work on the same tables, and not compete for CPU.

    Compute is also where the cost is. A warehouse has to be running and in use for the session to do work, and it consumes Snowflake credits while it runs. Warehouses can be started, stopped and resized at any time, even while running. Because storage is separate, stopping a warehouse doesn't touch the data.

    Checkpoint 4 of 5· Check yourself

    A data-loading warehouse is running a very heavy COPY INTO job. What effect does this have on an analyst's queries running on a different warehouse?

    Sources34

    4.The cloud services layer

    Storage holds the data and warehouses do the work, but something has to coordinate them. That is the cloud services layer: a set of services that ties Snowflake's components together to handle user requests, from sign-in to query dispatch. It runs on compute instances that Snowflake provisions from the cloud provider, not on your virtual warehouses.

    The documentation lists these services in this layer:

    - Security, authentication and access control - Snowflake Horizon Catalog - Infrastructure management with cloud platforms - Metadata management, including the SNOWFLAKE database and the Snowflake Information Schema - Query parsing and optimization - Regulatory compliance

    Put together, a query passes through all three layers. Cloud services authenticate the user and parse and optimize the SQL. A virtual warehouse in the compute layer executes it. The data comes from the database storage layer.

    Checkpoint 5 of 5· Match them up

    Match each responsibility to the Snowflake layer that handles it

    Tap a term, then the definition that fits it.

    Sources1

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Snowflake is a shared-nothing architecture, so each compute node permanently owns its slice of the data.Why is that wrong?

      Snowflake is a hybrid. Persisted data lives in a central repository that all compute nodes can reach, as in shared-disk, while queries run on MPP clusters whose nodes hold a portion of the data locally.

      Covered in A hybrid of shared-disk and shared-nothing

    2. 2.Virtual warehouses draw from a shared compute pool, so a heavy job on one warehouse slows queries on the others.Why is that wrong?

      Each warehouse is an independent compute cluster that doesn't share compute resources with other warehouses.

      Covered in The compute layer: virtual warehouses

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “offers the data management simplicity of a shared-disk architecture, but with the performance and scale-out benefits of a shared-nothing architecture”
      ↩︎ A hybrid of shared-disk and shared-nothing
      “each node in the cluster stores a portion of the entire data set locally”
      ↩︎ A hybrid of shared-disk and shared-nothing
      “You can’t install and run Snowflake locally or on private cloud infrastructures, whether on-premises or hosted.”
      ↩︎ A hybrid of shared-disk and shared-nothing
      “Snowflake manages all aspects of how this data is stored — including the organization, file size, structure, compression, metadata, and statistics.”
      ↩︎ The database storage layer
      “All data in Snowflake tables is automatically divided into micro-partitions, which are contiguous units of storage.”
      ↩︎ The database storage layer
      “The external storage isn’t part of Snowflake.”
      ↩︎ The database storage layer
      “The cloud services layer also runs on compute instances that are provisioned by Snowflake from the cloud provider.”
      ↩︎ The cloud services layer
      “Metadata management, including the SNOWFLAKE database and the Snowflake Information Schema”
      ↩︎ The cloud services layer
      “Snowflake separates storage and compute, which simplifies some traditional challenges of data engineering, such as infrastructure management and performance tuning.”
      ↩︎ Key concept
      “Snowflake’s architecture is a hybrid of traditional shared-disk and shared-nothing database architectures.”
      ↩︎ Exam trap 1
      “Each virtual warehouse is an independent compute cluster that doesn’t share compute resources with other virtual warehouses.”
      ↩︎ Exam trap 2
      “Snowflake’s architecture is a hybrid of traditional shared-disk and shared-nothing database architectures.”
      ↩︎ Prediction
      “Similar to shared-disk architectures, Snowflake uses a central data repository for persisted data that is accessible from all compute nodes in the platform.”
      ↩︎ Checkpoint
      “Snowflake reorganizes that data into its internally optimized, compressed, columnar format. Snowflake stores this optimized data in cloud storage.”
      ↩︎ Checkpoint
      “As a result, each virtual warehouse has no effect on the performance of other virtual warehouses.”
      ↩︎ Checkpoint
      “These services tie together all of the different components of Snowflake in order to process user requests, from sign-in to query dispatch.”
      ↩︎ Checkpoint
    2. 2.
      “all three layers of Snowflake’s architecture (storage, compute, and cloud services) are deployed and managed entirely on a selected cloud platform.”
      ↩︎ A hybrid of shared-disk and shared-nothing
    3. 3.
      “A warehouse provides the required resources, such as CPU, memory, and temporary storage, to perform the following operations in a Snowflake session:”
      ↩︎ The compute layer: virtual warehouses
      “To perform these operations, a warehouse must be running and in use for the session. While a warehouse is running, it consumes Snowflake credits.”
      ↩︎ The compute layer: virtual warehouses
    4. 4.
      “Warehouses are required for queries, as well as all DML operations, including loading data into tables.”
      ↩︎ The compute layer: virtual warehouses

    Continue to page 2 of 2

    Snowflake Editions Compared: Standard, Enterprise, Business Critical and VPS

    Spotted a mistake, or was something unclear? Tell us.