CertSafari

    Free Snowflake SnowPro Advanced: MLOps Engineer (MLA-B01) Sample Questions

    35 free sample questions from our bank of 360+, covering every exam domain, with answers and detailed explanations. Updated October 2026.

    Domain 1: Operationalize Data Preparation and Feature Engineering

    Subdomain 1.1: Construct distributed feature engineering pipelines.

    1.An engineer loads an 800-million-row feature table with `session.table("FEATURES")`, applies several `filter` and `with_column` calls, then runs `df.to_pandas()` to feed a local scaler. The notebook kernel runs out of memory. What is the best redesign?

    1. A.Increase the notebook kernel memory and call `to_pandas()` in a loop over `limit` and `offset` slices to process rows in chunks
    2. B.Keep the work as Snowpark DataFrame operations and persist results with `write.save_as_table` so it runs on the warehouse
    3. C.Replace the Snowpark calls with a Python loop that issues a `session.sql` SELECT per customer id and appends each result locally
    4. D.Export the table to CSV files in an internal stage with `COPY INTO`, then download the files and concatenate them using pandas
    Show answer & explanation

    Correct answer: B — Keep the work as Snowpark DataFrame operations and persist results with `write.save_as_table` so it runs on the warehouse

    • A. Incorrect: slicing with limit and offset still pulls every row to the client, repeats the scan per slice, and gives unstable ordering without a sort.
    • B. Correct: Snowpark DataFrames are lazy and compile into SQL, so transformations and the final write run in the warehouse and nothing large reaches the client.
    • C. Incorrect: issuing one query per customer id creates huge numbers of round trips and still concentrates all rows in client memory.
    • D. Incorrect: downloading unloaded CSV files moves the entire dataset out of Snowflake and recreates the same memory limitation locally.

    Subdomain 1.1: Construct distributed feature engineering pipelines.

    2.A team keeps raw documents and images for feature extraction in a stage. Several data scientists in different roles need to read the same files, and access must be controlled with grants. Which stage choice supports this requirement?

    1. A.A temporary stage, which persists across sessions so that all roles can read the files after the creator disconnects
    2. B.A table stage, which is shared automatically with every role that holds SELECT on any table in the schema
    3. C.A user stage, which every account user can read through the `@~` shortcut once the files have been uploaded
    4. D.A named internal stage, which is a schema-level object that privileges can be granted on to multiple roles
    Show answer & explanation

    Correct answer: D — A named internal stage, which is a schema-level object that privileges can be granted on to multiple roles

    • A. Incorrect: a temporary stage is dropped when the session ends, so it does not persist for other roles.
    • B. Incorrect: a table stage is tied to its table and requires ownership of that table to use; SELECT privilege does not grant access to it.
    • C. Incorrect: a user stage belongs to one user and cannot be accessed by other users, nor can privileges be granted on it.
    • D. Correct: a named internal stage is a schema-level object, so READ and WRITE privileges can be granted to roles for controlled sharing.

    Subdomain 1.5: Operationalize features as first-class data assets.

    3.Which storage technology backs the Snowflake Feature Store online store used for low-latency feature lookups?

    1. A.Hybrid tables, which provide row-oriented storage with indexed point-lookup performance suited to serving individual entity keys.
    2. B.Standard permanent tables with clustering keys on the entity column, which Snowflake queries through result caching for sub-10 millisecond lookups.
    3. C.External tables on cloud object storage, refreshed by Snowpipe notifications so serving reads fetch the latest files directly from the bucket.
    4. D.Materialized views over the offline feature table, which Snowflake maintains in memory so serving applications read precomputed results immediately.
    Show answer & explanation

    Correct answer: A — Hybrid tables, which provide row-oriented storage with indexed point-lookup performance suited to serving individual entity keys.

    • A. The online feature store is implemented on hybrid tables, whose row-based storage supports fast single-key reads that columnar tables are not optimized for.
    • B. Standard tables are columnar and optimized for scans, and result caching does not guarantee low-latency key lookups for changing values.
    • C. External tables read files from object storage with scan latency far higher than needed for online serving and are not used for the online store.
    • D. Materialized views are stored on disk like tables and are not the online store mechanism. They do not offer indexed point-lookup behaviour.

    Subdomain 1.5: Operationalize features as first-class data assets.

    4.A team wants feature transformations packaged so they can be versioned, tested and deployed independently of any model training code. Which practices support this? Select all that apply.(Select 3)

    1. A.Store Snowpark, SQL and dynamic table definitions in a dedicated Git repository with its own CI pipeline and release tags.
    2. B.Register each released transformation as a numbered feature view version, so models reference stable, named dependencies.
    3. C.Expose shared logic through reusable Snowpark functions packaged in a module imported by the feature repository only.
    4. D.Embed the feature code inside each model's training notebook so the notebook is the single place where all dependencies are visible.
    5. E.Export transformed features to CSV files in a shared drive and have each model load the latest file when it trains.
    Show answer & explanation

    Correct answers: A, B, C — Store Snowpark, SQL and dynamic table definitions in a dedicated Git repository with its own CI pipeline and release tags.; Register each released transformation as a numbered feature view version, so models reference stable, named dependencies.; Expose shared logic through reusable Snowpark functions packaged in a module imported by the feature repository only.

    • A. A separate repo gives features their own lifecycle, review and testing, so a model release does not drag feature changes along.
    • B. Versioned registration turns a transformation into a stable contract that models consume by name and version.
    • C. Packaging shared logic in a feature-owned module avoids copies in model notebooks and lets the logic be unit-tested independently.
    • D. Embedding logic in notebooks couples features to a model and duplicates it across consumers, defeating independent versioning.
    • E. CSV hand-offs lose lineage, versioning and access control, and they cannot provide online serving.

    Subdomain 1.3: Ensure temporal integrity and feature consistency.

    5.Which mechanism does the Snowflake Feature Store use when `generate_training_set` receives a `spine_timestamp_col` and feature views that define a `timestamp_col`?

    1. A.An equality join between the spine timestamp and each feature view timestamp, which keeps only rows that share an identical value
    2. B.A Time Travel `AT(TIMESTAMP => ...)` query against every feature view table, replaying the table as it stood at each spine row's time
    3. C.A windowed `QUALIFY` over a cross join of the spine and feature views, which ranks every historical row for every spine row
    4. D.An `ASOF JOIN` that matches each spine row to the latest feature row for the same entity key at or before the spine timestamp
    Show answer & explanation

    Correct answer: D — An `ASOF JOIN` that matches each spine row to the latest feature row for the same entity key at or before the spine timestamp

    • A. Incorrect. An equality join would drop most spine rows, because label times rarely match feature timestamps exactly. Point-in-time retrieval needs a nearest-prior match instead.
    • B. Incorrect. Time Travel is limited by retention and operates on one timestamp per query, not per spine row. The Feature Store does not rely on it for point-in-time lookups.
    • C. Incorrect. A cross join would be prohibitively expensive at training-set scale. Snowflake provides a purpose-built temporal join that avoids materializing that product.
    • D. Correct. The point-in-time lookup is implemented with an ASOF JOIN whose match condition keeps the closest feature row not later than the spine time, per entity key.

    Subdomain 1.3: Ensure temporal integrity and feature consistency.

    6.A feature view named `account_profile` is built from a `dim_account` table that is overwritten nightly with current values. The team backfills two years of labels and joins them with `generate_training_set`, and the model degrades sharply in production. What is the underlying problem and fix?

    1. A.The ASOF join rounds timestamps to the nearest day, so the team should switch to a sub-second `timestamp_col` data type with more precision
    2. B.Past labels get present-day attribute values, so build the feature view from a history table with effective-from `timestamp_col`
    3. C.The managed feature view refreshed too rarely, so lowering `refresh_freq` to one minute would restore the older historical values
    4. D.The feature view should be switched to an external feature view so Snowflake stops overwriting the stored values during refresh
    Show answer & explanation

    Correct answer: B — Past labels get present-day attribute values, so build the feature view from a history table with effective-from `timestamp_col`

    • A. Incorrect. Timestamp precision is not the issue here. The table holds only current state, so there is no historical row to match, whatever the precision.
    • B. Correct. A table that is overwritten has no history, so point-in-time lookups can only return current values. A history or SCD2-style source with an effective timestamp gives each label the values that applied then.
    • C. Incorrect. Refreshing more often captures future changes only. It cannot recreate attribute values from two years ago that were already overwritten.
    • D. Incorrect. Whether a view is managed or external does not retain history. History has to exist in the source data that the feature view reads.

    Subdomain 1.4: Configure automated ingestion and data quality.

    7.A team attaches data quality checks to a raw events table and needs two things in place before results start appearing. Which TWO actions are required?(Select 2)

    1. A.Run `ALTER TABLE ... ADD DATA METRIC FUNCTION SNOWFLAKE.CORE.NULL_COUNT ON (customer_id)` to associate the metric with the column.
    2. B.Run `ALTER TABLE ... SET DATA_METRIC_SCHEDULE = 'TRIGGER_ON_CHANGES'` (or another interval) so the table has a defined evaluation schedule.
    3. C.Run `ALTER SCHEMA ... ADD DATA METRIC FUNCTION` once, so every table in the schema inherits the NULL_COUNT check automatically on creation.
    4. D.Create a dedicated warehouse and set it as `DATA_METRIC_WAREHOUSE` on the table so scheduled metric runs always have compute available.
    5. E.Create a stream on the events table and a task that calls the DMF in a `SELECT` after each batch to persist the readings in a log table.
    Show answer & explanation

    Correct answers: A, B — Run `ALTER TABLE ... ADD DATA METRIC FUNCTION SNOWFLAKE.CORE.NULL_COUNT ON (customer_id)` to associate the metric with the column.; Run `ALTER TABLE ... SET DATA_METRIC_SCHEDULE = 'TRIGGER_ON_CHANGES'` (or another interval) so the table has a defined evaluation schedule.

    • A. Correct: a DMF only runs once it has been associated with the table column through an ADD DATA METRIC FUNCTION clause.
    • B. Correct: a table needs its DMF schedule defined, and the schedule applies to every metric function attached to that table.
    • C. Incorrect: DMFs are associated with individual tables, views, or similar objects and are not inherited from a schema.
    • D. Incorrect: scheduled DMFs run on serverless compute and there is no table-level warehouse setting that they use.
    • E. Incorrect: scheduled DMFs already log their results, so a hand-built stream and task duplicates the built-in mechanism.

    Subdomain 1.4: Configure automated ingestion and data quality.

    8.A team stores online features in hybrid tables and also consumes a shared table from a data provider. They want to attach data metric functions to both. Which TWO statements are accurate?(Select 2)

    1. A.DMFs cannot be attached to views, so any validation of a view-based external feature view needs to be rebuilt as a materialized table.
    2. B.DMFs cannot be set on a table received through a share, so the consumer should create a local table or view and attach the metric there.
    3. C.DMFs can be attached to streams, so change records can be validated before the consuming task merges them into the feature table.
    4. D.DMFs cannot be set on hybrid tables, so validation must run on the offline source or feature view that populates the online store.
    5. E.DMFs cannot be attached to dynamic tables, so features produced by a managed feature view can only be checked with a custom stored procedure.
    Show answer & explanation

    Correct answers: B, D — DMFs cannot be set on a table received through a share, so the consumer should create a local table or view and attach the metric there.; DMFs cannot be set on hybrid tables, so validation must run on the offline source or feature view that populates the online store.

    • A. Incorrect: views are supported DMF targets, so no materialization workaround is needed.
    • B. Correct: DMFs cannot be set on shared objects, so monitoring must be attached to an object owned by the consuming account.
    • C. Incorrect: streams are not supported for DMF association, so change records must be checked after landing in a table.
    • D. Correct: hybrid tables are among the unsupported object types for DMF association, so checks belong on the supported upstream objects.
    • E. Incorrect: dynamic tables are a supported target for DMFs, so managed feature view outputs can use them directly.

    Subdomain 1.4: Configure automated ingestion and data quality.

    9.A pricing model needs validation that no row has `discount_amount` greater than `list_price`, a rule involving two columns that the built-in system DMFs do not cover. What should the engineer do?

    1. A.Add a `CHECK (discount_amount <= list_price)` constraint to the table so that Snowflake rejects violating inserts during ingestion.
    2. B.Add a `CHECK (discount_amount <= list_price)` constraint to the table so that Snowflake rejects any violating inserts at the point of ingestion.
    3. C.Create a custom DMF taking a `TABLE(list_price NUMBER, discount_amount NUMBER)` argument that returns violating rows, then attach it to both.
    4. D.Attach the system `SNOWFLAKE.CORE.NULL_COUNT` DMF to both columns and treat a non-zero result as evidence of an inconsistent discount and price pair.
    Show answer & explanation

    Correct answer: C — Create a custom DMF taking a `TABLE(list_price NUMBER, discount_amount NUMBER)` argument that returns violating rows, then attach it to both.

    • A. Incorrect: Snowflake does not enforce CHECK constraints in this way, so violating rows would still load.
    • B. Incorrect: duplicate counts detect repeated values, which is unrelated to comparing one column's magnitude with another's.
    • C. Correct: custom DMFs can accept multiple columns and return a count, which suits cross-column business rules.
    • D. Incorrect: NULL_COUNT measures missing values only and says nothing about the relationship between the two numeric columns.

    Subdomain 1.2: Implement Snowflake Feature Store architecture and management.

    10.A retail team has a customer-level purchase aggregation in a Snowpark DataFrame keyed on `CUSTOMER_ID`. Several feature views will be built over different source tables, and all must be retrievable for the same business object. What is the correct way to model this in the Feature Store?

    1. A.Register one `Entity` named `CUSTOMER` with join_keys `["CUSTOMER_ID"]`, then pass it as `entities` on every feature view whose DataFrame has that column.
    2. B.Create a separate `Entity` per feature view with a unique name, since each feature view owns its own join keys and cannot share an entity with other views in the store.
    3. C.Skip entity registration and tag each feature view's source table with a `CUSTOMER_ID` tag, because the Feature Store infers join keys from column tags at lookup time.
    4. D.Define the `Entity` with join_keys taken from the label column of the training spine, so lookups at inference time use the target variable as the key for retrieval.
    Show answer & explanation

    Correct answer: A — Register one `Entity` named `CUSTOMER` with join_keys `["CUSTOMER_ID"]`, then pass it as `entities` on every feature view whose DataFrame has that column.

    • A. Correct. An Entity is a registered, tag-backed definition whose join keys must exist as columns in each feature view's DataFrame, which lets views and spines join consistently.
    • B. Incorrect. Entities are designed to be shared; many feature views can reference one entity, and duplicating them defeats centralized reuse and consistent joins.
    • C. Incorrect. Join keys are declared explicitly on the Entity object; the store does not infer them from tags you place on source tables.
    • D. Incorrect. Join keys identify the subject of the features, such as a customer, never the prediction label; using a label as a key would leak the target.

    Subdomain 1.2: Implement Snowflake Feature Store architecture and management.

    11.A managed feature view `USER_ACTIVITY` currently refreshes every 10 minutes on warehouse `ML_WH_L`. Finance wants cheaper compute and a 30-minute refresh, with no change to the feature logic or version. What is the appropriate action?

    1. A.Register a new version of `USER_ACTIVITY` with the same DataFrame, a refresh_freq of `30 minutes` and a smaller warehouse, then retire the original version.
    2. B.Drop the backing dynamic table and recreate it with `CREATE DYNAMIC TABLE` using a smaller warehouse and a `TARGET_LAG` of 30 minutes outside the Feature Store API.
    3. C.Call `update_feature_view` on the existing version, supplying `refresh_freq="30 minutes"` and a smaller `warehouse` while leaving the logic untouched.
    4. D.Suspend the feature view's dynamic table and run a scheduled task that calls `REFRESH` every 30 minutes on the larger warehouse to approximate the new cadence.
    Show answer & explanation

    Correct answer: C — Call `update_feature_view` on the existing version, supplying `refresh_freq="30 minutes"` and a smaller `warehouse` while leaving the logic untouched.

    • A. Incorrect. Nothing about the logic changes, so a new version is unnecessary overhead and forces every consumer to repoint; refresh settings can be changed in place.
    • B. Incorrect. Recreating the table manually leaves the Feature Store metadata out of sync and risks losing the registered feature view definition and tags.
    • C. Correct. Refresh frequency, warehouse and description are the mutable properties of a registered feature view, so cost tuning does not require a new version.
    • D. Incorrect. Manual suspension and tasks bypass the managed refresh, keep the expensive warehouse in use, and leave the registered refresh_freq unchanged.

    Domain 2: MLOps Infrastructure and Management

    Subdomain 2.1: Manage infrastructure for ML.

    12.A data scientist develops a training function in a local IDE and wants it to run on a GPU compute pool without rewriting it into a notebook. Which ML Jobs approach requires the least code change?

    1. A.Convert the function into a Snowflake Notebook cell and schedule the notebook with a task running on a warehouse
    2. B.Wrap the function in a stored procedure and call it on a Snowpark-optimized warehouse sized to the largest tier
    3. C.Package the function as a Streamlit app and deploy it to the compute pool as a long-running service
    4. D.Decorate the function with @remote, passing the compute pool name and a stage name for payload storage
    Show answer & explanation

    Correct answer: D — Decorate the function with @remote, passing the compute pool name and a stage name for payload storage

    • A. Incorrect: a task running a notebook on a warehouse stays on warehouse compute and requires restructuring the code into notebook cells.
    • B. Incorrect: stored procedures run on warehouses, not GPU compute pools, so they cannot use the GPU compute pool this scenario requires.
    • C. Incorrect: a Streamlit app is a UI service and not a mechanism for dispatching a training function as a finite job.
    • D. Correct: the ML Jobs @remote decorator ships a Python function to a compute pool, using a stage to hold the payload, and lets the function return its value to the caller.

    Subdomain 2.1: Manage infrastructure for ML.

    13.A team wants a retraining job to behave identically for months even as Snowflake publishes newer Container Runtime releases. Which setting supports this?

    1. A.Pin the job to a specific runtime version with the runtime_environment parameter instead of the latest default
    2. B.Keep the code in an internal stage, since Snowflake versions the stage and runtime together as a unit
    3. C.Set AUTO_SUSPEND_SECS to zero so that the pool image never refreshes between runs of the same job
    4. D.Use a larger instance family, because newer runtime releases are only deployed to smaller families
    Show answer & explanation

    Correct answer: A — Pin the job to a specific runtime version with the runtime_environment parameter instead of the latest default

    • A. Correct: Container Runtime is versioned, and runtime_environment lets a job choose a specific version rather than the latest default, giving repeatable behavior.
    • B. Incorrect: staging code does not pin the runtime; the two are versioned separately.
    • C. Incorrect: pool suspension does not control which runtime version a job uses.
    • D. Incorrect: runtime releases are not tied to instance family size.

    Subdomain 2.2: Utilize Snowflake Workspaces.

    14.A group already uses `snowflake-ml-python` modeling classes in a warehouse runtime and now wants to use an open-source package not supported there. What justifies moving to Container Runtime notebooks?

    1. A.Container Runtime bills nothing for compute while notebooks are idle, making it free for open-source experimentation.
    2. B.Container Runtime removes the need for any data to reside in Snowflake, so external storage becomes mandatory.
    3. C.Warehouse runtime cannot execute Python, so all Python workloads must move to Container Runtime regardless.
    4. D.Container Runtime lets them pip install open-source packages via an EAI and run them on CPU or GPU compute pools.
    Show answer & explanation

    Correct answer: D — Container Runtime lets them pip install open-source packages via an EAI and run them on CPU or GPU compute pools.

    • A. Incorrect. Compute pool nodes are billed while running.
    • B. Incorrect. Data stays in Snowflake and is not forced into external storage.
    • C. Incorrect. Warehouses run Python through Snowpark and stored procedures.
    • D. Correct. Package flexibility and GPU access are main reasons to choose Container Runtime.

    Subdomain 2.2: Utilize Snowflake Workspaces.

    15.Which of the following are supported ways to bring additional Python packages into a Notebook in Workspaces? (Select all that apply)(Select 3)

    1. A.Uploading a wheel file to the workspace or a stage and running `!pip install file_name.whl`
    2. B.Using an external access integration to pip install from public endpoints outside Snowflake
    3. C.Installing from the Snowflake-managed PyPI repository or a customer-managed artifact repository
    4. D.Editing the host operating system image on the compute node through SSH access to the pool
    5. E.Mounting a local laptop folder of packages through the Snowsight browser into the notebook
    Show answer & explanation

    Correct answers: A, B, C — Uploading a wheel file to the workspace or a stage and running `!pip install file_name.whl`; Using an external access integration to pip install from public endpoints outside Snowflake; Installing from the Snowflake-managed PyPI repository or a customer-managed artifact repository

    • A. Correct. Workspace or stage files such as wheels are a supported installation path.
    • B. Correct. EAIs enable access to a defined list of external endpoints.
    • C. Correct. Artifact repositories are a documented source.
    • D. Incorrect. Compute pool nodes are not accessible via SSH.
    • E. Incorrect. Local laptop folders are not mountable in the service.

    Subdomain 2.3: Track experiments and metadata.

    16.A team lead audits a pipeline for reproducibility. Which TWO properties of Snowflake Datasets support experiment reproducibility?(Select 2)

    1. A.Passing the Dataset's DataFrame to `log_model` links the dataset to the logged model for lineage
    2. B.Each version is an immutable point-in-time snapshot of the materialized DataFrame, stored as Parquet files
    3. C.Version names are generated by Snowflake as hashes and cannot be chosen by the user at creation time
    4. D.Dataset versions are stored as a single CSV file so the file checksum guarantees ordering of the rows
    5. E.Versions are refreshed automatically whenever the upstream table changes so models always see current rows
    Show answer & explanation

    Correct answers: A, B — Passing the Dataset's DataFrame to `log_model` links the dataset to the logged model for lineage; Each version is an immutable point-in-time snapshot of the materialized DataFrame, stored as Parquet files

    • A. Correct: `log_model` with Dataset-derived input connects source data, dataset, and model for lineage.
    • B. Correct: immutability means the same version always returns the same rows.
    • C. Incorrect: the version name is supplied to `create_from_dataframe`, up to 128 characters.
    • D. Incorrect: data is stored as evenly sized Parquet files, not a single CSV.
    • E. Incorrect: versions do not refresh; a new version must be created explicitly.

    Subdomain 2.3: Track experiments and metadata.

    17.A team lead is defining who can do what with Dataset objects used as training inputs and lineage anchors. Which TWO statements are accurate?(Select 2)

    1. A.A role needs only `USAGE` on a dataset to add new versions, because versions are immutable and cannot affect other readers
    2. B.A role needs `CREATE EXPERIMENT` on the schema to create a dataset, because datasets are stored as experiment artifacts
    3. C.A role needs `READ` on the underlying source table forever, because every Dataset read re-executes the original query
    4. D.A role needs `USAGE` or `OWNERSHIP` on a dataset to read its versions, but adding or deleting versions needs `OWNERSHIP`
    5. E.A role needs `CREATE DATASET` on the schema to create a dataset, such as with `create_from_dataframe`
    Show answer & explanation

    Correct answers: D, E — A role needs `USAGE` or `OWNERSHIP` on a dataset to read its versions, but adding or deleting versions needs `OWNERSHIP`; A role needs `CREATE DATASET` on the schema to create a dataset, such as with `create_from_dataframe`

    • A. Incorrect: adding versions modifies the dataset and needs OWNERSHIP.
    • B. Incorrect: datasets are their own schema-level objects with `CREATE DATASET`, not experiment artifacts.
    • C. Incorrect: a version is a materialized snapshot stored as Parquet, so reads do not re-run the source query.
    • D. Correct: reading requires USAGE or OWNERSHIP, while modifying versions requires OWNERSHIP.
    • E. Correct: creating a dataset requires `CREATE DATASET` on the schema.

    Domain 3: Model Serving and Deployment Operations

    Subdomain 3.1: Operate the Snowflake Model Registry.

    18.A data scientist logs a `CustomModel` whose `__init__` loads a 400 MB pickled preprocessor into `self.prep`, and also passes the same pickle path through `ModelContext`. The logged version is roughly double the expected size. What is the most likely cause?

    1. A.The preprocessor was captured directly on the instance instead of being read through the model context, so two copies are stored.
    2. B.Setting `relax_version` to its default of `True` makes the registry package every dependency twice, once pinned and once as a range.
    3. C.The `code_paths` argument was omitted, so the registry falls back to embedding the whole working directory alongside the model artifacts.
    4. D.Pickle files bundled through `ModelContext` are re-serialized with the Snowpark ML library by default, which doubles their stored footprint.
    Show answer & explanation

    Correct answer: A — The preprocessor was captured directly on the instance instead of being read through the model context, so two copies are stored.

    • A. Snowflake guidance is to always access model objects via the context. Capturing the object outside the context creates a duplicate serialized copy that inflates the artifact size.
    • B. `relax_version` only loosens version specifiers such as `==x.y.z` into ranges. It adds no packaged copies of artifacts, so it cannot explain a doubled model size.
    • C. Omitting `code_paths` just means no helper modules are bundled. The registry does not substitute the working directory, so this would not duplicate the pickle.
    • D. Files passed as paths are bundled as they are. Doubling comes from capturing the object separately, and `embed_local_ml_library` is a distinct option that defaults to `False`.

    Subdomain 3.1: Operate the Snowflake Model Registry.

    19.A model logged with `conda_dependencies=['conda-forge::xgboost==2.0.3']` and `target_platforms=['WAREHOUSE']` fails dependency validation when promoted for warehouse inference in production. What is the right fix?

    1. A.Set `embed_local_ml_library` to `True` so the conda-forge build is shipped inside the model, then log the same version again.
    2. B.Change `target_platforms` to `['SNOWPARK_CONTAINER_SERVICES']` only, then keep serving the model from warehouses through `m!PREDICT`.
    3. C.Pin `relax_version` to `False` so the exact conda-forge build is requested and the validation step stops looking for alternatives.
    4. D.Replace the conda-forge prefix with a version from the Snowflake conda channel, then log a new version targeting the warehouse.
    Show answer & explanation

    Correct answer: D — Replace the conda-forge prefix with a version from the Snowflake conda channel, then log a new version targeting the warehouse.

    • A. This option embeds a copy of the Snowpark ML library, not third-party packages. The unresolved conda-forge dependency would still fail validation for the warehouse.
    • B. Switching platforms would not let warehouse inference work, and the requirement is warehouse serving. The model must still be loadable where it is actually called.
    • C. `relax_version` governs how pinned versions are loosened and does not change which channel is searched. A package from conda-forge remains unavailable to warehouse execution.
    • D. Warehouse execution resolves packages from the Snowflake conda channel only. Choosing an available xgboost version from that channel lets the dependency resolve at registration and deployment.

    Subdomain 3.2: Implement inference deployment patterns.

    20.A data scientist runs `create_service` for a model with a public endpoint and the call fails with a privilege error. Which privileges must the role hold for this deployment? (Select all that apply)(Select 3)

    1. A.BIND SERVICE ENDPOINT, needed to create a service with a public endpoint.
    2. B.OWNERSHIP or READ privilege on the model being deployed from the registry.
    3. C.USAGE or OWNERSHIP on the compute pool that will host the model service.
    4. D.CREATE EXTERNAL FUNCTION on the schema where the registered model resides.
    5. E.MONITOR on the warehouse that was active when the model version was logged.
    Show answer & explanation

    Correct answers: A, B, C — BIND SERVICE ENDPOINT, needed to create a service with a public endpoint.; OWNERSHIP or READ privilege on the model being deployed from the registry.; USAGE or OWNERSHIP on the compute pool that will host the model service.

    • A. Correct: BIND SERVICE ENDPOINT is the account-level privilege needed for services that expose public endpoints.
    • B. Correct: the role must own or have READ on the model to deploy it as a service.
    • C. Correct: the role needs USAGE or OWNERSHIP on the compute pool to place the service on it.
    • D. Incorrect: the model service uses SPCS endpoints, not external functions, so this schema privilege is not needed.
    • E. Incorrect: the logging warehouse is not part of service creation, so no privilege on it is required.

    Subdomain 3.2: Implement inference deployment patterns.

    21.An SRE wants to observe request latency, throughput, and error rates for a model served through a real-time inference endpoint. Which statements are accurate? (Select all that apply)(Select 3)

    1. A.Platform metrics (CPU, memory, GPU use) and inference metrics are collected by the service controller.
    2. B.Latency, throughput, and error-rate metrics flow through an OpenTelemetry-powered collection pipeline.
    3. C.Metrics are made available through the Snowflake event table and drive the service's metrics dashboards.
    4. D.Metrics cover only service function calls, and direct HTTPS requests to the endpoint are not measured at all.
    5. E.Credit consumption is exposed only inside query history for the warehouse that called the endpoint service.
    Show answer & explanation

    Correct answers: A, B, C — Platform metrics (CPU, memory, GPU use) and inference metrics are collected by the service controller.; Latency, throughput, and error-rate metrics flow through an OpenTelemetry-powered collection pipeline.; Metrics are made available through the Snowflake event table and drive the service's metrics dashboards.

    • A. Correct: the controller collects platform and performance metrics for inference requests.
    • B. Correct: inference-specific metrics such as latency, throughput, and errors use an OpenTelemetry pipeline.
    • C. Correct: metrics reach the event table and the metrics dashboards for the inference service.
    • D. Incorrect: metrics cover REST endpoint requests; they are not limited to service function traffic.
    • E. Incorrect: serving credit use is reported through the MODEL_SERVING_USAGE_HISTORY view, not warehouse query history.

    Subdomain 3.3: Execute platform migrations.

    22.A data science group has MLflow pyfunc models in a Databricks workspace and must serve them from Snowflake. Which action is the most direct migration step?

    1. A.Load the pyfunc model in a Python environment and register it with `Registry.log_model`, supplying a signature or `sample_input_data` and pinned package versions.
    2. B.Stand up an MLflow tracking server in Snowpark Container Services and point Databricks at it, because the registry can only host models that an MLflow server proxies.
    3. C.Copy the Databricks metastore tables holding model metadata into a Snowflake schema; the registry then discovers the models automatically from those tables.
    4. D.Retrain every model from scratch with Snowflake ML estimators, since MLflow-flavored models cannot be imported into the registry under any circumstance.
    Show answer & explanation

    Correct answer: A — Load the pyfunc model in a Python environment and register it with `Registry.log_model`, supplying a signature or `sample_input_data` and pinned package versions.

    • A. Correct: the Model Registry supports MLflow models, so loading the artifact and calling log_model with a signature and pinned dependencies creates a native model version.
    • B. Incorrect: no proxy MLflow server is required; the registry stores the model itself rather than forwarding calls to an external tracking service.
    • C. Incorrect: copying metadata tables does not move the model artifact and the registry does not auto-discover models from user tables.
    • D. Incorrect: retraining is unnecessary because MLflow models can be logged to the registry directly.

    Subdomain 3.3: Execute platform migrations.

    23.A team used the open-source SHAP library on an external platform to explain predictions. Which Snowflake-native capability should they look for on registered models?

    1. A.Model explainability that computes Shapley values for registered versions, called as a model version method.
    2. B.A tag on the model version named EXPLAIN, which makes the registry attach feature importances to every response.
    3. C.The query profile of the prediction query, which lists each feature's contribution to the output value.
    4. D.Search optimization on the input table, which ranks features by how often each column is read.
    Show answer & explanation

    Correct answer: A — Model explainability that computes Shapley values for registered versions, called as a model version method.

    • A. Correct: the registry offers built-in Shapley-value explainability for logged models.
    • B. Incorrect: tags are metadata and do not compute explanations.
    • C. Incorrect: the query profile describes execution cost, not feature contributions.
    • D. Incorrect: search optimization is a performance feature unrelated to model explanations.

    Domain 4: Pipeline Orchestration and Automation (CI/CD)

    Subdomain 4.1: Orchestrate end-to-end ML workflows.

    24.Which statements about a finalizer task in a Snowflake task graph are true? (Select two.)(Select 2)

    1. A.Each root task can have only one finalizer task, and that finalizer task cannot itself have any child tasks below it in the graph.
    2. B.A finalizer still starts when the root task was skipped for the run, which lets it report on runs that never began executing at all.
    3. C.A finalizer runs only when every other task in the graph succeeded, which makes it the natural place for the model promotion step.
    4. D.A finalizer can carry its own SCHEDULE parameter, so it can run hourly on its own independent of the graph's root task schedule.
    5. E.A finalizer runs after all other tasks in the graph run have finished, succeeded or failed, so it suits cleanup such as dropping temporary tables.
    Show answer & explanation

    Correct answers: A, E — Each root task can have only one finalizer task, and that finalizer task cannot itself have any child tasks below it in the graph.; A finalizer runs after all other tasks in the graph run have finished, succeeded or failed, so it suits cleanup such as dropping temporary tables.

    • A. Correct. Both restrictions are part of the finalizer definition for a task graph.
    • B. Incorrect. If the root task is skipped, the finalizer does not start for that run.
    • C. Incorrect. A finalizer also runs after failures, so promotion placed there could publish a model from a broken run.
    • D. Incorrect. A finalizer is attached to the root task and runs as part of each graph run, not on a separate schedule.
    • E. Correct. This is the purpose of a finalizer: guaranteed post-run work regardless of the outcome of the other tasks.

    Subdomain 4.1: Orchestrate end-to-end ML workflows.

    25.A team that scripts deployments with SnowSQL is choosing a command-line tool for new ML pipeline automation. Which statement best describes how Snowflake CLI relates to SnowSQL?

    1. A.Snowflake CLI runs only interactively in a terminal session, whereas SnowSQL supports file execution suitable for unattended pipeline jobs.
    2. B.Snowflake CLI is a thin wrapper that forwards every command to SnowSQL, so SnowSQL must be installed on every CI runner.
    3. C.Snowflake CLI is the newer tool and adds command groups for stages, Git repositories, and Snowpark deployments beyond a SQL client.
    4. D.SnowSQL is the only client able to read config.toml, so Snowflake CLI projects must still depend on it for connection handling.
    Show answer & explanation

    Correct answer: C — Snowflake CLI is the newer tool and adds command groups for stages, Git repositories, and Snowpark deployments beyond a SQL client.

    • A. Incorrect. Snowflake CLI is designed for scripted use and supports running SQL files non-interactively.
    • B. Incorrect. Snowflake CLI is an independent tool and does not delegate to SnowSQL.
    • C. Correct. The CLI covers object and project workflows that SnowSQL, a SQL client, does not provide.
    • D. Incorrect. Config.toml is the Snowflake CLI configuration file; SnowSQL uses its own configuration.

    Subdomain 4.2: Configure CI/CD and version control.

    26.One deployment script must create `ML_DEV` objects on a dev account and `ML_PROD` objects on production with a larger warehouse, without copying the file. What is the best approach?

    1. A.Maintain long-lived `dev` and `prod` branches whose script copies hardcode their own database and warehouse names
    2. B.Set `ALTER SESSION` parameters for the environment name before the run and let Snowflake substitute them into object identifiers
    3. C.Reference Jinja variables such as `{{env}}` in the script and run `EXECUTE IMMEDIATE FROM ... USING (env => 'PROD')`
    4. D.Wrap the whole script in a stored procedure that builds each statement by string concatenation and runs it with dynamic SQL
    Show answer & explanation

    Correct answer: C — Reference Jinja variables such as `{{env}}` in the script and run `EXECUTE IMMEDIATE FROM ... USING (env => 'PROD')`

    • A. Incorrect. Divergent branches drift apart and defeat promoting the identical reviewed code; parameters belong in the invocation, not in forked files.
    • B. Incorrect. Session parameters are not substituted into script text; templating variables passed with `USING` are the mechanism.
    • C. Correct. `EXECUTE IMMEDIATE FROM` renders Jinja templating, so one versioned file serves every environment with values passed through `USING`.
    • D. Incorrect. This hides the DDL from review and testing, and it is unnecessary because templating is built into `EXECUTE IMMEDIATE FROM`.

    Subdomain 4.2: Configure CI/CD and version control.

    27.A data scientist's commit from a Snowflake workspace fails with an authorization error, though fetches from the same private repository work. What should be checked?

    1. A.Whether the API integration lists the repository branch names in its allowed prefixes parameter list
    2. B.Whether the stored token has write scope on the remote and the role has `WRITE` on the repository
    3. C.Whether the repository object was cloned with a shallow depth that disables pushes for older commits
    4. D.Whether the warehouse used by the workspace is large enough to complete the push within the timeout
    Show answer & explanation

    Correct answer: B — Whether the stored token has write scope on the remote and the role has `WRITE` on the repository

    • A. Incorrect. Allowed prefixes constrain URLs, not branches, and fetches already succeed through the integration.
    • B. Correct. Fetching requires only read access, while pushing needs write-capable credentials and the WRITE privilege.
    • C. Incorrect. Depth is not a setting of the repository object and has no bearing on authorization.
    • D. Incorrect. Pushing is not governed by warehouse size, and the error is about authorization.

    Subdomain 4.3: Implement retraining and troubleshooting.

    28.New labeled outcome rows are loaded into `labels_raw` by a pipeline at irregular times. A data scientist wants the retraining task graph to start only when new rows exist, and does not want to define a polling schedule on the root task. Which configuration is correct?

    1. A.Create a stream on `labels_raw`, then create the root task with WHEN SYSTEM$STREAM_HAS_DATA on that stream and omit SCHEDULE so it runs as a triggered task.
    2. B.Create the root task with a SCHEDULE of one minute and an AFTER clause that names the loading pipeline's task, so it fires each time that upstream task finishes loading rows.
    3. C.Create a stream on `labels_raw` and set the task's TRIGGER_ON_CHANGE parameter to TRUE, which makes the task execute once for every committed DML statement on the table.
    4. D.Create an alert with a one-minute SCHEDULE that counts rows in `labels_raw`, and set its action to ALTER TASK RESUME so the suspended root task starts running after each new load.
    Show answer & explanation

    Correct answer: A — Create a stream on `labels_raw`, then create the root task with WHEN SYSTEM$STREAM_HAS_DATA on that stream and omit SCHEDULE so it runs as a triggered task.

    • A. A triggered task has no SCHEDULE and uses a WHEN condition with SYSTEM$STREAM_HAS_DATA, so the graph starts only when the stream holds unconsumed changes.
    • B. Defining a one-minute SCHEDULE makes the root task a polling scheduled task, which is exactly what the requirement rules out, and AFTER is only valid for child tasks rather than the root.
    • C. There is no TRIGGER_ON_CHANGE task parameter; event-driven task execution is expressed through a stream condition in the WHEN clause.
    • D. An alert is a scheduled poller, and resuming a task does not make it execute a run, so this keeps a polling schedule and does not start the graph on new data.

    Subdomain 4.3: Implement retraining and troubleshooting.

    29.A data scientist submits a distributed training ML Job and wants it to run on four nodes. The job code calls `scale_cluster()` partway through, and sometimes the job stalls waiting for nodes. According to Snowflake's guidance, how should the cluster size be specified?

    1. A.Pass `target_instances` at job submission so the cluster size is fixed up front, instead of resizing with `scale_cluster()` inside the running job.
    2. B.Call `scale_cluster()` before every training epoch, because the documented approach is to grow the cluster incrementally as each epoch starts.
    3. C.Set MIN_NODES to four on the compute pool and omit any job-level setting, because a pool's MIN_NODES value is read as the instance count by the job.
    4. D.Increase `num_workers` in the dataloader so that the job requests one node per worker process automatically at runtime.
    Show answer & explanation

    Correct answer: A — Pass `target_instances` at job submission so the cluster size is fixed up front, instead of resizing with `scale_cluster()` inside the running job.

    • A. Snowflake recommends specifying the desired cluster size with `target_instances` when submitting the job, which avoids mid-job scaling stalls.
    • B. Resizing during execution is the approach that is discouraged, because it can leave the job waiting for nodes to be provisioned.
    • C. MIN_NODES keeps nodes warm in the pool but does not tell an individual job how many instances to use.
    • D. Dataloader workers are processes within a node and have no effect on how many compute pool nodes the job requests.

    Domain 5: Governance, Security, and Monitoring

    Subdomain 5.1: Enforce Snowflake security policies.

    30.A model-deployment role must be able to log new models into the schema ML.REGISTRY, while a separate application role must only call inference on an existing model. Which two grants implement this? (Choose two.)(Select 2)

    1. A.GRANT CREATE MODEL ON SCHEMA ml.registry to the deployment role, so that it can log new models and versions in that schema.
    2. B.GRANT OWNERSHIP ON MODEL fraud_model to the application role, so that it can invoke the model's methods for inference only.
    3. C.GRANT USAGE ON MODEL ml.registry.fraud_model to the application role, so that it can call model methods without modifying it.
    4. D.GRANT CREATE COMPUTE POOL ON ACCOUNT to the application role, so that its inference calls are allowed to run on a pool.
    5. E.GRANT MODIFY ON SCHEMA ml.registry to the application role, so that it can read models registered by other roles in it.
    Show answer & explanation

    Correct answers: A, C — GRANT CREATE MODEL ON SCHEMA ml.registry to the deployment role, so that it can log new models and versions in that schema.; GRANT USAGE ON MODEL ml.registry.fraud_model to the application role, so that it can call model methods without modifying it.

    • A. Correct. CREATE MODEL on the schema is the privilege required to log models and versions into the registry schema.
    • B. Incorrect. OWNERSHIP lets the role alter, drop, and re-grant the model, which is far beyond inference-only access.
    • C. Correct. USAGE on the model lets a role call its methods and see it without changing it.
    • D. Incorrect. CREATE COMPUTE POOL is an account-level privilege for creating pools and is not needed to call model methods.
    • E. Incorrect. MODIFY is not the privilege that enables reading models, and it would be broader than inference-only access.

    Subdomain 5.1: Enforce Snowflake security policies.

    31.Which statements about row access policies are correct? (Choose three.)(Select 3)

    1. A.A single row access policy can be attached to many tables and views so that one definition is maintained centrally.
    2. B.A table can have several row access policies, each applying to a different role and combined with logical AND.
    3. C.When an object has both policy types, the row access policy is evaluated before any masking policies are applied.
    4. D.Row access policies can be placed on dynamic tables, which lets feature view data be filtered by role as well.
    5. E.Future grants can be used to give roles the APPLY privilege on row access policies created later in a schema.
    6. F.A row access policy can be attached to a tag so that every table carrying the tag is filtered automatically.
    Show answer & explanation

    Correct answers: A, C, D — A single row access policy can be attached to many tables and views so that one definition is maintained centrally.; When an object has both policy types, the row access policy is evaluated before any masking policies are applied.; Row access policies can be placed on dynamic tables, which lets feature view data be filtered by role as well.

    • A. Correct. One policy definition can protect multiple tables and views.
    • B. Incorrect. An object can have only one row access policy at a time, so policies cannot be combined this way.
    • C. Correct. Snowflake evaluates the row access policy first and then masking policies.
    • D. Correct. Row access policies are supported on dynamic tables, which back feature views.
    • E. Incorrect. Future grants on row access policies are not supported.
    • F. Incorrect. Row access policies are attached to tables and views, not to tags.

    Subdomain 5.2: Monitor model health and compliance.

    32.An MLOps team wants a degradation alert that does not require keeping a warehouse running for a check that rarely triggers. Which privilege lets the alert owner role create an alert that runs on serverless compute?

    1. A.`USAGE` on a dedicated X-Small warehouse that the alert references in its `WAREHOUSE` parameter.
    2. B.`EXECUTE MANAGED ALERT` at account level, plus `EXECUTE ALERT` and `CREATE ALERT` on the schema.
    3. C.`MONITOR EXECUTION` on the account, which lets alerts borrow compute from the cloud services layer.
    4. D.`CREATE COMPUTE POOL` on the account, because serverless alerts run as containers in a compute pool.
    Show answer & explanation

    Correct answer: B — `EXECUTE MANAGED ALERT` at account level, plus `EXECUTE ALERT` and `CREATE ALERT` on the schema.

    • A. Incorrect: USAGE on a warehouse applies when the alert names a warehouse, which is the opposite of the serverless option.
    • B. Correct: serverless alerts omit the warehouse and need the account-level EXECUTE MANAGED ALERT privilege, plus the usual alert privileges.
    • C. Incorrect: MONITOR EXECUTION grants visibility into query execution and does not provide compute for alerts.
    • D. Incorrect: compute pools belong to Snowpark Container Services and are unrelated to alert execution.

    Subdomain 5.2: Monitor model health and compliance.

    33.An engineer wants to track an extra numeric business column, `basket_value`, in the monitor without it being treated as a model input. Which setting fits?

    1. A.List `basket_value` in `ID_COLUMNS`, which tracks any numeric column and exposes it as a statistic.
    2. B.List `basket_value` in `SEGMENT_COLUMNS`, which tracks numeric columns and computes drift for them separately.
    3. C.List `basket_value` in `CUSTOM_METRIC_COLUMNS`, which tracks extra NUMBER columns outside the feature set.
    4. D.List `basket_value` in `PREDICTION_SCORE_COLUMNS`, which tracks numeric columns beyond the model output.
    Show answer & explanation

    Correct answer: C — List `basket_value` in `CUSTOM_METRIC_COLUMNS`, which tracks extra NUMBER columns outside the feature set.

    • A. Incorrect: ID columns identify rows uniquely and are not a mechanism for tracking numeric business values.
    • B. Incorrect: segment columns must be STRING and only slice metrics by category.
    • C. Correct: CUSTOM_METRIC_COLUMNS holds extra NUMBER columns to track that are not treated as model features.
    • D. Incorrect: prediction score columns describe model outputs, so listing a business column there would corrupt performance metrics.

    Subdomain 5.3: Manage ML cost attribution and resource optimization.

    34.For which compute pool states does Snowpark Container Services bill credits?

    1. A.Billing applies per container hour for running services, with the pool nodes themselves free while the services are stopped and the pool is IDLE.
    2. B.Billing applies at a flat daily rate per compute pool regardless of state, and suspending the pool only pauses the node-hour portion of the charge.
    3. C.Billing applies while nodes are running or transitioning, such as IDLE and ACTIVE, and stops entirely when the pool is in the SUSPENDED state.
    4. D.Billing applies only while at least one service is scheduled, so a pool in the IDLE state with no running services accrues no compute credits.
    Show answer & explanation

    Correct answer: C — Billing applies while nodes are running or transitioning, such as IDLE and ACTIVE, and stops entirely when the pool is in the SUSPENDED state.

    • A. Incorrect: billing is by pool node, not by container, so stopping services alone does not remove the node charge.
    • B. Incorrect: there is no flat daily rate; charges follow the number and type of nodes while the pool is not suspended.
    • C. Correct: charges track the nodes provisioned in the pool, so IDLE and ACTIVE pools bill while a SUSPENDED pool, having no nodes, incurs no compute charge.
    • D. Incorrect: an IDLE pool still has nodes provisioned and is billed even though no service is running on it.

    Subdomain 5.3: Manage ML cost attribution and resource optimization.

    35.Beyond compute pool node charges, which additional cost categories apply to a Snowpark Container Services ML deployment? (Select all that apply.)(Select 3)

    1. A.Storage for service logs captured in an event table, which grows as containers emit output.
    2. B.A separate per-image pull fee charged each time a service container starts from a repository.
    3. C.Block storage volumes and their snapshots when services mount persistent volumes.
    4. D.Storage for container images held in an image repository, which is built on Snowflake stage storage.
    5. E.A licensing surcharge per GPU node that is billed outside the credit and storage consumption model.
    Show answer & explanation

    Correct answers: A, C, D — Storage for service logs captured in an event table, which grows as containers emit output.; Block storage volumes and their snapshots when services mount persistent volumes.; Storage for container images held in an image repository, which is built on Snowflake stage storage.

    • A. Correct: event table log data is stored in Snowflake and billed as storage.
    • B. Incorrect: no per-pull fee exists; image storage is the repository cost.
    • C. Correct: persistent volumes use block storage and snapshots, which are billed as storage.
    • D. Correct: image repositories use stage storage, which is billed like other stored data.
    • E. Incorrect: GPU node usage is billed through compute pool credits; there is no extra licensing line item.

    Want the full experience?

    These are just samples. Practice the full Snowflake SnowPro Advanced: MLOps Engineer (MLA-B01) question bank in quiz mode — free, no signup, with domain practice and exam simulation.