CertSafari

    Free Databricks Certified Machine Learning Professional Sample Questions

    35 free sample questions from our bank of 352+, covering every exam domain, with answers and detailed explanations. Updated September 2026.

    Domain 1: Model Development Using Spark ML

    Subdomain 1.2: Scaling and Tuning

    1.An engineer configures `TorchDistributor.init(num_processes=4, local_mode=True, use_gpu=True)` and calls `.run(train_fn)` on a Databricks cluster with a driver that has 4 GPUs and several worker nodes attached. Where does the training actually execute?

    1. A.All 4 processes run on the single driver machine, since `local_mode=True` confines TorchDistributor to launching multiple processes on one node.
    2. B.The 4 processes are spread one per worker node across the cluster, since `local_mode=True` always maps each process to a distinct physical node.
    3. C.TorchDistributor ignores `local_mode` when `use_gpu=True` and always distributes processes across every available GPU in the entire cluster.
    4. D.The call raises an error because `num_processes` must equal the total worker node count whenever `local_mode` is explicitly set to `True`.
    Show answer & explanation

    Correct answer: A — All 4 processes run on the single driver machine, since `local_mode=True` confines TorchDistributor to launching multiple processes on one node.

    • A. Correct: `local_mode=True` tells TorchDistributor to launch the specified number of processes on a single machine, so all 4 GPU processes run on the driver rather than being spread across workers.
    • B. Spreading one process per worker node is the behavior of multi-node mode (`local_mode=False`), not local mode, which intentionally keeps every process on one machine.
    • C. `use_gpu` only controls whether processes use GPU devices; it does not override `local_mode`, which independently governs whether execution stays on one node or spans the cluster.
    • D. `num_processes` in local mode specifies how many processes run on the single local machine and has no required relationship to the number of worker nodes in the cluster.

    Subdomain 1.2: Scaling and Tuning

    2.A hyperparameter search currently uses Hyperopt's `fmin` with `SparkTrials`. Which of the following are accurate reasons, per current Databricks guidance, to migrate this workload away from Hyperopt? (Select all that apply.)(Select 3)

    1. A.Hyperopt has been removed from Databricks Runtime ML versions released after 16.4 LTS, so it is no longer available by default on current runtimes.
    2. B.Optuna with `MlflowSparkStudy` provides a currently supported path for Spark-distributed hyperparameter trials with MLflow-backed shared state.
    3. C.Ray Tune is named alongside Optuna as a currently recommended option for distributed hyperparameter tuning on Databricks.
    4. D.Hyperopt was deprecated because it never supported any form of parallel trial execution, unlike every currently recommended tuning library.
    5. E.Migrating away from Hyperopt is only necessary for teams using GPU clusters, since CPU-only clusters continue to include Hyperopt indefinitely.
    6. F.Migrating away from Hyperopt additionally requires abandoning MLflow tracking entirely, since Optuna is documented as incompatible with logging results to MLflow.
    Show answer & explanation

    Correct answers: A, B, C — Hyperopt has been removed from Databricks Runtime ML versions released after 16.4 LTS, so it is no longer available by default on current runtimes.; Optuna with `MlflowSparkStudy` provides a currently supported path for Spark-distributed hyperparameter trials with MLflow-backed shared state.; Ray Tune is named alongside Optuna as a currently recommended option for distributed hyperparameter tuning on Databricks.

    • A. Correct: Hyperopt's removal from Databricks Runtime ML after version 16.4 LTS is exactly the documented reason current-runtime users lose access to it and must migrate.
    • B. Correct: `MlflowSparkStudy` combined with `MlflowStorage` is the documented current replacement path for Spark-distributed trials with shared, MLflow-backed state.
    • C. Correct: current guidance names Ray Tune, alongside Optuna, as a supported option for distributed hyperparameter tuning going forward.
    • D. Incorrect: Hyperopt did support parallel trial execution through `SparkTrials`; the deprecation is about runtime removal, not an inherent inability to run trials in parallel.
    • E. Incorrect: the removal of Hyperopt from Databricks Runtime ML after 16.4 LTS applies generally, not only to GPU clusters, so CPU-only clusters on current runtimes are affected too.
    • F. Incorrect: Optuna integrates directly with MLflow through `MlflowStorage` and `MlflowSparkStudy`, so migrating away from Hyperopt does not require giving up MLflow tracking.

    Subdomain 1.4: Advanced Feature Store Concepts

    3.A recommendation service needs millisecond-latency reads against a feature table for online inference. The team already has the feature table in Unity Catalog and wants to provision the low-latency copy using the Databricks SDK. What is the correct approach?

    1. A.Build an `OnlineTableSpec` naming the table and keys, wrap it in an `OnlineTable`, and call `online_tables.create_and_wait(table=t)`.
    2. B.Call `spark.sql("CACHE TABLE feature_table")` in the notebook, which SDK docs describe as the supported way to provision online tables.
    3. C.Set the feature table's `delta.enableChangeDataFeed` property to `true` and query it directly from the endpoint, skipping any online table object.
    4. D.Use `dbutils.fs.cp` to copy the feature table's underlying Parquet files onto the serving cluster's local disk before the endpoint starts.
    Show answer & explanation

    Correct answer: A — Build an `OnlineTableSpec` naming the table and keys, wrap it in an `OnlineTable`, and call `online_tables.create_and_wait(table=t)`.

    • A. This is the documented SDK pattern: an `OnlineTableSpec` describes the source table and keys, an `OnlineTable` object wraps that spec, and `create_and_wait` provisions the low-latency online copy.
    • B. `CACHE TABLE` caches a DataFrame in cluster memory for the life of that Spark session; it does not create a durable, low-latency online table and is not part of the online table provisioning API.
    • C. Enabling Change Data Feed is a prerequisite for some publishing modes, but it does not by itself create an online table; a serving endpoint cannot query an unprovisioned online table just because CDF is on.
    • D. Copying raw Parquet files onto local disk bypasses the managed online table infrastructure entirely and does not produce anything a serving endpoint or `FeatureLookup` can query.

    Subdomain 1.4: Advanced Feature Store Concepts

    4.A team is designing a Structured Streaming job that reads raw events and writes aggregated features into a feature table via `fe.write_table(mode="merge")`. Which of the following are genuine design considerations for this streaming feature pipeline?(Select 3)

    1. A.Choosing a checkpoint location so the streaming query can recover its progress and avoid reprocessing or dropping events after a restart.
    2. B.Applying a watermark on the event-time column to bound how long the job retains state for late-arriving events.
    3. C.Picking a trigger interval that matches how fresh the feature values need to be, such as a fixed processing time or `availableNow`.
    4. D.Setting the write mode to `"overwrite"` instead of `"merge"`, since only `"overwrite"` is supported once a streaming source is involved.
    5. E.Disabling Change Data Feed on the feature table, since streaming writes are documented as incompatible with Change Data Feed.
    6. F.Repartitioning the streaming DataFrame down to a single partition before each write, which is required for merge writes to succeed.
    Show answer & explanation

    Correct answers: A, B, C — Choosing a checkpoint location so the streaming query can recover its progress and avoid reprocessing or dropping events after a restart.; Applying a watermark on the event-time column to bound how long the job retains state for late-arriving events.; Picking a trigger interval that matches how fresh the feature values need to be, such as a fixed processing time or `availableNow`.

    • A. A checkpoint location lets Structured Streaming persist offsets and state, so the job can resume correctly after a restart instead of reprocessing or losing events.
    • B. A watermark on event time bounds how long windowed aggregation state is retained, which is necessary to keep the job's memory and state store from growing without limit as late events trickle in.
    • C. The trigger interval controls how often micro-batches run and therefore how fresh the written feature values are, so it should be chosen to match the freshness the downstream use case needs.
    • D. `"merge"` is the only mode `write_table` supports for both batch and streaming input; there is no requirement or need to switch to `"overwrite"` for streaming sources.
    • E. Change Data Feed is unrelated to whether the write side is streaming; it is a source-table property used for downstream incremental consumption and is not something streaming writes require disabling.
    • F. Merge writes do not require single-partition output; forcing one partition would only hurt throughput and is not a documented requirement of `write_table`.

    Subdomain 1.3: Advanced MLflow Usage

    5.A training script computes accuracy, precision, recall, and F1 score at the end of each epoch and currently makes four separate `mlflow.log_metric()` calls per epoch inside a tight loop over 200 epochs, resulting in 800 individual tracking-server calls. An engineer wants to reduce the number of network round trips while logging the exact same information. Which change accomplishes this?

    1. A.Replace the four `log_metric` calls with a single `mlflow.log_metrics({"accuracy": acc, "precision": prec, "recall": rec, "f1": f1}, step=epoch)` call.
    2. B.Replace the four calls with `mlflow.log_param({"accuracy": acc, "precision": prec, "recall": rec, "f1": f1})`, since parameters accept dictionaries natively.
    3. C.Wrap the existing four calls inside `mlflow.start_run(nested=True)` for each epoch, which batches all metric writes from that nested run into one request.
    4. D.Call `mlflow.log_artifact()` once per epoch with a JSON file containing all four values, since artifacts always require fewer network calls than metrics.
    Show answer & explanation

    Correct answer: A — Replace the four `log_metric` calls with a single `mlflow.log_metrics({"accuracy": acc, "precision": prec, "recall": rec, "f1": f1}, step=epoch)` call.

    • A. `log_metrics` accepts a dictionary and sends all the named values for a given step in a single tracking-server call, which cuts four calls down to one per epoch while logging identical data.
    • B. `log_param` stores fixed configuration values and is not intended for per-epoch numeric series; using it here also does not solve the batching goal, since it is still a per-call operation.
    • C. Opening a nested run per epoch adds run-management overhead and does not batch metric writes into fewer calls; it would in fact multiply the number of tracking-server interactions.
    • D. Uploading a JSON artifact each epoch still requires one call per epoch and adds file-storage overhead, and it loses the native metric charting that `log_metrics` provides in the UI.

    Subdomain 1.3: Advanced MLflow Usage

    6.After training, an engineer has both a single serialized preprocessing object, `scaler.pkl`, and an entire local folder, `evaluation_report/`, containing several plots and a metrics CSV. They want to log the single file and the whole folder to the current MLflow run as artifacts. Which pair of calls is correct?

    1. A.`mlflow.log_artifact("scaler.pkl")` for the single file, and `mlflow.log_artifacts("evaluation_report")` for the folder, since the plural form uploads a directory's contents.
    2. B.`mlflow.log_artifacts("scaler.pkl")` for the single file, and `mlflow.log_artifact("evaluation_report")` for the folder, since the plural form is required for any file with an extension.
    3. C.`mlflow.log_artifact("scaler.pkl")` and `mlflow.log_artifact("evaluation_report")`, since the singular form automatically detects whether the path is a file or a directory.
    4. D.`mlflow.log_model("scaler.pkl")` for the single file, and `mlflow.log_artifacts("evaluation_report")` for the folder, since only `log_model` supports serialized Python objects.
    Show answer & explanation

    Correct answer: A — `mlflow.log_artifact("scaler.pkl")` for the single file, and `mlflow.log_artifacts("evaluation_report")` for the folder, since the plural form uploads a directory's contents.

    • A. `log_artifact` uploads one local file, while `log_artifacts` walks a local directory and uploads its full contents, so pairing them this way matches each call to its intended input type.
    • B. This reverses the actual behavior: `log_artifacts` expects a directory path and `log_artifact` expects a single file path, regardless of whether that file has a recognizable extension.
    • C. `log_artifact` does not accept a directory path; passing a folder to it raises an error, since only `log_artifacts` is designed to recurse through a directory's contents.
    • D. `log_model` is for logging a trained model in a specific ML flavor format with a signature, not for arbitrary serialized files, and it does not accept a directory of mixed report files.

    Subdomain 1.3: Advanced MLflow Usage

    7.A team enables `mlflow.autolog()` at the top of their training script for a scikit-learn `RandomForestClassifier`, which automatically logs the model's hyperparameters, training metrics, and the fitted model artifact. They also want to additionally log a custom business metric, projected cost savings, that autologging has no way of knowing about. Which statements about combining autologging with custom logging are correct?(Select 2)

    1. A.A custom call like `mlflow.log_metric("projected_cost_savings", value)` can be added inside the same active run that `autolog()` is populating, without disabling autologging.
    2. B.Autologging and manual `log_metric` calls write to the same active run object, so both the automatic and the custom metrics appear together on one run's page.
    3. C.Calling `mlflow.autolog()` locks the active run against any further logging calls, so custom metrics must be logged in a completely separate follow-up run.
    4. D.The custom `projected_cost_savings` metric must be named with an `autolog_` prefix, or the autologging integration will silently discard it during the run.
    5. E.Enabling autologging replaces the need to call `mlflow.start_run()` at all, since autologging manages its own separate run outside of the script's control.
    6. F.Custom metrics added after `mlflow.autolog()` is enabled overwrite any automatically logged metric that happens to share the same name, since manual calls always take precedence.
    Show answer & explanation

    Correct answers: A, B — A custom call like `mlflow.log_metric("projected_cost_savings", value)` can be added inside the same active run that `autolog()` is populating, without disabling autologging.; Autologging and manual `log_metric` calls write to the same active run object, so both the automatic and the custom metrics appear together on one run's page.

    • A. Autologging populates whatever run is active without preventing additional manual logging calls, so a custom metric like cost savings can be logged into that same run alongside the autologged values.
    • B. Both autologging and manual calls target the currently active MLflow run, so a reviewer opening that run's page sees the automatically captured metrics together with any manually logged custom ones.
    • C. Autologging does not lock the run against further calls; it simply adds its own logging calls under the hood while leaving the run fully open to any additional manual logging the script performs.
    • D. There is no naming convention required for custom metrics to coexist with autologged ones; any metric name not already used by the autologging integration logs normally.
    • E. Autologging still requires an active run, either one explicitly started with `start_run()` or one it creates automatically if none is active; it does not remove the need for run management.
    • F. If a custom metric name collides with one autologging already logs, the later call simply overwrites the earlier value for that metric name, which is not guaranteed to favor either source and can be surprising rather than a documented precedence rule.

    Subdomain 1.1: Model Development Using Spark ML

    8.A machine learning engineer is assembling a Spark ML `Pipeline` with a `StringIndexer`, a `OneHotEncoder`, a `VectorAssembler`, and a `LogisticRegression` estimator as its stages, in that order. What happens when `pipeline.fit(trainingData)` is called?

    1. A.Each stage runs in order: `Transformer` stages transform immediately, `Estimator` stages fit and are replaced by fitted transformers, giving a `PipelineModel`.
    2. B.Only the final `LogisticRegression` stage executes; the earlier `StringIndexer`, `OneHotEncoder`, and `VectorAssembler` stages act as inert configuration only.
    3. C.All four stages are fit and executed in parallel across the cluster, and their outputs are merged into a single feature vector before logistic regression begins.
    4. D.The call raises an exception because a `Pipeline` cannot mix indexing, encoding, and assembling stages together with a classification estimator in one object.
    Show answer & explanation

    Correct answer: A — Each stage runs in order: `Transformer` stages transform immediately, `Estimator` stages fit and are replaced by fitted transformers, giving a `PipelineModel`.

    • A. `Pipeline.fit()` runs each stage in sequence: transformer stages like `OneHotEncoder` and `VectorAssembler` apply their transform immediately, while estimator stages like `StringIndexer` and `LogisticRegression` are fit on the data produced by prior stages and become fitted transformers inside the returned `PipelineModel`.
    • B. Every stage in the list actually executes; the earlier stages are not skipped or treated as inert configuration, since each one produces the transformed columns the next stage depends on.
    • C. Pipeline stages run sequentially because each stage typically consumes the output column of the previous one, so they cannot be executed in parallel as independent, unrelated steps.
    • D. Mixing indexing, encoding, assembling, and a classifier in one `Pipeline` is the standard pattern the API is designed for, so no exception is raised for combining these stage types.

    Subdomain 1.1: Model Development Using Spark ML

    9.A team previously used Hyperopt's `SparkTrials` to distribute hyperparameter search for a scikit-learn model across a Databricks cluster. They are now setting up a new project on a current Databricks Runtime ML version. What should they know about this approach?

    1. A.Hyperopt is removed from Databricks Runtime ML after version 16.4 LTS, and current Databricks guidance recommends Optuna, with MLflow 3 integration, or Ray Tune instead.
    2. B.Hyperopt remains the only Databricks-recommended distributed tuning library, and `SparkTrials` is required for any hyperparameter search run on a Databricks cluster today.
    3. C.Hyperopt's `SparkTrials` was renamed to `MlflowSparkStudy` in current runtimes, but its API and distributed execution model are otherwise unchanged.
    4. D.Hyperopt is only deprecated for Spark ML estimators; it remains fully supported and recommended for tuning single-node scikit-learn models on current runtimes.
    Show answer & explanation

    Correct answer: A — Hyperopt is removed from Databricks Runtime ML after version 16.4 LTS, and current Databricks guidance recommends Optuna, with MLflow 3 integration, or Ray Tune instead.

    • A. Databricks removed Hyperopt from Databricks Runtime ML after the 16.4 LTS release, and current documentation points teams toward Optuna, which integrates with MLflow 3 through `MlflowStorage` and `MlflowSparkStudy`, or toward Ray Tune for distributed search.
    • B. This is the opposite of current guidance: Hyperopt is deprecated on current runtimes, and teams starting new projects are directed toward Optuna or Ray Tune rather than continuing to rely on `SparkTrials`.
    • C. `MlflowSparkStudy` is part of the newer Optuna-with-MLflow-3 integration, not a renamed version of Hyperopt's `SparkTrials`; the two come from different libraries with different APIs.
    • D. Hyperopt's removal from Databricks Runtime ML after 16.4 LTS applies regardless of whether it is tuning a Spark ML estimator or a single-node scikit-learn model, so it is not selectively deprecated by model type.

    Subdomain 1.1: Model Development Using Spark ML

    10.A multiclass product-category classifier has ten roughly balanced classes. The team wants to report a single overall correctness metric as well as a metric that accounts for per-class support when averaging precision across the ten classes. Select the valid `metricName` values for `MulticlassClassificationEvaluator` that provide these two readings.(Select 2)

    1. A.`"accuracy"`, which reports the overall fraction of rows correctly classified across all ten classes.
    2. B.`"weightedPrecision"`, which averages per-class precision weighted by each class's support in the dataset.
    3. C.`"f1"`, which reports the harmonic mean of precision and recall, averaged across the ten classes.
    4. D.`"areaUnderROC"`, which reports the area under the receiver operating characteristic curve for the classifier.
    5. E.`"rmse"`, which reports the root mean squared error between the predicted and true class labels.
    6. F.`"r2"`, which reports the proportion of variance in the true class label explained by the model.
    Show answer & explanation

    Correct answers: A, B — `"accuracy"`, which reports the overall fraction of rows correctly classified across all ten classes.; `"weightedPrecision"`, which averages per-class precision weighted by each class's support in the dataset.

    • A. `"accuracy"` is a valid `MulticlassClassificationEvaluator` metric name and directly reports the overall fraction of correctly classified rows across all classes, matching the request for a single correctness metric.
    • B. `"weightedPrecision"` is a valid metric name that averages per-class precision weighted by class support, matching the request for a support-weighted precision average across the ten classes.
    • C. `"f1"` is a valid `MulticlassClassificationEvaluator` metric, but it reports a precision/recall harmonic mean rather than a support-weighted precision average, so it does not satisfy the second requirement specifically.
    • D. `"areaUnderROC"` belongs to `BinaryClassificationEvaluator` for two-class problems; it is not a valid `metricName` on `MulticlassClassificationEvaluator` for a ten-class problem.
    • E. `"rmse"` is a `RegressionEvaluator` metric for continuous numeric targets; it is not a valid `metricName` on `MulticlassClassificationEvaluator` for categorical class labels.
    • F. `"r2"` is a `RegressionEvaluator` metric describing explained variance for continuous targets; it is not defined for a multiclass categorical classification problem.

    Subdomain 1.2: Scaling and Tuning

    11.A single-node model currently takes 6 hours to train on a Databricks cluster driver, and the team is deciding whether to move to a bigger single-node instance type or to a horizontally scaled multi-node cluster. Which consideration most strongly favors staying with vertical scaling for this workload?

    1. A.The training algorithm is inherently sequential and cannot be partitioned across workers, so a larger single machine avoids the coordination overhead of a distributed job that would not actually parallelize the computation.
    2. B.The dataset no longer fits in the memory or disk of any available single-node instance type, so the team should scale out to a cluster whose combined worker memory and distributed storage can hold the full training set.
    3. C.The training job already scales near-linearly with additional workers on the current distributed framework, so the team should add more worker nodes to the multi-node cluster and expect training time to drop in proportion to the extra compute.
    4. D.The organization wants training throughput to keep growing indefinitely as data volume increases over the next several years, so the team should build out a horizontally scaled cluster now that can keep adding worker nodes without re-architecting the pipeline later.
    Show answer & explanation

    Correct answer: A — The training algorithm is inherently sequential and cannot be partitioned across workers, so a larger single machine avoids the coordination overhead of a distributed job that would not actually parallelize the computation.

    • A. Correct: when an algorithm cannot be meaningfully partitioned, distributing it across nodes adds network and coordination overhead without unlocking real parallel speedup, so a bigger single machine is the more effective lever.
    • B. A dataset too large for any single instance is a textbook argument for horizontal scaling, not vertical scaling, since no single machine can hold the data regardless of size.
    • C. Near-linear scaling with added workers is precisely the property that makes horizontal scaling attractive, so this observation favors adding nodes rather than favoring a bigger single machine.
    • D. Indefinite future growth is an argument for an elastic, horizontally scalable architecture rather than for committing to ever-larger single instances, which eventually hit a hardware ceiling.

    Subdomain 1.2: Scaling and Tuning

    12.A data scientist calls `MlflowSparkStudy(study_name="tuning", storage=mlflow_storage)` and then `mlflow_study.optimize(objective, n_trials=40, n_jobs=8)` without specifying a sampler. Which statement correctly describes what happens by default?

    1. A.Optuna defaults to the tree-structured Parzen estimator sampler, so the 40 trials are proposed adaptively based on prior results while 8 run concurrently across executors.
    2. B.Optuna defaults to pure random search across every hyperparameter dimension, so the 40 trials sample independently and adaptive samplers like TPE only take effect once passed explicitly to `optimize`.
    3. C.Optuna requires an explicit sampler argument in the `optimize` call, so this configuration raises a `ValueError` immediately and none of the 40 requested trials begin executing across the 8 workers.
    4. D.Optuna defaults to exhaustive grid search over the parameter ranges implied by the `suggest_float` and `suggest_categorical` calls in the objective, enumerating every combination across the 8 workers.
    Show answer & explanation

    Correct answer: A — Optuna defaults to the tree-structured Parzen estimator sampler, so the 40 trials are proposed adaptively based on prior results while 8 run concurrently across executors.

    • A. Correct: Optuna studies default to the TPESampler, an adaptive Bayesian-style sampler, so trials are proposed based on the outcomes of earlier trials rather than sampled independently at random.
    • B. Plain random search is available as an explicit sampler choice, but it is not Optuna's default; the default sampler already performs adaptive, history-informed suggestion.
    • C. No sampler argument is required to start a study; omitting it simply causes Optuna to fall back to its default TPESampler rather than raising an error.
    • D. Grid search is a separate, explicitly selectable Optuna sampler; the default sampler explores the continuous and categorical space adaptively rather than exhaustively enumerating a grid.

    Subdomain 1.4: Advanced Feature Store Concepts

    13.In a `FeatureLookup` used for a point-in-time join, what does the value passed to `timestamp_lookup_key` represent?

    1. A.The name of the column in the label DataFrame containing the timestamp against which the feature table's time series values are matched.
    2. B.The name of the column in the feature table that stores the Delta Lake commit timestamp marking when each row was physically written to the table.
    3. C.A fixed cutoff date applied uniformly to every row in the training set, using the maximum label timestamp instead of each row's own event time.
    4. D.The refresh interval that the Databricks Feature Engineering client uses when re-publishing the feature table to the online store for serving.
    Show answer & explanation

    Correct answer: A — The name of the column in the label DataFrame containing the timestamp against which the feature table's time series values are matched.

    • A. `timestamp_lookup_key` names the column on the label/spine DataFrame holding each row's own timestamp, which the as-of join compares against the feature table's timestamp key to find the correct historical value per row.
    • B. Write time to the Delta table is unrelated to `timestamp_lookup_key`; the parameter refers to a column on the label DataFrame, not metadata about when feature rows were physically ingested.
    • C. The parameter is per-row, not a single fixed cutoff — every label row can carry its own distinct timestamp, and the join is evaluated independently for each one.
    • D. Publishing cadence for online tables is controlled by scheduling policies like `TRIGGERED` or `CONTINUOUS`, which are unrelated to `timestamp_lookup_key` and to training-set construction entirely.

    Domain 2: MLOps Model Lifecycle Management

    Subdomain 2.1: MLOps Model Lifecycle Management

    14.After a model pipeline is deployed, ML engineers manage the running jobs and serving endpoints in the production workspace, while data scientists who built the model can view logs and metrics for debugging but cannot modify production assets. Which environment does this describe?

    1. A.Staging, where automated integration tests run before any human reviews the pipeline.
    2. B.Development, where data scientists hold full read-write control over every asset they create.
    3. C.Production, where ML engineers own pipelines and data scientists keep read-only visibility.
    4. D.A disaster-recovery workspace that only activates when the primary production workspace fails.
    Show answer & explanation

    Correct answer: C — Production, where ML engineers own pipelines and data scientists keep read-only visibility.

    • A. Staging is for automated pipeline testing before promotion, not for ML engineers managing live jobs while data scientists hold read-only debugging access.
    • B. Development grants data scientists broad read-write control, which is the opposite of the read-only visibility described for the environment in this scenario.
    • C. This matches the production environment's access pattern: ML engineers own and operate the deployed pipeline, while data scientists keep read-only visibility for monitoring and debugging rather than write access.
    • D. The scenario describes normal day-to-day operation of the live pipeline, not a failover workspace that only activates during an outage.

    Subdomain 2.1: MLOps Model Lifecycle Management

    15.A platform team needs a way to orchestrate the sequence of steps in a CI/CD deployment pipeline, running validation, comparison, and promotion steps automatically once code merges to the release branch. Which Databricks feature is responsible for this lifecycle activity?

    1. A.Lakehouse federation, which lets queries reach external data sources without moving data into Databricks.
    2. B.MLflow Model Registry, which tracks model version lineage and alias assignments after training completes.
    3. C.Lakeflow Jobs (Databricks Workflows), which orchestrates the execution of pipeline steps in sequence.
    4. D.Cluster policies, which restrict the compute configurations a workspace user is allowed to launch.
    Show answer & explanation

    Correct answer: C — Lakeflow Jobs (Databricks Workflows), which orchestrates the execution of pipeline steps in sequence.

    • A. Lakehouse federation is about querying external data sources; it has no role in orchestrating the steps of a CI/CD deployment pipeline.
    • B. The Model Registry records lineage and alias state after each step completes, but it does not itself sequence or trigger the validation, comparison, and promotion steps.
    • C. Lakeflow Jobs is the orchestration engine that runs validation, comparison, and promotion steps in sequence once code merges, which is exactly the coordination activity described.
    • D. Cluster policies constrain compute configuration choices; they do not orchestrate the order of pipeline steps in a CI/CD deployment.

    Subdomain 2.1: MLOps Model Lifecycle Management

    16.Which of the following feature-to-activity pairings are correct within the Databricks model lifecycle management process? (Select all that apply.)(Select 3)

    1. A.Lakeflow Jobs are used to orchestrate the steps of a CI/CD promotion pipeline across environments.
    2. B.Lakeflow Spark Declarative Pipelines are used to serve real-time model predictions to applications.
    3. C.Databricks Repos are used to assign the Challenger alias to a newly validated model version.
    4. D.MLflow Tracking is used to log and compare parameters and metrics during model development.
    5. E.Unity Catalog model aliases are used to mark which model version is currently serving as Champion.
    Show answer & explanation

    Correct answers: A, D, E — Lakeflow Jobs are used to orchestrate the steps of a CI/CD promotion pipeline across environments.; MLflow Tracking is used to log and compare parameters and metrics during model development.; Unity Catalog model aliases are used to mark which model version is currently serving as Champion.

    • A. Lakeflow Jobs is the orchestration tool used to run and sequence the steps of a CI/CD promotion pipeline as code moves across environments.
    • B. Serving real-time predictions is the role of a Model Serving endpoint, not Lakeflow Spark Declarative Pipelines, which builds data and feature pipelines instead.
    • C. Alias assignment happens through Unity Catalog's model registry mechanisms, not through Repos, which only manages source code synchronization with a Git remote.
    • D. MLflow Tracking is the tool used to log and compare parameters, metrics, and artifacts across experiment runs while a model is being developed.
    • E. Unity Catalog model aliases, such as Champion, are exactly the mechanism used to mark which model version is currently designated for production serving.

    Subdomain 2.4: Automated Retraining

    17.A team is setting up Lakehouse Monitoring on an inference table for a new fraud model and must choose what data to load into the baseline table so future drift comparisons are meaningful. What should the baseline table contain?

    1. A.A representative sample of the training or validation data whose distribution reflects what the model was built to expect.
    2. B.The very first day of raw, unprocessed production traffic captured before any data cleaning or filtering was applied.
    3. C.A synthetic dataset generated to cover every possible feature value, regardless of how likely those values are in production.
    4. D.The current day's live inference table itself, refreshed continuously so the baseline always matches the newest traffic.
    Show answer & explanation

    Correct answer: A — A representative sample of the training or validation data whose distribution reflects what the model was built to expect.

    • A. A baseline built from the training or validation distribution gives drift metrics a stable, meaningful reference point representing what the model expects, which is what current production data is compared against.
    • B. Unprocessed raw traffic from before any cleaning does not represent the distribution the model was trained and validated on, so comparisons against it would flag differences that have nothing to do with real drift.
    • C. A synthetic dataset covering unlikely feature combinations skews the baseline away from realistic production behavior, making drift statistics measure divergence from an artificial reference rather than the model's actual training assumptions.
    • D. Continuously refreshing the baseline from the live table eliminates any fixed reference point, so drift can never be detected because the baseline would always move in lockstep with the data being compared.

    Subdomain 2.4: Automated Retraining

    18.An automated retraining pipeline promotes a new Champion by reassigning the alias, and the inference job runs nightly, always loading whatever model version currently holds the "Champion" alias. Why does this pattern let retraining automation update production behavior without any change to the inference job's code or configuration?

    1. A.Because the inference job resolves the alias to a model version at run time, picking up a new assignment on its next run.
    2. B.Because Unity Catalog automatically restarts and reruns every job that references a model whenever its Champion alias changes.
    3. C.Because the inference job caches the model version permanently, so alias changes only take effect after a manual cache clear.
    4. D.Because aliases are tied to the job's cluster configuration and node pool rather than to the registered model itself.
    Show answer & explanation

    Correct answer: A — Because the inference job resolves the alias to a model version at run time, picking up a new assignment on its next run.

    • A. The alias is resolved to whichever version currently holds it each time the inference job runs, so reassigning the Champion alias is enough for the next scheduled run to load the new version with no code or config edits.
    • B. Unity Catalog does not automatically restart or re-trigger jobs when an alias changes; the inference job simply resolves the alias the next time it happens to run on its own schedule or trigger.
    • C. There is no permanent caching of a specific model version tied to the alias that requires a manual clear; each run resolves the alias fresh, which is exactly what lets the new version take effect automatically.
    • D. Aliases are a property of the registered model and its versions in Unity Catalog, not of the job's cluster configuration, so cluster settings have nothing to do with how the alias is resolved.

    Subdomain 2.4: Automated Retraining

    19.A team wants their automated retraining pipeline to react specifically to changes in the data itself, rather than to a fixed calendar or to unrelated model registry events. Select the three trigger mechanisms that react to a data-related change.(Select 3)

    1. A.A Table update trigger, which fires when the referenced Delta table receives a new commit.
    2. B.A File arrival trigger, which fires when new objects land in a monitored storage location.
    3. C.A SQL alert on the drift metrics table that calls the Jobs API once a drift threshold is breached.
    4. D.A Model update trigger, which fires on Unity Catalog registry events such as a new model version.
    5. E.A Continuous trigger, which fires as soon as the job's own previous run completes or fails.
    6. F.A Scheduled trigger, which fires at fixed calendar times regardless of whether any data changed.
    Show answer & explanation

    Correct answers: A, B, C — A Table update trigger, which fires when the referenced Delta table receives a new commit.; A File arrival trigger, which fires when new objects land in a monitored storage location.; A SQL alert on the drift metrics table that calls the Jobs API once a drift threshold is breached.

    • A. This is correct: a Table update trigger fires specifically because the underlying data table changed, which is a direct reaction to a data-related event.
    • B. This is correct: a File arrival trigger reacts to new data objects showing up in storage, which is also a direct, data-driven event rather than a calendar or registry-based one.
    • C. This is correct: a SQL alert tied to the drift metrics table reacts to a measured change in the data's distribution, making it a data-driven trigger even though it is implemented through an alert-to-API call rather than a native Jobs trigger type.
    • D. A Model update trigger reacts to registry events like a new model version or alias assignment, which are events about the model artifact, not about the underlying training or production data changing.
    • E. A Continuous trigger reacts only to the prior run of the same job finishing, which has nothing to do with whether the underlying data has actually changed.
    • F. A Scheduled trigger fires purely based on calendar time, so it runs on its fixed cadence whether or not any data has changed at all.

    Subdomain 2.2: Validation Testing

    20.A Databricks ML team is asked to justify, in an architecture review, why they organize their Python project with functions in importable modules, unit tests in a corresponding test suite, and a staging integration test suite that runs the assembled pipeline. Which justification is most accurate?

    1. A.Modules make functions independently testable, unit tests catch function defects early, and staging integration tests confirm the pipeline works as a whole.
    2. B.Modules, unit tests, and integration tests are three names for identical testing work performed on different days of the calendar year.
    3. C.Modules and unit tests are optional conveniences with no real effect on defects, and only the staging integration suite meaningfully catches bugs.
    4. D.Modules exist only to satisfy Scala packaging requirements, unit tests exist only for coverage metrics, and integration tests exist only for compliance audits.
    Show answer & explanation

    Correct answer: A — Modules make functions independently testable, unit tests catch function defects early, and staging integration tests confirm the pipeline works as a whole.

    • A. This matches the documented reasoning: modules let Python functions be imported and tested outside notebooks, unit tests are added continuously during development to catch function-level issues early, and staging integration tests run the whole pipeline together to catch cross-stage defects before production.
    • B. Modules, unit tests, and integration tests are distinct concepts serving different purposes, module organization enables testability, unit tests check individual functions, and integration tests check the assembled pipeline, not different names for one activity.
    • C. Unit tests catch function-level defects early and cheaply, which is a meaningful and distinct benefit from integration testing; treating them as having no effect ignores the documented value of catching issues before the more expensive integration stage.
    • D. Modules are a general Python organization pattern, not a Scala-specific requirement; coverage metrics and compliance audits are not the stated reasons for unit or integration testing, which instead exist to catch actual functional defects.

    Subdomain 2.2: Validation Testing

    21.During a retrospective, a team lists several claims about their organization of functions, unit tests, and integration tests in staging. Which of these claims are consistent with documented Databricks testing guidance? (Select all that apply.)(Select 3)

    1. A.Integration testing in staging is meant to confirm that feature engineering, training, and inference work correctly together, not just individually.
    2. B.For Python code, storing functions outside notebooks in modules improves both reusability and compatibility with standard test frameworks.
    3. C.Unit tests are meant to replace integration tests entirely, since a sufficiently thorough set of unit tests always implies pipeline-level correctness.
    4. D.Staging is designed to closely mirror production so that the full test suite can run under representative conditions before promotion.
    5. E.For SQL logic, functions are best organized by embedding the SQL directly inside Python notebook cells rather than as SQL user-defined functions.
    Show answer & explanation

    Correct answers: A, B, D — Integration testing in staging is meant to confirm that feature engineering, training, and inference work correctly together, not just individually.; For Python code, storing functions outside notebooks in modules improves both reusability and compatibility with standard test frameworks.; Staging is designed to closely mirror production so that the full test suite can run under representative conditions before promotion.

    • A. This matches documented guidance on integration testing: it exists to validate that pipelines including feature engineering, training, and inference function correctly together, beyond what testing each in isolation would show.
    • B. This matches documented guidance for Python: moving functions into modules outside notebooks makes them usable both inside and outside notebooks and lets standard test frameworks run against them more directly than against notebook cells.
    • C. This is not accurate; passing unit tests for every individual function does not guarantee that the functions work correctly together as a connected pipeline, which is exactly why integration testing exists as a distinct, necessary stage.
    • D. This matches documented guidance: staging is built to mirror production as closely as feasible and is where the complete testing suite runs before code or artifacts are promoted further.
    • E. This contradicts documented guidance for SQL, which recommends organizing SQL logic as SQL user-defined functions stored within a schema, not by embedding SQL inside Python notebook cells.

    Subdomain 2.3: Environment Architectures

    22.An MLOps engineer is authoring a Databricks Asset Bundle (DAB) and needs the bundle deploy to create and manage a real-time serving endpoint for a Unity Catalog model, fully described in version-controlled YAML alongside the training job. Under which top-level `resources` key should this endpoint be declared?

    1. A.model_serving_endpoints
    2. B.serving_endpoints
    3. C.model_endpoints
    4. D.endpoints
    Show answer & explanation

    Correct answer: A — model_serving_endpoints

    • A. model_serving_endpoints is the resource key Databricks Asset Bundles use for declaring a Model Serving endpoint, including its served_entities and traffic_config, in the resources section.
    • B. serving_endpoints is not a recognized Databricks Asset Bundle resource key; the bundle schema uses the fuller model_serving_endpoints name for this resource type.
    • C. model_endpoints is not a valid Databricks Asset Bundle resource key; no such shortened name exists in the resources schema.
    • D. endpoints alone is too generic and is not the resource key the bundle schema expects for Model Serving endpoints.

    Subdomain 2.3: Environment Architectures

    23.A platform team is designing environment isolation for a highly regulated ML project and wants layered controls beyond Unity Catalog namespace separation alone. Which of the following genuinely add a distinct layer of isolation or governance beyond catalog-level separation? (Select all that apply.)(Select 3)

    1. A.Deploy production compute to its own Databricks workspace with its own network configuration, separate from the workspace used for dev and staging.
    2. B.Use a distinct service principal per environment for CI/CD deployments, so a compromised or misconfigured dev credential cannot deploy or modify prod resources.
    3. C.Scope `grants` on production `registered_models` resources to only the specific principals that need access, rather than granting broadly to every workspace user.
    4. D.Reuse one shared service principal with broad admin rights across dev, staging, and prod, since a single credential is simpler for the CI/CD pipeline to manage.
    5. E.Rely solely on catalog-level separation, skipping workspace-level or identity-level controls, since Unity Catalog namespaces alone are sufficient for regulated workloads.
    6. F.Give every data scientist the same permission level in every target, since consistent permissions across environments simplifies onboarding regardless of sensitivity.
    Show answer & explanation

    Correct answers: A, B, C — Deploy production compute to its own Databricks workspace with its own network configuration, separate from the workspace used for dev and staging.; Use a distinct service principal per environment for CI/CD deployments, so a compromised or misconfigured dev credential cannot deploy or modify prod resources.; Scope `grants` on production `registered_models` resources to only the specific principals that need access, rather than granting broadly to every workspace user.

    • A. A separate workspace with its own network configuration adds an isolation layer at the compute and network level, which catalog separation inside a shared workspace does not provide on its own.
    • B. Distinct per-environment service principals limit the blast radius of a compromised or misconfigured dev credential, adding an identity-level control on top of namespace separation.
    • C. Scoping grants tightly to the specific principals that need access on production model resources adds a fine-grained authorization layer beyond simply placing the model in a separate catalog.
    • D. A single shared service principal with broad rights across every environment removes the identity-level boundary between environments, so a dev-side compromise could reach production resources.
    • E. Catalog separation alone governs data and metadata namespace boundaries; it does not address compute network isolation or scoped CI/CD identities, so relying on it exclusively leaves those layers uncovered for a regulated workload.
    • F. Giving every data scientist identical permissions regardless of environment ignores that production sensitivity typically calls for tighter access than a dev sandbox, undermining layered governance rather than adding to it.

    Subdomain 2.3: Environment Architectures

    24.A reviewer is doing a final check on a bundle's ML-specific resources — `model_serving_endpoints`, `experiments`, and `registered_models` — before its first production deployment. Which of the following are legitimate things to verify at this stage? (Select all that apply.)(Select 3)

    1. A.Confirm the served entity references the model by alias (such as `@Champion`) rather than a version number hardcoded to whatever was true when the YAML was written.
    2. B.Confirm `catalog_name` and `schema_name` on the `registered_models` resource resolve to the intended production namespace once the prod target's variables are applied.
    3. C.Confirm `grants` on the production `registered_models` resource are scoped to the specific principals that need access, not granted broadly by default.
    4. D.Confirm the `experiments` resource's `comment` field is left blank, since Databricks documentation states a populated comment field is known to break MLflow run logging.
    5. E.Confirm the `model_serving_endpoints` resource has `traffic_config` omitted entirely, since Model Serving endpoints route all traffic automatically without any configuration.
    6. F.Confirm `entity_version` is hardcoded to the highest currently existing version so the endpoint never needs to be redeployed for any future promotion.
    Show answer & explanation

    Correct answers: A, B, C — Confirm the served entity references the model by alias (such as `@Champion`) rather than a version number hardcoded to whatever was true when the YAML was written.; Confirm `catalog_name` and `schema_name` on the `registered_models` resource resolve to the intended production namespace once the prod target's variables are applied.; Confirm `grants` on the production `registered_models` resource are scoped to the specific principals that need access, not granted broadly by default.

    • A. Checking that the served entity tracks an alias rather than a version frozen at authoring time confirms the endpoint will automatically follow future promotions instead of silently going stale.
    • B. Verifying that catalog_name and schema_name resolve correctly once target variables are applied catches a misconfigured or unresolved variable before it registers the model into the wrong namespace in production.
    • C. Verifying grants are scoped to only the principals that actually need access is exactly the least-privilege check a reviewer should make before a production model resource goes live.
    • D. A populated comment field has no effect on MLflow run logging; comment is purely descriptive metadata, so this is not a real risk worth checking for.
    • E. Traffic does not route itself automatically without configuration; traffic_config's routes are what define how requests are split across served entities, so omitting it entirely is not a valid or harmless state to leave unchecked.
    • F. Hardcoding entity_version to the current highest version freezes the endpoint on that version and requires a manual redeploy for every future promotion, which is the opposite of what a reviewer should confirm is in place.

    Subdomain 2.5: Drift Detection and Lakehouse Monitoring

    25.A monitor's label column points at `repaid_on_time`, populated only for loans over 60 days old. This week's performance metrics show a sharp accuracy drop for the newest fully-labeled cohort, while input drift metrics stay stable. What should the team check first?

    1. A.Whether the label's definition or collection process changed, since stable inputs point to the label side.
    2. B.Whether the Kolmogorov-Smirnov threshold was too strict, since a quiet drift table always needs loosening.
    3. C.Whether to switch from a labeled cohort window to snapshot profiling, required whenever metrics disagree.
    4. D.Whether to delete the baseline table, since a performance drop without drift proves it is invalid.
    Show answer & explanation

    Correct answer: A — Whether the label's definition or collection process changed, since stable inputs point to the label side.

    • A. With input features stable but real accuracy down, a change in how the label is defined, recorded, or delayed is a plausible cause worth checking first, since it need not show up as input drift.
    • B. A stable drift table does not mean the threshold is miscalibrated; loosening it until drift appears would manufacture a false signal rather than diagnose the real cause.
    • C. There is no rule forcing a switch to snapshot profiling when metrics disagree; inference profiling is exactly the type built to support labeled-cohort performance metrics.
    • D. Deleting the baseline is not a documented response to this mismatch and would remove the drift reference point without addressing the accuracy drop's cause.

    Subdomain 2.5: Drift Detection and Lakehouse Monitoring

    26.A team monitoring a forecasting model wants to compare accuracy between two customer segments receiving different pricing treatments, using the monitor's existing output. Which steps achieve this? (Select all that apply.)(Select 3)

    1. A.Configure a slicing expression on the segment column so profile and drift tables report per-segment rows.
    2. B.Ensure the label column is configured so joined ground truth produces segment-level metrics.
    3. C.Compare the resulting segment-level performance rows to spot a meaningful accuracy difference.
    4. D.Configure two separate baseline tables per segment, since slicing cannot apply to inference profiles.
    5. E.Disable the label column, since segment-level accuracy only needs raw predictions, never ground truth.
    Show answer & explanation

    Correct answers: A, B, C — Configure a slicing expression on the segment column so profile and drift tables report per-segment rows.; Ensure the label column is configured so joined ground truth produces segment-level metrics.; Compare the resulting segment-level performance rows to spot a meaningful accuracy difference.

    • A. Slicing by the segment column is the mechanism that produces separate profile and drift rows per segment for comparison.
    • B. The label column configuration is what lets ground truth be joined in, which is required to compute performance metrics at all, including per segment.
    • C. Comparing the segment-level performance rows against each other is exactly how the team surfaces whether accuracy differs meaningfully between the two treatment groups.
    • D. Slicing expressions apply to inference profile types just as they do to time series or snapshot profiles; separate baselines per segment are not required.
    • E. Performance metrics fundamentally require ground-truth labels; disabling the label column would prevent segment-level accuracy from being computed at all.

    Subdomain 2.5: Drift Detection and Lakehouse Monitoring

    27.A team is reviewing how the built-in drift tests compute statistical significance across two numerical columns and one categorical column, all compared against the baseline table. Which statements are correct? (Select all that apply.)(Select 3)

    1. A.Both numerical columns are evaluated with the Kolmogorov-Smirnov test against the baseline distribution.
    2. B.The categorical column is evaluated with the chi-squared test comparing category counts to baseline.
    3. C.A lower p-value on either test indicates a more statistically significant difference from the baseline.
    4. D.A high p-value on any test guarantees the column has zero practical difference from the baseline.
    5. E.The two tests can be swapped freely between numerical and categorical columns without changing results.
    Show answer & explanation

    Correct answers: A, B, C — Both numerical columns are evaluated with the Kolmogorov-Smirnov test against the baseline distribution.; The categorical column is evaluated with the chi-squared test comparing category counts to baseline.; A lower p-value on either test indicates a more statistically significant difference from the baseline.

    • A. Kolmogorov-Smirnov is the built-in test applied to numerical columns, so both numerical columns are evaluated with it against the baseline.
    • B. Chi-squared is the built-in test applied to categorical columns, comparing observed category counts to the baseline distribution.
    • C. For both tests, a lower p-value indicates the observed difference from baseline is less likely to be due to chance, meaning stronger statistical significance.
    • D. A high p-value only means the test failed to find statistically significant evidence of a difference; it does not guarantee there is zero practical difference, especially with a small sample.
    • E. Kolmogorov-Smirnov operates on continuous distributions and chi-squared operates on category counts, so they are not interchangeable between column types.

    Subdomain 2.2: Validation Testing

    28.A platform team is setting up dev, staging, and production workspaces for an ML project and deciding what kind of testing belongs in each. Which statement correctly matches a testing activity to the environment where it is described as primarily occurring?

    1. A.Development and production share one testing suite, while staging is used exclusively for manual business sign-off.
    2. B.Development supports exploratory work without strict access controls, while staging hosts the full testing suite before production.
    3. C.Production hosts the full testing suite so that any defects are caught using real production data before rollback is needed.
    4. D.Staging is reserved exclusively for manual exploratory testing, while automated unit and integration tests run solely in development.
    Show answer & explanation

    Correct answer: B — Development supports exploratory work without strict access controls, while staging hosts the full testing suite before production.

    • A. Development and production are described as distinct environments with different access controls and purposes, and staging's defining role is to run the full automated test suite, not just manual sign-off.
    • B. The development environment is where data scientists experiment with models and pipelines without strict access controls, while staging is designed to mirror production as closely as feasible and is where the complete testing suite runs before promotion.
    • C. Running the full testing suite in production would mean validating code against live production traffic without the staging gate that exists specifically to catch defects before they reach production.
    • D. Staging is where the automated unit and integration test suite is described as running, not development; development is oriented toward exploratory work, not the formal test suite.

    Domain 3: Model Deployment

    Subdomain 3.1: Deployment Strategies

    29.A team configuring a Model Serving endpoint for a new high-traffic use case is reviewing available compute and scaling settings. Which of the following are valid, documented configuration levers for the endpoint? (Select all that apply.)(Select 4)

    1. A.Selecting a workload type such as CPU, CPU_MEDIUM, CPU_LARGE, or a GPU tier based on the model's compute needs.
    2. B.Choosing a concurrency tier such as Small, Medium, or Large that bounds how many simultaneous requests the endpoint handles.
    3. C.Enabling or disabling scale-to-zero, trading idle-cost savings against cold-start latency on the first request after idling.
    4. D.Setting `min_provisioned_concurrency` and `max_provisioned_concurrency` in multiples of four to bound autoscaling range.
    5. E.Manually assigning each incoming request to a specific virtual machine by IP address to guarantee consistent routing.
    Show answer & explanation

    Correct answers: A, B, C, D — Selecting a workload type such as CPU, CPU_MEDIUM, CPU_LARGE, or a GPU tier based on the model's compute needs.; Choosing a concurrency tier such as Small, Medium, or Large that bounds how many simultaneous requests the endpoint handles.; Enabling or disabling scale-to-zero, trading idle-cost savings against cold-start latency on the first request after idling.; Setting `min_provisioned_concurrency` and `max_provisioned_concurrency` in multiples of four to bound autoscaling range.

    • A. Workload type selection, including CPU and GPU tiers, is a documented compute configuration option for a Model Serving endpoint.
    • B. Concurrency tiers are a documented setting that bounds how much simultaneous request traffic the endpoint can handle at once.
    • C. Scale-to-zero is a documented toggle that saves idle cost at the expense of cold-start latency on the next incoming request.
    • D. Provisioned concurrency bounds set in multiples of four are documented settings for controlling the endpoint's autoscaling range.
    • E. Model Serving does not expose per-request manual VM or IP assignment; scaling and routing are managed by the platform, not by hand-assigning machines.

    Subdomain 3.1: Deployment Strategies

    30.A platform team is defining guardrails for how deployment strategies should be chosen across Model Serving endpoints. Which of the following reflect sound practice based on the tradeoffs of canary and blue-green rollouts? (Select all that apply.)(Select 3)

    1. A.High-traffic, revenue-critical endpoints should default to a canary rollout so a regression is caught while affecting only a small fraction of requests.
    2. B.A low-traffic internal tool with lenient latency needs may reasonably use a simpler blue-green cutover, since the blast radius of a mistake is smaller.
    3. C.Every endpoint, regardless of traffic volume or business criticality, should always use the exact same fixed 50/50 traffic split for every rollout.
    4. D.Rollout decisions should account for how quickly a regression can be detected and how much traffic is exposed during that detection window.
    5. E.Because Model Serving is documented as highly available, rollout strategy is irrelevant and any version can be swapped in at 100% traffic with no monitoring.
    Show answer & explanation

    Correct answers: A, B, D — High-traffic, revenue-critical endpoints should default to a canary rollout so a regression is caught while affecting only a small fraction of requests.; A low-traffic internal tool with lenient latency needs may reasonably use a simpler blue-green cutover, since the blast radius of a mistake is smaller.; Rollout decisions should account for how quickly a regression can be detected and how much traffic is exposed during that detection window.

    • A. Defaulting critical, high-traffic endpoints to a gradual canary rollout limits how much traffic is exposed to a bad new version before it is caught.
    • B. A lower-stakes internal tool can reasonably accept the simplicity of a full blue-green swap because the cost of a mistake affecting all its traffic is lower.
    • C. A single fixed split ignores that different endpoints carry different risk profiles; sound guardrails size the rollout to the endpoint's criticality and traffic instead.
    • D. Detection speed and exposure during the detection window are exactly the factors that should drive whether a slower canary ramp or a faster cutover is appropriate.
    • E. High availability describes serving infrastructure uptime, not whether an untested model version is safe to expose to all traffic at once; staged rollout remains necessary.

    Subdomain 3.1: Deployment Strategies

    31.Which of the following are true, documented characteristics of Model Serving that a team should factor into a deployment-strategy decision for a high-traffic production model? (Select all that apply.)(Select 3)

    1. A.Model Serving is designed for high-availability, low-latency production use and can support a very large volume of queries per second with low overhead.
    2. B.Model Serving can deploy custom Python models, Databricks-hosted foundation models, and certain external models through one unified interface.
    3. C.Model Serving requires every served model to be scored exclusively through a Spark cluster that stays attached to the calling notebook session.
    4. D.Model Serving exposes each deployed model as a documented REST API that any application running outside Databricks can integrate against.
    5. E.Model Serving guarantees that scale-to-zero endpoints never incur any added cold-start latency whatsoever on the very first request after idling.
    Show answer & explanation

    Correct answers: A, B, D — Model Serving is designed for high-availability, low-latency production use and can support a very large volume of queries per second with low overhead.; Model Serving can deploy custom Python models, Databricks-hosted foundation models, and certain external models through one unified interface.; Model Serving exposes each deployed model as a documented REST API that any application running outside Databricks can integrate against.

    • A. High availability, low latency, and a documented high-throughput ceiling are core, accurate characteristics of the platform relevant to a high-traffic decision.
    • B. Supporting custom models, foundation models, and external models behind one interface is a documented, accurate capability of Model Serving.
    • C. Model Serving does not require a caller-attached Spark cluster to score a model; it exposes a serverless REST endpoint instead of that.
    • D. Exposing each model as a REST API that external applications can call is a core, documented feature of Model Serving worth relying on.
    • E. Scale-to-zero endpoints do incur cold-start latency on the first request after idling, so a claim of a zero-latency guarantee is inaccurate.

    Subdomain 3.1: Deployment Strategies

    32.Reviewing sample interview notes from an ML platform candidate, an interviewer flags a few statements as technically wrong about deploying custom models on Model Serving. Which of the following statements are INCORRECT? (Select all that apply.)(Select 3)

    1. A."You must use `mlflow.sklearn.log_model` for every model type; `pyfunc` models cannot be deployed to Model Serving at all."
    2. B."A `pyfunc` model's `predict` method can call arbitrary Python logic, including custom preprocessing, before returning a prediction."
    3. C."Unity Catalog will register any model regardless of whether it has a signature, so signatures are optional for serving deployment."
    4. D."Loading a large reference file once at container startup belongs in `load_context`, not inside `predict` on every call."
    5. E."Registering a model to Unity Catalog automatically creates and starts a serving endpoint for that registered model."
    Show answer & explanation

    Correct answers: A, C, E — "You must use `mlflow.sklearn.log_model` for every model type; `pyfunc` models cannot be deployed to Model Serving at all."; "Unity Catalog will register any model regardless of whether it has a signature, so signatures are optional for serving deployment."; "Registering a model to Unity Catalog automatically creates and starts a serving endpoint for that registered model."

    • A. This is incorrect: `pyfunc` models built by subclassing `PythonModel` are a documented, standard way to deploy custom logic to Model Serving, not something excluded.
    • B. This statement is accurate: a `pyfunc` model's `predict` method is exactly where arbitrary custom preprocessing and inference logic can live and run.
    • C. This is incorrect: Unity Catalog requires a valid signature before accepting a model registration, so signatures are not optional there at all.
    • D. This statement is accurate: `load_context` is the documented place for one-time startup work like loading a large reference file once.
    • E. This is incorrect: registering a model to Unity Catalog does not itself create a serving endpoint; endpoint creation remains a separate, explicit step.

    Subdomain 3.2: Custom Model Serving

    33.A data scientist wants `mlflow.register_model()` to place a newly logged custom model into a Unity Catalog three-level namespace instead of the legacy workspace Model Registry. What must they configure before calling `register_model`?

    1. A.Call `mlflow.set_registry_uri("databricks-uc")` and pass a model name formatted as `catalog.schema.model_name`.
    2. B.Call `mlflow.set_tracking_uri("databricks-uc")` and pass a model name formatted as `schema.model_name.version`.
    3. C.Set the `MLFLOW_REGISTRY_MODE` environment variable to `unity_catalog` and pass the model's run ID as the model name.
    4. D.Call `mlflow.set_experiment("unity_catalog")` before logging the model so the run is associated with Unity Catalog by default.
    Show answer & explanation

    Correct answer: A — Call `mlflow.set_registry_uri("databricks-uc")` and pass a model name formatted as `catalog.schema.model_name`.

    • A. Setting the registry URI to `databricks-uc` points `register_model` at Unity Catalog, and Unity Catalog model names must follow the three-level `catalog.schema.model_name` convention rather than the flat names used by the workspace registry.
    • B. `set_tracking_uri` controls where experiment runs and metrics are logged, not where models are registered, and it does not accept `databricks-uc` as a valid tracking destination.
    • C. There is no `MLFLOW_REGISTRY_MODE` environment variable in MLflow, and a run ID is not a valid model name for registration in either registry.
    • D. `set_experiment` only selects which experiment new runs are logged under; it has no effect on which model registry a subsequent `register_model` call targets.

    Subdomain 3.2: Custom Model Serving

    34.Before a custom pyfunc model can be registered under a Unity Catalog three-level name such as `ml_prod.fraud.scoring_model`, which of the following conditions must already be true? (Select all that apply.)(Select 3)

    1. A.The `ml_prod` catalog and `fraud` schema already exist in Unity Catalog, or the caller has permission to create them.
    2. B.The caller has `CREATE MODEL` or equivalent privilege on the target schema so the registration operation is authorized.
    3. C.The MLflow client's registry URI is set to `databricks-uc` so `register_model` targets Unity Catalog rather than the workspace registry.
    4. D.The model was trained using a Unity Catalog-managed Feature Store table, since Unity Catalog rejects models trained on any other data source.
    5. E.The model's Python class was written using only Databricks Runtime ML's built-in algorithms, since Unity Catalog rejects custom pyfunc classes.
    Show answer & explanation

    Correct answers: A, B, C — The `ml_prod` catalog and `fraud` schema already exist in Unity Catalog, or the caller has permission to create them.; The caller has `CREATE MODEL` or equivalent privilege on the target schema so the registration operation is authorized.; The MLflow client's registry URI is set to `databricks-uc` so `register_model` targets Unity Catalog rather than the workspace registry.

    • A. The catalog and schema in the three-level name must exist, or the registering identity needs permission to create them, before a model can be registered under that namespace.
    • B. Unity Catalog enforces standard privilege checks, so the caller needs the appropriate model-creation privilege on the target schema for the registration call to succeed.
    • C. Pointing the MLflow registry URI at `databricks-uc` is required so that `register_model` resolves the three-level name against Unity Catalog instead of the legacy workspace registry.
    • D. Unity Catalog model registration has no requirement that training data originated from a Feature Store table; models trained on any data source can be registered as long as the model itself is properly logged.
    • E. Unity Catalog fully supports custom pyfunc models built with arbitrary Python logic; it does not restrict registration to built-in Databricks Runtime ML algorithms.

    Subdomain 3.2: Custom Model Serving

    35.A team wants to update the served model version on a live custom model endpoint without causing a visible outage for callers currently sending requests. Which of the following are valid techniques for achieving this? (Select all that apply.)(Select 3)

    1. A.Add the new model version as a second served entity on the same endpoint and initially route a small percentage of traffic to it.
    2. B.Update the endpoint configuration via the REST API or the SDK to adjust `traffic_percentage` between the served entities over time.
    3. C.Monitor request latency and error rate for the new served entity closely before increasing its share of the traffic further.
    4. D.Delete the endpoint entirely and recreate it again from scratch every time a new model version needs to be served.
    5. E.Ask every calling application to pause its own retries manually while an engineer swaps the model file on disk.
    Show answer & explanation

    Correct answers: A, B, C — Add the new model version as a second served entity on the same endpoint and initially route a small percentage of traffic to it.; Update the endpoint configuration via the REST API or the SDK to adjust `traffic_percentage` between the served entities over time.; Monitor request latency and error rate for the new served entity closely before increasing its share of the traffic further.

    • A. Adding the new version as an additional served entity with a small initial traffic share is the documented mechanism for introducing a new model version without an abrupt cutover.
    • B. Adjusting `traffic_percentage` across served entities through the REST API or SDK is how the gradual shift from an old to a new version is actually performed.
    • C. Watching latency and error metrics for the new entity before increasing its share lets the team detect a regression before most traffic depends on the new version.
    • D. Deleting and recreating the entire endpoint causes exactly the outage the scenario is trying to avoid, and discards the multi-entity, gradual-rollout capability the endpoint already provides.
    • E. Coordinating manual pauses across every calling application is not a Model Serving mechanism at all, and served model files are not swapped on disk outside of MLflow's deployment workflow.

    Want the full experience?

    These are just samples. Practice the full Databricks Certified Machine Learning Professional question bank in quiz mode — free, no signup, with domain practice and exam simulation.