CertSafari

    Free Databricks Certified Machine Learning Associate Sample Questions

    35 free sample questions from our bank of 330+, covering every exam domain, with answers and detailed explanations. Updated September 2026.

    Domain 1: Databricks Machine Learning

    Subdomain 1.2: Identify the advantages of using ML runtimes

    1.A team maintains a model training pipeline that must produce identical results when re-run six months later, including the exact versions of scikit-learn, XGBoost, and their transitive dependencies used at training time. Which characteristic of Databricks Runtime for Machine Learning most directly supports this requirement?

    1. A.Databricks Runtime for Machine Learning automatically upgrades scikit-learn and XGBoost to their latest release on every cluster restart, keeping the environment current with upstream fixes.
    2. B.Each Databricks Runtime for Machine Learning release pins a tested, documented set of library versions, so re-running a job on the same runtime version reproduces the same dependency set.
    3. C.Databricks Runtime for Machine Learning stores every notebook's output as a Delta table, which independently guarantees identical library versions across separate training runs.
    4. D.Databricks Runtime for Machine Learning removes the need to track library versions at all, because MLflow autologging reconstructs the original environment from the model artifact alone.
    Show answer & explanation

    Correct answer: B — Each Databricks Runtime for Machine Learning release pins a tested, documented set of library versions, so re-running a job on the same runtime version reproduces the same dependency set.

    • A. Databricks Runtime for Machine Learning does not silently auto-upgrade libraries on restart; a given runtime release keeps a fixed library set, which is what makes it reproducible rather than a moving target.
    • B. Each Databricks Runtime for Machine Learning release ships a specific, tested, and documented set of library versions, so pinning a job to that runtime version reproduces the same scikit-learn, XGBoost, and dependency versions on a later run.
    • C. Saving notebook output to Delta tables preserves data, not the software environment; it does not record or guarantee which library versions produced that output on a separate run.
    • D. MLflow autologging can record library versions as part of a model's metadata, but it does not eliminate the need to track versions during training or reconstruct an environment on its own without a pinned runtime to install against.

    Subdomain 1.4: Identify the advantages AutoML brings to the model development process

    2.A data science manager wants to compare the parameters, metrics, and artifacts of every model AutoML tried for a pricing-prediction problem, including runs that were not selected as the best model. Which Databricks AutoML behavior makes this possible?

    1. A.AutoML logs every trial as an MLflow run, capturing parameters, metrics, and artifacts so the manager can compare all trials, not only the best one, in the MLflow UI.
    2. B.AutoML keeps detailed records only for the single best-performing trial and permanently discards the parameters and metrics from every other trial once training finishes.
    3. C.AutoML stores trial results in a local CSV file on the cluster's driver node, which is deleted automatically when the cluster terminates after the run.
    4. D.AutoML emails a static summary report of the best trial to the workspace administrator, without persisting any run data that could be reopened or compared.
    Show answer & explanation

    Correct answer: A — AutoML logs every trial as an MLflow run, capturing parameters, metrics, and artifacts so the manager can compare all trials, not only the best one, in the MLflow UI.

    • A. Every AutoML trial is tracked as its own MLflow run with parameters, metrics, and artifacts, so the manager can open the MLflow experiment and compare all trials side by side, not just the one AutoML selected as best.
    • B. AutoML does not discard the non-winning trials; each one remains logged as an MLflow run with its own parameters and metrics available for later comparison.
    • C. Trial results are not stashed in an ephemeral local CSV on the driver — they are written to MLflow, which persists independently of the cluster's lifecycle.
    • D. AutoML does not communicate results only through a one-time emailed summary; the run data for every trial stays queryable and reopenable in MLflow after the job finishes.

    Subdomain 1.3: Identify how AutoML facilitates model/feature selection.

    3.Which open-source libraries does Databricks AutoML draw its trial algorithms from when running a classification experiment?

    1. A.TensorFlow and PyTorch implementations of deep learning classifiers tuned automatically by AutoML runs.
    2. B.Spark MLlib's distributed classification implementations, run natively across the cluster's worker nodes.
    3. C.scikit-learn, XGBoost, and LightGBM implementations of classification algorithms tried during trials.
    4. D.Prophet and ARIMA implementations, the same forecasting libraries AutoML uses for time series experiments.
    Show answer & explanation

    Correct answer: C — scikit-learn, XGBoost, and LightGBM implementations of classification algorithms tried during trials.

    • A. AutoML's classification trials are not built on TensorFlow or PyTorch deep learning models; its algorithm pool for classification and regression comes from classical machine learning libraries instead.
    • B. AutoML does not train its trials using Spark MLlib's distributed algorithm implementations; it distributes trials of single-node scikit-learn, XGBoost, and LightGBM models across the cluster instead.
    • C. Databricks AutoML evaluates classification and regression trials using open source algorithms from scikit-learn, XGBoost, and LightGBM, training and comparing multiple model types from these libraries before ranking the trials.
    • D. Prophet and ARIMA are the libraries AutoML uses specifically for forecasting experiments, not classification; classification trials draw from a different set of algorithm libraries.

    Subdomain 1.1: Identify the best practices of an MLOps strategy

    4.A team runs identical training pipeline code in dev, staging, and production Databricks workspaces, and retraining is cheap because the dataset is small. Following Databricks' recommended MLOps strategy, how should a validated model move into production?

    1. A.Promote the pipeline code into the production workspace so the production job retrains on production data and registers its own fresh model version.
    2. B.Copy the serialized model file produced in staging directly into the production catalog, skipping retraining so the production version matches staging exactly.
    3. C.Export the staging MLflow run's artifact and re-upload it as a new run in production, keeping the original training code only in the staging workspace.
    4. D.Point the production serving endpoint at the staging workspace's tracking server so production always serves whatever model staging last trained.
    Show answer & explanation

    Correct answer: A — Promote the pipeline code into the production workspace so the production job retrains on production data and registers its own fresh model version.

    • A. This matches Databricks' code-first promotion guidance: promoting the pipeline through environments and retraining in production keeps the deployed model consistent with production data and infrastructure rather than a possibly stale artifact.
    • B. Copying a serialized artifact between environments bypasses retraining on production data and skips the governance and reproducibility benefits of running the pipeline where it will serve, so it is not the recommended approach when retraining is cheap.
    • C. Re-uploading an artifact as a new run without rerunning the training code in production leaves the production environment without a reproducible pipeline, undermining lineage and future retraining.
    • D. Having production depend on staging's tracking server couples environments together and removes the isolation MLOps practices rely on, since a change in staging would silently affect production.

    Subdomain 1.5: Identify the benefits of creating feature store tables at the account level in Unity Catalog in Databricks vs at the workspace level

    5.A data science team in Workspace A builds a demand-forecasting feature table for a retail company. The marketing analytics team, working in Workspace B, is attached to the same Unity Catalog metastore and wants to reuse those exact features for a promotion-targeting model instead of recomputing them. The features were originally created as an account-level table in Unity Catalog rather than in a workspace-level feature store. What is the primary benefit this gives the marketing team?

    1. A.Because the feature table lives in Unity Catalog, it is accessible to every workspace attached to the same metastore, so the marketing team can read and reuse it through a normal `catalog.schema.table` reference without copying data.
    2. B.Because the feature table is stored in a workspace-level Hive metastore, Databricks grants read access to any user across the entire account regardless of which metastore their workspace is attached to.
    3. C.Because the feature table was registered with the legacy `FeatureStoreClient`, MLflow automatically forwards its schema to every workspace's local model registry so downstream teams inherit access.
    4. D.Because the feature table was written with Auto Loader, Databricks automatically replicates the underlying files into every workspace's own DBFS mount, letting each team keep an independent physical copy of the data rather than one shared source.
    Show answer & explanation

    Correct answer: A — Because the feature table lives in Unity Catalog, it is accessible to every workspace attached to the same metastore, so the marketing team can read and reuse it through a normal `catalog.schema.table` reference without copying data.

    • A. Unity Catalog objects are scoped to the metastore, not to a single workspace, so any workspace attached to that metastore can query the table directly by its three-level name. This is exactly why account-level feature tables enable reuse across teams without duplicating pipelines.
    • B. A workspace-level Hive metastore is scoped to a single workspace by definition, so it cannot grant account-wide access — this is the opposite of how workspace-level feature stores behave. It also inverts the premise of the question, which already states the table is account-level.
    • C. The legacy `FeatureStoreClient` is workspace-scoped and does not propagate schemas into other workspaces' model registries; registry entries are not how table read access is granted. This mixes up model registry replication with the actual mechanism, metastore-level table access.
    • D. Auto Loader is an ingestion tool for incrementally loading files into a table; it has no role in replicating a Unity Catalog table's files into other workspaces' DBFS mounts. This describes a mechanism that does not exist and misattributes cross-workspace access to the wrong feature.

    Subdomain 1.7: Write data to a feature store table

    6.A data scientist runs `fe.write_table(name="loan_features", df=updates_df, mode="merge")` against a feature table whose primary key is the composite `(loan_id, as_of_date)`, and the call fails during validation. `updates_df` contains `loan_id`, `credit_score`, and `utilization_ratio`, but no `as_of_date` column. What is the most likely cause of the failure?

    1. A.Merge mode is only supported for tables with a single-column primary key, so a composite-key table like `loan_features` must always use `mode="overwrite"` instead.
    2. B.The DataFrame's column order does not match the feature table's stored column order, and `write_table` requires primary key columns to appear first in the schema.
    3. C.The feature table has `publish_table` online replication enabled, and any offline write is blocked until the online store finishes syncing the previous batch.
    4. D.The DataFrame is missing one of the feature table's primary key columns, and `write_table` requires every primary key column to be present so it can determine which rows to update or insert.
    Show answer & explanation

    Correct answer: D — The DataFrame is missing one of the feature table's primary key columns, and `write_table` requires every primary key column to be present so it can determine which rows to update or insert.

    • A. This claim about merge mode is invented: merge mode works the same way regardless of whether the primary key is a single column or a composite key, so restricting composite-key tables to overwrite mode is not an actual constraint of the API.
    • B. Column order in the incoming DataFrame does not matter to `write_table`; it matches columns by name against the feature table's schema, so this is not the reason the call fails.
    • C. Publishing to an online store does not place a lock on offline writes to the feature table, so a prior online sync would not block this `write_table` call.
    • D. Because the feature table's primary key is `(loan_id, as_of_date)`, every write must supply both columns so the client can identify which existing rows to update and which new rows to insert; omitting `as_of_date` leaves the operation unable to resolve the primary key, which is why validation fails.

    Subdomain 1.6: Create a feature store table in Unity Catalog

    7.A team is building a feature table for fraud detection where feature values change throughout the day and the training set must only use values that were known at each transaction's timestamp, to avoid leaking future information. Which approach correctly creates this table with `FeatureEngineeringClient.create_table`?

    1. A.Include the timestamp column in `primary_keys` with the entity key, and pass that column to `timeseries_columns` so lookups respect point-in-time correctness.
    2. B.Pass the timestamp column only to the `schema` argument as a plain feature column, letting `create_training_set` infer time ordering from row insertion order automatically.
    3. C.Create the table without a timestamp column, and rely on `fe.write_table` in `merge` mode to always overwrite older feature values before each training run.
    4. D.Add the timestamp column as a `tags` entry on the table so downstream jobs know which column represents event time during scoring and inference.
    Show answer & explanation

    Correct answer: A — Include the timestamp column in `primary_keys` with the entity key, and pass that column to `timeseries_columns` so lookups respect point-in-time correctness.

    • A. Correct. Time series feature tables require the time column to be part of the primary key and declared through `timeseries_columns`, which is what enables point-in-time lookups during training-set creation.
    • B. Incorrect. Treating the timestamp as a plain schema column gives it no special meaning; `create_training_set` has no mechanism to infer time ordering from insertion order alone.
    • C. Incorrect. Omitting the timestamp key removes any way to look up feature values as of a specific point in time, and `merge` mode write behavior does not substitute for that.
    • D. Incorrect. `tags` store descriptive key-value metadata for humans and catalog search; they are not read by the training-set or lookup logic that enforces point-in-time correctness.

    Subdomain 1.11: Identify the best run using the MLflow Client API.

    8.What is the primary purpose of using the MLflow Client API's `search_runs` method with `order_by` and `max_results` parameters, compared to browsing the MLflow Tracking UI's runs table?

    1. A.It lets code retrieve a ranked, limited set of runs by metric value, enabling automated best-run selection without manual inspection in the UI.
    2. B.It permanently deletes every run that does not rank within the requested `max_results` window, keeping only the best-performing runs in the experiment.
    3. C.It replaces the MLflow Tracking UI entirely, since the UI cannot display metrics, parameters, or artifacts once the client API has been used on an experiment.
    4. D.It uploads the ranked run results into Unity Catalog as a new registered model version, replacing the separate manual model registration step needed otherwise.
    Show answer & explanation

    Correct answer: A — It lets code retrieve a ranked, limited set of runs by metric value, enabling automated best-run selection without manual inspection in the UI.

    • A. `search_runs` with `order_by` and `max_results` is built to filter and rank runs programmatically, which is exactly what automated pipelines need to select a best run without a human browsing the UI.
    • B. `search_runs` only queries and returns run metadata; it never deletes runs that happen to fall outside the requested result window.
    • C. The Tracking UI keeps displaying metrics, parameters, and artifacts for an experiment regardless of whether the client API has also been used to query the same runs.
    • D. Retrieving the best run through `search_runs` does not register anything; registering a model into Unity Catalog is a separate, explicit call to `mlflow.register_model`.

    Subdomain 1.14: Register a model using the MLflow Client API in the Unity Catalog registry

    9.A workspace administrator with full workspace admin rights repeatedly gets a permission-denied error when running `mlflow.register_model(model_uri, "ops.ml.models.fraud_detector")` against the `ops.ml.models` schema. What privilege must be granted for the registration to succeed?

    1. A.Grant `CREATE MODEL` on the `ops.ml.models` schema in addition to `USE CATALOG` and `USE SCHEMA`, since workspace admin rights do not carry Unity Catalog object privileges.
    2. B.Add the user as a metastore admin in the Unity Catalog metastore, since only metastore-level roles can create new registered models in any schema.
    3. C.Grant `CAN_MANAGE` on the cluster running the notebook, since compute-level permissions control whether registry write calls are allowed to execute.
    4. D.Grant `MODIFY` on the `ops` catalog only, since catalog-level `MODIFY` privilege is sufficient to create objects in every schema it contains.
    Show answer & explanation

    Correct answer: A — Grant `CREATE MODEL` on the `ops.ml.models` schema in addition to `USE CATALOG` and `USE SCHEMA`, since workspace admin rights do not carry Unity Catalog object privileges.

    • A. Unity Catalog privileges are separate from workspace admin rights, so registering a model requires `USE CATALOG` and `USE SCHEMA` to traverse the namespace plus `CREATE MODEL` on the schema to create the object itself. Missing any of these three produces a permission-denied error regardless of workspace-level roles.
    • B. Metastore admin is a broad administrative role for managing the metastore itself; it is not the minimal privilege intended for day-to-day model registration and is not required just to create a registered model in a schema the user should already have scoped access to.
    • C. Cluster compute permissions govern who can attach to or manage a cluster; they do not grant or restrict access to Unity Catalog securable objects like schemas or registered models.
    • D. Catalog-level `MODIFY` alone does not substitute for the schema-level `CREATE MODEL` privilege needed to create a new registered model object inside a specific schema; the schema-level grant is still required.

    Subdomain 1.12: Manually log metrics, artifacts, and models in an MLflow Run.

    10.A data scientist is training a scikit-learn classifier over 20 epochs inside a `with mlflow.start_run():` block on Databricks and wants the MLflow UI to plot how validation accuracy changes across epochs, not just show a single final number. Which manual logging call, placed inside the training loop, achieves this?

    1. A.Call `mlflow.log_metric("val_accuracy", value, step=epoch)` once per epoch, passing the epoch number as `step` so the metric history builds one point per epoch.
    2. B.Call `mlflow.log_metrics({"val_accuracy": value})` once per epoch without a `step` argument, since batching a dictionary automatically tracks call order as the chart's x-axis.
    3. C.Call `mlflow.log_param("val_accuracy", value)` at the end of every epoch, since parameters and metrics are both rendered as line charts in the run's UI page.
    4. D.Call `mlflow.log_artifact("accuracy_history.json")` once after training completes, uploading a JSON file of every epoch's value so the UI renders it as a metrics chart automatically.
    Show answer & explanation

    Correct answer: A — Call `mlflow.log_metric("val_accuracy", value, step=epoch)` once per epoch, passing the epoch number as `step` so the metric history builds one point per epoch.

    • A. This is correct because passing an explicit `step` value on each call is exactly how MLflow builds a per-step metric history for a key, which the UI then renders as a trend line across epochs.
    • B. This is incorrect because omitting `step` logs every call at the default step of 0, so repeated calls create multiple points stacked at the same step rather than a clean per-epoch trend.
    • C. This is incorrect because parameters are meant for static configuration values, are not designed to accept repeated updates to the same key across a run, and are not charted as a time series in the UI.
    • D. This is incorrect because uploaded files are stored as opaque artifacts; MLflow does not parse artifact contents and turn them into interactive metric charts.

    Subdomain 1.13: Identify information available in the MLFlow UI

    11.An experiment has 200 runs logged over several weeks of model iteration. A data scientist wants to view, directly in the MLflow UI run table, only the runs where `params.model_type` is `"random_forest"` and the logged `metrics.f1_score` is above 0.85. What should they do?

    1. A.Type a filter expression such as `params.model_type = "random_forest" and metrics.f1_score > 0.85` into the run table's search box, which MLflow parses and applies directly against the logged parameters and metrics.
    2. B.Click each column header to sort the run table alphabetically by parameter value, then manually scroll through all 200 rows to identify which runs meet both the model type and score conditions.
    3. C.Export the entire experiment to a CSV file from the Artifacts tab and open it in a spreadsheet application, since the run table's own search box is believed to have no way to combine a parameter and a metric condition.
    4. D.Open the Compare Runs view and drag a Parallel Coordinates Plot threshold line at `f1_score = 0.85`, which highlights matching lines but does not filter by the `model_type` parameter.
    Show answer & explanation

    Correct answer: A — Type a filter expression such as `params.model_type = "random_forest" and metrics.f1_score > 0.85` into the run table's search box, which MLflow parses and applies directly against the logged parameters and metrics.

    • A. The MLflow UI's search box accepts a SQL-like filter syntax referencing `params.<name>` and `metrics.<name>` combined with comparison operators and `and`, so this exact expression narrows the run table to runs matching both conditions at once.
    • B. Column-header sorting only reorders visible rows by one field at a time; it provides no way to jointly satisfy a parameter equality condition and a metric threshold without manually scanning every row.
    • C. The run table supports combined parameter and metric filtering natively through its search box, so exporting to a spreadsheet to combine two conditions is unnecessary and the Artifacts tab does not offer a full experiment CSV export in this way.
    • D. The Parallel Coordinates Plot can visually emphasize a metric range along one axis, but it does not apply a `model_type` equality condition, so it would not isolate runs matching both criteria.

    Subdomain 1.17: Set or remove a tag for a model

    12.A registered model named `pricing_model` carries a `status=deprecated` tag from an earlier retirement decision that has since been reversed. The team wants that tag gone entirely, not just cleared to an empty value, and the tag lives at the registered-model level, not on any single version. Which call removes it correctly?

    1. A.`client.delete_registered_model_tag(name="pricing_model", key="status")` to remove the key from the registered model.
    2. B.`client.set_registered_model_tag(name="pricing_model", key="status", value="")` to overwrite the value with an empty string.
    3. C.`client.delete_model_version_tag(name="pricing_model", version=1, key="status")` to remove the key from the first model version.
    4. D.`client.transition_model_version_stage(name="pricing_model", version=1, stage="Production")` to move the version out of the deprecated stage.
    Show answer & explanation

    Correct answer: A — `client.delete_registered_model_tag(name="pricing_model", key="status")` to remove the key from the registered model.

    • A. `delete_registered_model_tag` removes the key entirely from the registered model's tag set, which matches the requirement to drop the `status` key rather than merely blank out its value.
    • B. Setting the value to an empty string still leaves the `status` key present with an empty value, so it would still show up in tag listings and tag-based searches instead of disappearing.
    • C. This targets a tag on a specific model version rather than the registered model, so it would not touch the model-level `status` tag the team wants removed.
    • D. Unity Catalog's Model Registry does not use the legacy workspace-registry stage concept, and even if a stage transition existed here, changing a version's stage would not delete a tag on the registered model.

    Domain 2: Data Processing

    Subdomain 2.1: Compute summary statistics on a Spark DataFrame using .summary() or dbutils data summaries

    13.A data scientist samples a Spark DataFrame down to a small pandas DataFrame inside a notebook and wants a rich, browsable report with per-column data types, missing-value counts, and a small histogram for each numeric column, rendered inline. Which call is designed for this?

    1. A.`dbutils.data.summarize(pandas_df)`, because the function accepts pandas DataFrames as well as Spark DataFrames and renders an interactive report with types, null counts, and histograms per column.
    2. B.`pandas_df.describe()`, because pandas' built-in summary method renders the same interactive report, including histograms and missing-value counts, directly inside the notebook cell.
    3. C.`spark.createDataFrame(pandas_df).summary()`, because converting back to a Spark DataFrame is required to get histogram visualizations that pandas objects cannot render on their own.
    4. D.`dbutils.data.summarize(pandas_df, precise=True)`, because setting `precise` to `True` is what enables the histogram and missing-value visualizations, which are otherwise left out of the report entirely.
    Show answer & explanation

    Correct answer: A — `dbutils.data.summarize(pandas_df)`, because the function accepts pandas DataFrames as well as Spark DataFrames and renders an interactive report with types, null counts, and histograms per column.

    • A. `dbutils.data.summarize()` works on both pandas and Spark DataFrames and produces an interactive, browsable notebook report with per-column type, null-count, and histogram information, matching exactly what was asked for.
    • B. pandas' `.describe()` returns a plain numeric table of count, mean, stddev, and quartiles with no histograms, type breakdown, or interactive rendering, so it does not meet the visualization requirement.
    • C. `summary()` on a Spark DataFrame returns a plain text-based statistics table, not an interactive report with histograms, so converting the pandas object back to Spark first does not add the requested visualizations.
    • D. The `precise` flag controls whether statistics are computed exactly or approximately; it does not gate whether the histogram and visualization elements of the report are shown, which appear regardless of this setting.

    Subdomain 2.2: Remove outliers from a Spark DataFrame based on standard deviation or IQR

    14.In the standard IQR-based outlier rule, bounds are set as `Q1 - k*IQR` and `Q3 + k*IQR`. What is the effect of raising `k` from `1.5` to `3` on which values get flagged as outliers?

    1. A.Raising `k` to `3` widens the fences away from the quartiles, so fewer, more extreme values beyond that boundary get flagged, matching 'extreme' outliers versus the broader set at `1.5`.
    2. B.Raising `k` to `3` narrows the fences closer to the median, so more values near the center of the distribution get flagged, which increases the total count of rows removed as outliers on the same data.
    3. C.Raising `k` to `3` shifts the fences from being based on the interquartile range to being based on the full range of the data, so the dataset's minimum and maximum values always become the new fence boundary.
    4. D.Raising `k` to `3` has no effect on which values get flagged, since the multiplier only changes the reported IQR statistic itself and does not feed into the upper or lower fence calculation at all.
    Show answer & explanation

    Correct answer: A — Raising `k` to `3` widens the fences away from the quartiles, so fewer, more extreme values beyond that boundary get flagged, matching 'extreme' outliers versus the broader set at `1.5`.

    • A. A larger multiplier pushes both fences further away from the quartiles, so fewer, more extreme points fall outside them; `k=1.5` is the common threshold for moderate outliers and `k=3` for extreme ones.
    • B. Increasing the multiplier widens the fences rather than narrowing them, since `k*IQR` grows as `k` grows, so more values fall inside the boundary and fewer get flagged, not more.
    • C. The fence formula continues to use `Q1`, `Q3`, and the IQR regardless of the value of `k`; changing `k` never switches the calculation to use the dataset's minimum or maximum instead.
    • D. The multiplier `k` is multiplied directly against the IQR inside the fence formula, so changing it from `1.5` to `3` directly changes where the upper and lower bounds fall.

    Subdomain 2.4: Compare two categorical or two continuous features using the appropriate method

    15.Which method is the standard approach for measuring the strength and direction of the linear relationship between two continuous features?

    1. A.Pearson correlation coefficient
    2. B.Chi-square test of independence
    3. C.One-way ANOVA
    4. D.One-hot encoding
    Show answer & explanation

    Correct answer: A — Pearson correlation coefficient

    • A. The Pearson correlation coefficient is the standard statistic for measuring the strength and direction of a linear relationship between two continuous features.
    • B. This test compares observed versus expected frequency counts between two categorical features and does not apply to continuous numeric features.
    • C. ANOVA compares the mean of a continuous variable across groups defined by a categorical variable; it does not measure the relationship between two continuous features.
    • D. This technique converts categorical labels into binary indicator columns and is unrelated to measuring the relationship between two continuous features.

    Subdomain 2.3: Create visualizations for categorical or continuous features

    16.An analyst suspects the continuous `delivery_time_hours` column has a handful of unusually large values that could be outliers, and wants a single plot that shows the median, the interquartile range, and any points that fall far outside that range. Which visualization satisfies this?

    1. A.A box plot of `delivery_time_hours`, drawing the median line, the interquartile box, and points beyond the whiskers as outliers.
    2. B.A histogram of `delivery_time_hours`, binning the values and counting rows per bin without marking a median or flagging outliers.
    3. C.A bar chart of `delivery_time_hours`, showing one bar per unique delivery time and computing no quartiles or median at all.
    4. D.A pie chart of `delivery_time_hours`, dividing the total delivery hours across every row into proportional wedges only.
    Show answer & explanation

    Correct answer: A — A box plot of `delivery_time_hours`, drawing the median line, the interquartile box, and points beyond the whiskers as outliers.

    • A. A box plot is built specifically from the median and the interquartile range, and it plots any value beyond 1.5 times the IQR from the box as an individual outlier point, matching every requirement in the scenario.
    • B. A histogram summarizes the count of values per bin, which shows overall shape but does not compute or draw a median line or explicitly flag individual outlier points.
    • C. A bar chart of a continuous column like delivery time treats each raw value as a separate category and provides no summary statistic such as a median or interquartile range.
    • D. A pie chart shows proportional shares of a total and has no mechanism for representing quartiles, a median, or outlier points along a numeric range.

    Subdomain 2.5: Compare and contrast imputing missing values with the mean or median or mode value

    17.A data scientist is deciding between mean and median imputation for a household `annual_income` feature that is right-skewed by a small number of very high earners. A colleague argues mean imputation is fine because it will not change the overall column average. Which response correctly evaluates that argument?

    1. A.The argument is flawed, because keeping the column average unchanged is not the same as filling in a typical, realistic income for most missing households.
    2. B.The argument is correct, because preserving the exact column mean after imputation is the primary goal of any imputation strategy, regardless of skew.
    3. C.The argument is irrelevant, because annual income should always be imputed with the mode instead of either the mean or the median for any dataset.
    4. D.The argument is correct, because a right-skewed distribution has no meaningful median at all, so the mean is presented as the only statistic that can be computed reliably here.
    Show answer & explanation

    Correct answer: A — The argument is flawed, because keeping the column average unchanged is not the same as filling in a typical, realistic income for most missing households.

    • A. Mean imputation trivially preserves the column average by construction, but the goal of imputation is a plausible value per household, and the mean is skewed high by the top earners, misrepresenting most missing rows.
    • B. Preserving the exact mean is a side effect of mean imputation, not the actual goal; the relevant question is whether the imputed value is representative, and skew undermines that for the mean.
    • C. Annual income is a continuous numeric feature, so the mode is not the natural choice; this claim also does not engage with the colleague's argument about the mean at all.
    • D. A median is well defined for any ordered numeric distribution including a right-skewed one, so this claim about the median being unavailable is factually incorrect.

    Subdomain 2.6: Impute missing values with the mode, mean, or median value

    18.A data engineer writes the following code to fill missing values in a heavily right-skewed `transaction_amount` column on a Spark DataFrame, expecting the median to be used: ``` from pyspark.ml.feature import Imputer imputer = Imputer(inputCols=["transaction_amount"], outputCols=["transaction_amount_filled"]) model = imputer.fit(df) ``` What is the actual result of running this code, and what change is needed to get the intended behavior?

    1. A.`Imputer` defaults to `strategy='mean'`, so the outlier-sensitive mean is used instead of the median; the code must set `strategy='median'` explicitly to match the intended behavior.
    2. B.`Imputer` defaults to `strategy='median'`, so the code already fills the column with the median as intended, and no change to the constructor arguments is required.
    3. C.The code raises an error at `fit` time because `outputCols` must exactly match the names given in `inputCols`, so the two column-name lists need to be made identical before any strategy can be applied.
    4. D.`Imputer` defaults to `strategy='mode'`, so the most frequent transaction amount is used instead of the median; the code must set `strategy='median'` explicitly to fix this.
    Show answer & explanation

    Correct answer: A — `Imputer` defaults to `strategy='mean'`, so the outlier-sensitive mean is used instead of the median; the code must set `strategy='median'` explicitly to match the intended behavior.

    • A. `pyspark.ml.feature.Imputer` defaults to the mean strategy when `strategy` is not passed, so this code silently fills the skewed column with an outlier-sensitive mean rather than the intended median, and `strategy='median'` must be added explicitly.
    • B. The default strategy for `pyspark.ml.feature.Imputer` is mean, not median, so this code does not already produce the intended behavior and would fill the skewed column with a value pulled toward its outliers.
    • C. `inputCols` and `outputCols` are allowed to differ, which is how a new filled column is created alongside the original; this code's mismatched names are valid and do not cause a `fit`-time error.
    • D. Mode is not the default strategy for `pyspark.ml.feature.Imputer`; the default is mean, so describing this failure as a mode-based fill misidentifies both the actual behavior and the fix needed.

    Subdomain 2.8: Identify and explain the model types or data sets for which one-hot encoding is or is not appropriate.

    19.A model trained with `OneHotEncoder(handle_unknown='ignore')` on a `subscription_tier` feature is deployed for real-time scoring in Databricks Model Serving. In production, a request arrives containing a new tier value, `enterprise-plus`, that did not exist in the training data. What happens when this request is scored?

    1. A.The encoder outputs a row of all zeros for `subscription_tier`, so the model scores the request using no information from that feature rather than raising an error.
    2. B.The encoder raises a `ValueError` and the scoring request fails, because `handle_unknown='ignore'` only suppresses warnings during `.fit()` and has no effect during `.transform()`.
    3. C.The encoder automatically creates a new binary column for `enterprise-plus` on the fly, expanding the model's input dimensionality to match the newly observed category.
    4. D.The encoder maps `enterprise-plus` to the most frequent training category, silently substituting it in place of the unseen value before the row reaches the model.
    Show answer & explanation

    Correct answer: A — The encoder outputs a row of all zeros for `subscription_tier`, so the model scores the request using no information from that feature rather than raising an error.

    • A. Correct — with `handle_unknown='ignore'` set at fit time, any category not seen during training is encoded as all zeros across that feature's one-hot columns at transform time, letting the pipeline continue scoring instead of raising an error, though the model then has no signal from that feature for this row.
    • B. Incorrect — `handle_unknown='ignore'` specifically changes behavior at `.transform()` time; the default `handle_unknown='error'` is what raises a `ValueError` on unseen categories, so describing the parameter as fit-time-only reverses its actual purpose.
    • C. Incorrect — a fitted `OneHotEncoder` has a fixed set of output columns determined during `.fit()`; it cannot expand the number of columns at transform time, since that would change the input shape the downstream model was trained to expect.
    • D. Incorrect — `OneHotEncoder` has no built-in logic for substituting the most frequent category for unseen values; that behavior would require a separate imputation step, and `handle_unknown='ignore'` produces all zeros, not a substituted category.

    Subdomain 2.7: Use one-hot encoding for categorical features

    20.A data scientist is building a gradient-boosted tree ensemble to classify support tickets using a `category` column with 40 distinct values. A colleague suggests one-hot encoding it the same way they would for the logistic regression baseline. What should the data scientist consider before doing so?

    1. A.Tree-based models split on individual feature values, so encoding `category` with `OrdinalEncoder` instead often keeps the feature compact while the trees still learn useful splits.
    2. B.One-hot encoding is required before fitting any scikit-learn estimator, tree-based or otherwise, so the suggestion to encode `category` the same way is the only technically valid option here.
    3. C.Gradient-boosted trees cannot process categorical data in any encoded form, so the `category` column must be dropped from the feature set entirely before training.
    4. D.One-hot encoding a 40-level column produces exactly the same split quality as leaving it as ordinal integers, so the choice between the two encodings has no effect on this model.
    Show answer & explanation

    Correct answer: A — Tree-based models split on individual feature values, so encoding `category` with `OrdinalEncoder` instead often keeps the feature compact while the trees still learn useful splits.

    • A. Because tree ensembles choose split thresholds on individual columns rather than forming a linear combination of features, an ordinal-encoded categorical column can still be split on effectively, avoiding the wide, sparse feature set that one-hot encoding a 40-level column would create.
    • B. Some scikit-learn estimators and gradient-boosting implementations can accept categorical or ordinal-encoded inputs directly, so one-hot encoding every categorical column before every estimator is not a universal requirement.
    • C. Gradient-boosted trees can consume encoded categorical features just fine, whether ordinal integers or one-hot columns, so dropping the column outright discards a feature the model could otherwise use.
    • D. The two encodings are not interchangeable in effect: one-hot encoding gives the tree a separate binary split per category while ordinal encoding forces splits along a single learned numeric ordering, which can change which splits the model finds and how many are needed.

    Domain 3: Model Development

    Subdomain 3.3: Compare estimators and transformers

    21.A data scientist builds a scikit-learn `Pipeline` with steps `[('scaler', StandardScaler()), ('pca', PCA(n_components=10)), ('clf', LogisticRegression())]`. A teammate asks why `StandardScaler` and `PCA` can occupy the middle positions while `LogisticRegression` can only occupy the last position. What is the correct explanation?

    1. A.Every step before the last one must implement `fit` and `transform` so its output can feed the next step, while only the final step is allowed to be a plain estimator like a classifier.
    2. B.`Pipeline` requires every step, including the last one, to implement `transform`, so `pipe.transform()` always returns a fitted feature matrix no matter which estimator sits at the end.
    3. C.Only the first step in a `Pipeline` needs a `fit` method; each later step, including `LogisticRegression`, must implement `transform` to reshape features before the classifier scores them.
    4. D.`Pipeline` runs all steps independently and in parallel on the raw input, so `StandardScaler` and `PCA` never actually see the output of each other before `LogisticRegression` is trained.
    Show answer & explanation

    Correct answer: A — Every step before the last one must implement `fit` and `transform` so its output can feed the next step, while only the final step is allowed to be a plain estimator like a classifier.

    • A. This is correct: scikit-learn's `Pipeline` contract requires every step except the last to implement both `fit` and `transform`, because each step's transformed output becomes the next step's input, and only the final step is allowed to be a non-transformer such as a classifier or regressor.
    • B. This is backwards; `Pipeline` does not require the last step to implement `transform` at all. When the last step is a classifier like `LogisticRegression`, the pipeline exposes `predict` instead, and calling `.transform()` on it raises an `AttributeError`.
    • C. This misstates the rule: it is every step except the last that needs `transform`, not just the first, and `LogisticRegression` is deliberately exempt from implementing `transform` precisely because it is the final estimator that produces predictions.
    • D. `Pipeline` executes its steps sequentially, not in parallel; each step's `fit_transform` output is passed directly as the input to the next step, which is exactly why `PCA` operates on the scaled output of `StandardScaler` rather than on raw features.

    Subdomain 3.1: Use ML foundations to select the appropriate algorithm for a given model scenario

    22.A data scientist on Databricks is predicting a continuous target, total monthly revenue, from five numeric features. Diagnostic plots show a roughly straight-line relationship between each feature and revenue, and the business asks for a model whose coefficients can be quoted directly to stakeholders as "revenue changes by X per unit of feature Y." Which model best fits this scenario?

    1. A.Linear regression, because it fits a continuous target with a linear relationship to the predictors and produces coefficients that directly quantify each feature's effect on the outcome.
    2. B.Logistic regression, because it estimates class probabilities from a linear combination of features and outputs coefficients that translate into odds ratios for reporting.
    3. C.A random forest regressor, because averaging many decision trees reduces variance and produces a feature importance ranking stakeholders can read as each feature's relative contribution to revenue.
    4. D.A k-means clustering model, because grouping similar revenue records together and presenting the resulting cluster centers as typical revenue profiles per feature satisfies the reporting request.
    Show answer & explanation

    Correct answer: A — Linear regression, because it fits a continuous target with a linear relationship to the predictors and produces coefficients that directly quantify each feature's effect on the outcome.

    • A. This is correct because linear regression is designed for continuous targets with linear predictor relationships, and its fitted coefficients have a direct, interpretable meaning: the expected change in the target per unit change in a feature.
    • B. This is incorrect because logistic regression models the probability of a categorical (typically binary) outcome, not a continuous target like monthly revenue, so its coefficients describe log-odds rather than a direct revenue effect.
    • C. This is incorrect because a random forest is an ensemble of decision trees without a single coefficient per feature; feature importance ranks relative influence but cannot be quoted as a per-unit revenue change the way a linear coefficient can.
    • D. This is incorrect because k-means is an unsupervised clustering technique that groups observations by similarity and does not model or predict a target variable like revenue at all.

    Subdomain 3.2: Identify methods to mitigate data imbalance in training data

    23.A team building a churn model on Databricks decides to apply SMOTE to the minority (churned) class before training. A reviewer asks how SMOTE actually generates its additional training rows. Which description is accurate?

    1. A.It creates new synthetic rows by interpolating between a minority-class sample and one of its nearest minority neighbors.
    2. B.It copies existing minority-class rows verbatim and appends the copies to the training set until the classes are balanced.
    3. C.It removes majority-class rows nearest to the decision boundary so the remaining classes end up closer in size.
    4. D.It reweights the loss function so minority-class errors count more, without adding or removing any rows from the dataset.
    Show answer & explanation

    Correct answer: A — It creates new synthetic rows by interpolating between a minority-class sample and one of its nearest minority neighbors.

    • A. SMOTE generates a new point along the line segment between a minority-class sample and one of its k-nearest minority-class neighbors, producing genuinely new feature values rather than exact copies. This interpolation is the defining mechanism of the technique.
    • B. Verbatim duplication describes random oversampling, not SMOTE; SMOTE is specifically designed to avoid exact duplicates by synthesizing new feature vectors instead of repeating rows.
    • C. Removing majority-class rows near the boundary describes an undersampling strategy such as Tomek links or edited nearest neighbors, not SMOTE, which only adds minority-class rows.
    • D. Adjusting the loss function without changing the dataset describes class weighting, a separate mitigation technique that leaves the row count untouched, unlike SMOTE.

    Subdomain 3.6: Perform random or grid search or Bayesian search as a method for tuning hyperparameters.

    24.A data scientist is configuring Hyperopt's `fmin` and wants the search to use a Bayesian-style approach that adapts its choice of the next hyperparameter set based on the outcomes of earlier trials, rather than sampling each trial independently of the others. Which `algo` argument accomplishes this?

    1. A.`algo=tpe.suggest`, which implements the Tree-structured Parzen Estimator and models prior trial outcomes to propose the next, more promising candidate.
    2. B.`algo=rand.suggest`, which draws each candidate independently from the defined search space and therefore reaches a comparable result using less per-trial computation overall.
    3. C.`algo=hp.choice`, which defines a categorical hyperparameter dimension within the search space and does not control how candidates are proposed across successive trials.
    4. D.`algo=Trials()`, which stores the history of completed evaluations for later inspection but does not itself influence how the next candidate gets selected.
    Show answer & explanation

    Correct answer: A — `algo=tpe.suggest`, which implements the Tree-structured Parzen Estimator and models prior trial outcomes to propose the next, more promising candidate.

    • A. This is correct: `tpe.suggest` is Hyperopt's Bayesian-style algorithm, and it explicitly conditions each new candidate on the distribution of hyperparameter values that scored well versus poorly in earlier trials.
    • B. `rand.suggest` implements plain random search, drawing each trial's candidate independently of every prior trial, which is the opposite of the adaptive, history-informed behavior being asked for.
    • C. `hp.choice` is used inside the search-space definition to declare a categorical parameter; it plays no role in the `algo` argument and does not determine how successive candidates are proposed.
    • D. `Trials()` is the object passed to `fmin` to record evaluation history for later inspection; it is not a valid value for `algo` and does not itself decide how the next candidate is chosen.

    Subdomain 3.10: Identify the number of models being trained in conjunction with a grid-search and cross-validation process.

    25.In scikit-learn's `GridSearchCV`, what determines the total number of model fits performed during the search phase, before any optional final refit on the full training data?

    1. A.The product of the number of hyperparameter combinations in the grid and the number of cross-validation folds specified by `cv`
    2. B.The number of hyperparameter combinations in the grid alone, since each combination is scored once regardless of the fold count
    3. C.The number of cross-validation folds alone, since every hyperparameter combination shares the same set of fold splits
    4. D.The sum of the number of hyperparameter combinations and the number of cross-validation folds specified by `cv`
    Show answer & explanation

    Correct answer: A — The product of the number of hyperparameter combinations in the grid and the number of cross-validation folds specified by `cv`

    • A. `GridSearchCV` trains a distinct model for every hyperparameter combination on every fold's training split, so the total search-phase fit count is the product of the combination count and the fold count, matching scikit-learn's documented behavior.
    • B. This ignores that each combination must still be trained separately on each fold to produce a cross-validated score. Fold count is not incidental; it directly multiplies the number of fits performed for every combination in the grid.
    • C. This ignores that a distinct model is trained for every hyperparameter combination within each fold, not just once per fold. Sharing the same fold splits across combinations does not mean the combinations share trained models.
    • D. Adding the combination count and fold count instead of multiplying them misrepresents the nested structure of the search, where every combination is independently retrained across all folds rather than folds and combinations being counted separately.

    Subdomain 3.8: Describe the benefits and downsides of using cross-validation over a train-validation split.

    26.A data scientist notices that a single train-validation split on a customer dataset happened to place nearly all of one important customer segment into the training partition and almost none into the validation partition, making the validation score misleading. Which statement correctly describes how k-fold cross-validation relates to this specific problem?

    1. A.Cross-validation rotates which rows serve as the test fold, so no single arbitrary partition dominates the result, and a stratified k-fold variant additionally preserves segment proportions in every fold.
    2. B.Cross-validation completely eliminates this segment-imbalance problem by removing all randomness from fold assignment, guaranteeing every fold contains an identical distribution of every customer segment.
    3. C.Cross-validation does not affect this problem at all, because folds are drawn from the dataset using the exact same random assignment process a single train-validation split already uses.
    4. D.Cross-validation only fixes this problem once the fold count is raised to match the number of customer segments present, so fold count must track segment count rather than dataset size.
    Show answer & explanation

    Correct answer: A — Cross-validation rotates which rows serve as the test fold, so no single arbitrary partition dominates the result, and a stratified k-fold variant additionally preserves segment proportions in every fold.

    • A. This is correct: because each row eventually serves as test data across the rotation of folds, no single arbitrary partition dominates the reported metric, and `StratifiedKFold` goes further by explicitly preserving the proportion of each class or segment within every fold, directly addressing the described imbalance.
    • B. Standard k-fold cross-validation still assigns rows to folds using randomized partitioning by default and does not guarantee identical segment distributions across folds unless a stratified variant is explicitly used.
    • C. Cross-validation differs meaningfully from a single split because it averages results across multiple rotating partitions rather than relying on one fixed assignment, so it is not equivalent to repeating the same random process once.
    • D. The number of folds is a general hyperparameter typically chosen based on dataset size and compute budget (commonly 5 or 10), not a value that must be matched to the number of customer segments present in the data.

    Subdomain 3.9: Perform cross-validation as a part of model fitting.

    27.A binary classification dataset is severely imbalanced, with 95% negative and 5% positive labels. An engineer wants to run k-fold cross-validation while keeping the same 95/5 class ratio inside every fold. Which cross-validator should they use?

    1. A.StratifiedKFold
    2. B.KFold
    3. C.ShuffleSplit
    4. D.LeaveOneOut
    Show answer & explanation

    Correct answer: A — StratifiedKFold

    • A. `StratifiedKFold` builds each fold so the proportion of each class matches the proportion in the full dataset, which keeps the rare positive class represented in every fold's train and test split.
    • B. Plain `KFold` splits rows into folds without regard to class labels, so with a 95/5 imbalance it can easily produce folds where the minority class is underrepresented or missing entirely.
    • C. `ShuffleSplit` generates random train/test partitions with possible overlap between iterations and does not account for class proportions, so it offers no built-in guarantee of preserving the 95/5 ratio.
    • D. `LeaveOneOut` holds out a single row per iteration and trains on the rest, which is computationally expensive on larger data and does nothing to balance class proportions within folds.

    Subdomain 3.12: Use common regression metrics: RMSE, MAE, R-squared, etc.

    28.After training a regression model to predict house prices, a machine learning engineer evaluates it on a holdout set and finds an R-squared value of -0.15. What does this result indicate about the model?

    1. A.The model fits the holdout data worse than a naive baseline that always predicts the average house price, so the learned relationship is actively misleading rather than simply weak
    2. B.The model explains about fifteen percent less variance than a perfect predictor, which is a mediocre but still broadly acceptable result for an early modeling iteration
    3. C.The negative score signals a data leakage bug in the training pipeline, since the coefficient of determination is mathematically bounded between zero and one
    4. D.Root mean squared error and mean absolute error must both equal zero for this model, because a negative coefficient of determination only occurs when squared error terms fully cancel out
    Show answer & explanation

    Correct answer: A — The model fits the holdout data worse than a naive baseline that always predicts the average house price, so the learned relationship is actively misleading rather than simply weak

    • A. R-squared compares the model's squared error against the squared error of always predicting the mean; a negative value means the fitted model performs worse than that trivial baseline, which is a genuine warning sign rather than a rounding artifact.
    • B. R-squared is not a percentage scaled from zero to one hundred in this way, and a negative value does not mean the model is merely fifteen percent worse than perfect; it means the model underperforms the mean-prediction baseline.
    • C. The coefficient of determination is only bounded above by one; it is unbounded below and can legitimately go negative whenever a model fits worse than the mean, so a negative value alone does not indicate a pipeline bug.
    • D. Root mean squared error and mean absolute error are unrelated in magnitude to whether R-squared is negative, and neither metric is forced to zero by a negative coefficient of determination.

    Subdomain 3.15: Assess the impact of model complexity and the bias variance tradeoff on model performance

    29.A data scientist sets `n_neighbors=1` for a k-nearest neighbors classifier and observes that predictions change drastically when a single training point near the decision boundary is removed. Which statement best explains this behavior and how to address it?

    1. A.With `n_neighbors=1`, model complexity is effectively maximized because each prediction depends on one nearby point, producing high variance; increasing `n_neighbors` trades some variance for bias.
    2. B.With `n_neighbors=1`, the model averages over too many neighbors, producing high bias; decreasing `n_neighbors` further would smooth the boundary and reduce sensitivity to individual points.
    3. C.The instability comes from an unscaled feature space rather than model complexity; standardizing the features while keeping `n_neighbors=1` would make predictions stable regardless of neighbor count.
    4. D.k-nearest neighbors is a parametric model, so its complexity is fixed by the feature count rather than `n_neighbors`; the instability must come from a bug in the training pipeline.
    Show answer & explanation

    Correct answer: A — With `n_neighbors=1`, model complexity is effectively maximized because each prediction depends on one nearby point, producing high variance; increasing `n_neighbors` trades some variance for bias.

    • A. A single-neighbor classifier draws a highly irregular decision boundary that hugs each individual training point, so removing or adding one point can flip nearby predictions — the hallmark of high variance. Increasing `n_neighbors` averages over more points, smoothing the boundary and accepting some bias in exchange for lower variance.
    • B. With only one neighbor there is no averaging happening at all, so this describes the opposite regime from what `n_neighbors=1` actually produces. Decreasing `n_neighbors` below 1 is not possible and would not smooth anything even if it were.
    • C. Feature scaling affects which points are considered nearest, but it does not remove the fundamental instability caused by basing every prediction on a single neighbor. Even on perfectly scaled features, a one-neighbor model remains maximally sensitive to individual training points.
    • D. k-nearest neighbors is a classic example of a non-parametric method, and `n_neighbors` is precisely the hyperparameter that controls its effective complexity. The observed instability is expected behavior for this setting, not evidence of a pipeline defect.

    Subdomain 3.14: Identify the need to exponentiate log-transformed variables before calculating evaluation metrics or interpreting predictions

    30.A data scientist computes MAE directly on a model's `np.log1p`-transformed predictions and the corresponding `np.log1p`-transformed actual values, then tells the business team "our model is off by an average of 0.18." What is wrong with this statement?

    1. A.An MAE of 0.18 is measured in log-transformed units, not the target's original units, so it does not describe how far off predictions are in real terms.
    2. B.MAE cannot be computed on log-transformed data at all, so the reported value of 0.18 is not a valid metric and the whole calculation must be redone.
    3. C.The statement is correct as written, because MAE computed on log-transformed values is always numerically identical to MAE on the original, untransformed target.
    4. D.The value 0.18 is implausibly small for any regression metric, which suggests the model was almost certainly trained against the wrong target column.
    Show answer & explanation

    Correct answer: A — An MAE of 0.18 is measured in log-transformed units, not the target's original units, so it does not describe how far off predictions are in real terms.

    • A. This is correct: 0.18 is an average absolute difference in log-space, which does not translate linearly into a dollar or unit difference; exponentiating predictions and actuals first is needed before the error can be described in the target's real-world units.
    • B. MAE is a perfectly valid calculation on log-transformed arrays; the numeric computation itself succeeds, the problem is only that the resulting figure is in the wrong units for the business audience.
    • C. MAE computed in log space and MAE computed in original units are generally different numbers, especially for right-skewed targets, because the log transform compresses large values disproportionately.
    • D. A small MAE value like 0.18 is exactly what is expected when errors are measured on a compressed log scale, so its size alone does not indicate a wrong target column was used.

    Domain 4: Model Deployment

    Subdomain 4.2: Deploy a custom model to a model endpoint

    31.A team has trained a custom scikit-learn fraud-detection model and registered it in Unity Catalog. The application must return a prediction within 50 milliseconds of each transaction as it happens, and transactions arrive continuously throughout the day. Which approach should the team use to serve predictions?

    1. A.Deploy the registered model to a Model Serving endpoint as a served entity, then have the transaction service call the endpoint's REST API synchronously for each incoming transaction.
    2. B.Schedule a nightly Lakeflow job that loads the model with `mlflow.pyfunc.load_model` and scores the day's transactions in bulk through a vectorized pandas UDF on a Spark DataFrame.
    3. C.Register the model inside a Lakeflow Spark Declarative Pipeline flow so it scores transaction records as they land in a bronze streaming table before downstream aggregation runs on them.
    4. D.Load the model inside a notebook and rerun the scoring cell interactively whenever an analyst opens a flagged transaction, recording the resulting score by hand in a shared tracking table.
    Show answer & explanation

    Correct answer: A — Deploy the registered model to a Model Serving endpoint as a served entity, then have the transaction service call the endpoint's REST API synchronously for each incoming transaction.

    • A. A Model Serving endpoint keeps the model loaded and ready to respond to individual REST calls with single-digit-to-double-digit millisecond overhead, which matches a per-transaction, sub-50ms latency requirement.
    • B. A nightly batch job scores transactions long after they occur, so it cannot meet a requirement that each transaction be scored the moment it arrives.
    • C. A streaming pipeline flow scores records as micro-batches land in a table, which adds pipeline trigger and table-commit latency that is far larger than the required per-transaction response time.
    • D. Manually rerunning a notebook cell depends on an analyst being available and cannot deliver a consistent, automated response within a fixed millisecond budget for every transaction.

    Subdomain 4.3: Use pandas to perform batch inference

    32.A data engineer has a Delta table with 50 million transaction rows and a model registered at `models:/fraud_detector/Production`. They need to score every row and write predictions to a new Delta table while distributing the scoring work across the cluster. Which approach should they use?

    1. A.Call `mlflow.pyfunc.load_model("models:/fraud_detector/Production")` on the driver, convert the table to pandas with `.toPandas()`, then run a single `.predict()` call over the whole DataFrame at once.
    2. B.Load the model with `mlflow.pyfunc.spark_udf(spark, "models:/fraud_detector/Production")` and add predictions to the Spark DataFrame via `withColumn`, letting executors score partitions in parallel.
    3. C.Register the model on a real-time Model Serving endpoint and loop over the 50 million rows in the driver notebook, sending one REST request per transaction to the endpoint.
    4. D.Skip Spark and call `ai_query()` from a SQL notebook cell against a deployed serving endpoint, since it removes the need to load the model or write any DataFrame or pandas UDF code here.
    Show answer & explanation

    Correct answer: B — Load the model with `mlflow.pyfunc.spark_udf(spark, "models:/fraud_detector/Production")` and add predictions to the Spark DataFrame via `withColumn`, letting executors score partitions in parallel.

    • A. Collecting 50 million rows into a single pandas DataFrame with `.toPandas()` pulls the entire dataset into driver memory, which risks an out-of-memory failure and abandons the cluster's distributed executors entirely. It also removes any parallelism, since one `.predict()` call runs sequentially on the driver.
    • B. Wrapping the registered model as a pandas UDF with `mlflow.pyfunc.spark_udf` and applying it through `withColumn` lets each executor load the model once and score its own partition in parallel, which is the standard pattern for distributed batch inference over a large Delta table.
    • C. Issuing 50 million individual REST calls from a single driver notebook is extremely slow, adds network latency per row, and does not use the cluster's compute for scoring at all, making it impractical at this scale.
    • D. This SQL function can call a served model per row, but it requires a live serving endpoint to already be deployed and bypasses pandas or Spark DataFrame scoring altogether, so it does not match a pandas-based batch scoring workflow over an existing Delta table.

    Subdomain 4.5: Deploy and query a model for realtime inference

    33.A team is deploying a large image-classification deep learning model to a Model Serving endpoint and finds that CPU-based serving does not meet the required response latency under production load. Which change to the served entity's configuration is most appropriate?

    1. A.Change the served entity's `workload_type` to a GPU option such as `GPU_SMALL` or `GPU_MEDIUM` so inference runs on accelerator hardware sized for the model.
    2. B.Increase the `entity_version` number on the served entity, since a higher version number automatically provisions faster underlying serving hardware.
    3. C.Disable `scale_to_zero_enabled` and leave `workload_type` on `CPU`, since keeping compute always warm resolves latency caused by GPU-only architectures.
    4. D.Add a second served entity with the identical CPU configuration and split traffic evenly, since doubling CPU replicas replaces the need for GPU acceleration.
    Show answer & explanation

    Correct answer: A — Change the served entity's `workload_type` to a GPU option such as `GPU_SMALL` or `GPU_MEDIUM` so inference runs on accelerator hardware sized for the model.

    • A. Setting `workload_type` to a GPU option moves inference onto accelerator hardware, which is the configuration change that addresses latency for a compute-heavy deep learning model that CPU serving cannot meet.
    • B. `entity_version` identifies which trained model artifact is served; incrementing it does not change or upgrade the underlying compute hardware assigned to the served entity.
    • C. Keeping CPU compute warm avoids cold-start delay but does not add the parallel compute throughput a GPU provides, so it does not resolve latency caused by an under-powered CPU workload type.
    • D. Adding more CPU replicas increases throughput for concurrent requests but does not reduce the per-request inference time the way GPU acceleration does for a compute-heavy model.

    Subdomain 4.4: Identify how streaming inference is performed with Delta Live Tables

    34.A single Lakeflow pipeline defines two flows: one applies AUTO CDC logic to merge inserts, updates, and deletes into a `customers` dimension table, and the other continuously appends model predictions onto incoming clickstream events using `mlflow.pyfunc.spark_udf`. Which statement best describes how these two flows relate to each other?

    1. A.Each flow independently reads its own source, applies its own transformation, and writes to its own target dataset, so the CDC merge flow and the model-scoring flow can both run in the same pipeline.
    2. B.Both flows must write into the same target table, because a single pipeline is limited to exactly one streaming table target, so the CDC output and the scored predictions have to be unioned before being written.
    3. C.The AUTO CDC flow always executes only after every other flow in the pipeline has completed, because change-data-capture flows are scheduled last regardless of the dependency graph between datasets.
    4. D.Applying a model UDF inside a flow automatically converts that flow's target into a materialized view, because Lakeflow treats any UDF-based transformation as forcing a full, batch-style recompute.
    Show answer & explanation

    Correct answer: A — Each flow independently reads its own source, applies its own transformation, and writes to its own target dataset, so the CDC merge flow and the model-scoring flow can both run in the same pipeline.

    • A. This is correct: a flow reads from a source, applies its own transformation, and writes to its own target, so an AUTO CDC merge flow and a model-scoring flow can coexist as independent flows within one pipeline.
    • B. A pipeline is not limited to a single streaming table target; each flow can write to its own distinct target dataset, so unioning the CDC output with scored predictions is not required.
    • C. Flow execution order is derived from the dependency graph between datasets, not a fixed rule that change-data-capture flows always run last regardless of what they depend on.
    • D. Using a UDF, including a model-scoring UDF, inside a flow does not change the target's dataset type; a streaming table continues to process data incrementally regardless of the transformation applied.

    Subdomain 4.6: Split data between endpoints for realtime interference

    35.A team is rolling out a new pricing model as a canary. They start it at 5% of endpoint traffic, and over the following days plan to raise its share to 25%, then 50%, then 100% as confidence grows, while the old version keeps serving the remainder at each stage. What is the main operational advantage of doing this rollout through the endpoint's traffic split instead of switching client applications to a new endpoint URL at each stage?

    1. A.Client applications keep calling the same endpoint URL throughout the rollout, so each traffic_config update changes routing behind the scenes without any client redeployment.
    2. B.The endpoint automatically reverts the traffic_config to its previous percentages whenever the new version average response time increases by any measurable amount during the rollout.
    3. C.Traffic-split rollouts remove the need for the new pricing model to be registered in Unity Catalog before it can receive any share of the endpoint's live traffic.
    4. D.Splitting traffic through one endpoint guarantees the new and old pricing models will return numerically identical predictions for any request routed during the rollout.
    Show answer & explanation

    Correct answer: A — Client applications keep calling the same endpoint URL throughout the rollout, so each traffic_config update changes routing behind the scenes without any client redeployment.

    • A. Because the served entities and their traffic split live inside one endpoint's configuration, raising the canary's share from 5% to 25% to 50% only requires updating that endpoint's traffic_config, so calling applications never need to change the URL they invoke at any stage.
    • B. Databricks does not automatically roll back traffic percentages based on observed latency changes; adjusting or reverting the split in response to performance is a manual decision the team makes using metrics from tools like inference tables.
    • C. A served entity must still reference a model registered in Unity Catalog before it can be added to the endpoint and receive any traffic share, so registration remains a prerequisite regardless of the rollout strategy.
    • D. Splitting traffic routes each request to one model or the other; it does not make the two different model versions produce identical outputs, since they are distinct trained artifacts making independent predictions.

    Want the full experience?

    These are just samples. Practice the full Databricks Certified Machine Learning Associate question bank in quiz mode — free, no signup, with domain practice and exam simulation.