What you will be able to do
- Explain how AutoML searches across algorithms and hyperparameters and picks the best model by a primary metric
- Use AutoML parameters (primary_metric, exclude_frameworks, timeout_minutes, exclude_cols, imputers, split_col) to steer model and feature selection
- Add features from feature tables to an AutoML experiment with feature_store_lookups
- Use AutoML's generated notebooks and Shapley values to review and reproduce what AutoML selected
Key concept
AutoML as an automated, inspectable model search — You supply a dataset, a problem type and a target. AutoML then prepares the data, trains and tunes many candidate models across several algorithm families, ranks them on one metric and gives you editable source notebooks, so you can see exactly what was chosen and why.
1.What AutoML automates, and what it hands back
Choosing a model normally means many manual decisions: which algorithm family to try, which hyperparameters to tune, how to handle missing values and how to compare the results fairly. Databricks AutoML automates that loop. You provide the dataset and say which kind of problem it is. AutoML then works through a fixed sequence:
1. It cleans and prepares your data. 2. It orchestrates distributed model training and hyperparameter tuning across multiple algorithms. 3. It finds the best model, using open source evaluation algorithms from scikit-learn, xgboost, LightGBM, Prophet and ARIMA. 4. It presents the results, together with generated source-code notebooks for each trial.
Step 2 is the model selection: AutoML doesn't commit to one algorithm, it compares several. Step 1 is where feature handling starts: AutoML decides how to treat each column. Step 4 makes both steps inspectable. The exam calls these generated notebooks "glassbox" notebooks. The documentation itself says the notebooks let you "review, reproduce, and modify the code."
You can start an experiment from a low-code UI (Experiments → the Classification, Regression or Forecasting card → Start training) or from the Python API (databricks.automl.classify, databricks.automl.regress). Both paths expose the same choices: dataset, target, evaluation metric and stopping conditions. AutoML relies on the libraries preinstalled in Databricks Runtime for Machine Learning. The docs warn that adding, removing, upgrading or downgrading those libraries causes run failures. Also, in Databricks Runtime 18.0 ML and above, AutoML is no longer included as a built-in library.
Checkpoint 1 of 8· Put it in order
Put the stages of an AutoML run in the order the documentation lists them.
- 1.Presents the results and generates source code notebooks for each trial
- 2.Orchestrates distributed model training and hyperparameter tuning across multiple algorithms
- 3.Cleans and prepares your data
- 4.Finds the best model using open source evaluation algorithms
Data preparation comes first, then the multi-algorithm training and tuning, then selection of the best model, and finally the results and generated notebooks.
“Orchestrates distributed model training and hyperparameter tuning across multiple algorithms.”Source: docs.databricks.com
Sources1
2.Model selection: candidate algorithms, one ranking metric, a time budget
For classification and regression, AutoML draws its candidates from a fixed set of algorithm families. The decision tree, random forest, logistic regression and linear regression (stochastic gradient descent) models are based on scikit-learn. XGBoost and LightGBM come from their own libraries. Forecasting uses a different set (Prophet, Auto-ARIMA and, on serverless, DeepAR).
| Classification models | Regression models |
|---|---|
| Decision trees | Decision trees |
| Random forests | Random forests |
| Logistic regression | Linear regression with stochastic gradient descent |
| XGBoost | XGBoost |
| LightGBM | LightGBM |
Every trial gets a score, and the score comes from a single primary metric. In the API this is primary_metric, described as the metric "used to evaluate and rank model performance". Classification defaults to f1, with log_loss, precision, accuracy and roc_auc as alternatives. Regression defaults to r2, with mae, rmse and mse as alternatives. This metric decides which model counts as "best". When a run completes, the top row of the runs table shows the best model based on the primary metric. Change the metric and you may get a different winner.
You can also narrow the field. exclude_frameworks (in the UI: excluding training frameworks under Advanced Configuration) removes whole families from the search. Its possible values are sklearn, lightgbm and xgboost, and the default is an empty list, meaning all frameworks are considered.
exclude_cols: Optional[List[str]] = None, # :re[DBR] 10.3 ML and above
exclude_frameworks: Optional[List[str]] = None, # :re[DBR] 10.3 ML and above
feature_store_lookups: Optional[List[Dict]] = None, # :re[DBR] 11.3 LTS ML and above
imputers: Optional[Dict[str, Union[str, Dict[str, Any]]]] = None, # :re[DBR] 10.4 LTS ML and aboveThe last control is how long the search runs. The API parameter timeout_minutes defaults to 120 minutes, with a minimum of 5. The docs say "longer timeouts allow AutoML to run more trials and identify a model with better accuracy." The older max_trials parameter is not supported in Databricks Runtime 11.0 ML and above. On those runtimes the number of trials is not a stopping condition. For classification and regression, AutoML also stops early: it stops training and tuning when the validation metric is no longer improving.
Checkpoint 2 of 8· Check yourself
On Databricks Runtime 14 ML, a data scientist wants an AutoML classification run to explore more candidate models before it stops. Which change does the documentation support?
On 11.0 ML and above, max_trials is not supported and the trial count is not a stopping condition. The time budget is what you control, and a longer timeout lets AutoML run more trials.
“Longer timeouts allow AutoML to run more trials and identify a model with better accuracy.”Source: docs.databricks.com
Checkpoint 3 of 8· Exam question
A data scientist launches a Databricks AutoML classification experiment on a customer churn dataset. The dataset includes a `loyalty_tier_code` column stored as integers (1, 2, 3), but these integers represent categorical membership tiers rather than a continuous quantity. AutoML's semantic type detection treats the column as numeric, and the resulting models underperform because they treat tier 3 as three times tier 1. What should the data scientist do before rerunning the experiment?
Correct answer: D — Add a semantic type annotation marking `loyalty_tier_code` as categorical so AutoML treats its values as discrete labels instead of an ordered numeric quantity.
- A. Excluding the column removes it from training entirely, discarding a field that likely has real predictive value for churn once it is typed correctly. This trades away useful signal instead of fixing the underlying type-detection problem.
- B. The `imputers` parameter only controls how missing values are filled in; it has no effect on whether a column is interpreted as numeric or categorical. There is no indication of missing values in this scenario, so this does not address the issue.
- C. Feature table lookups are for joining in precomputed features from Feature Engineering in Unity Catalog tables that are not already in the training dataframe. The tier column already exists in the dataset, so this mechanism is unrelated to correcting its detected type.
- D. Semantic type annotations let a user override AutoML's automatic type detection for a specific column, so marking `loyalty_tier_code` as categorical stops it from being modeled as an ordered numeric quantity. This directly fixes the root cause: the column's values were being interpreted with a false ordinal relationship.
3.Feature selection: which columns AutoML uses and how it interprets them
You can deselect customer_id (in Databricks Runtime 10.3 ML and above you choose which columns AutoML trains on). You cannot remove the prediction target. You also cannot remove the time column used to split the data. AutoML handles the target as the label, not as a feature.
AutoML's feature handling starts with the columns you let it see. In the UI, once you pick a table its schema appears, and you can choose which columns to train on. The API equivalent is exclude_cols, a "list of columns to ignore during AutoML calculations". Two columns are protected in the UI: you cannot remove the column selected as the prediction target or the time column used to split the data.
Next, AutoML decides how to read each remaining column. In Databricks Runtime 9.1 LTS ML and above, AutoML tries to detect whether a column has a semantic type different from the Spark or pandas data type in the schema, and treats it as the detected type. String or integer columns holding dates become timestamps. String columns holding numbers become numeric. From 10.1 ML, numeric columns that contain categorical IDs are treated as a categorical feature, and string columns containing English text are treated as a text feature. Detection is best effort.
Missing values are handled per column. Unless you say otherwise, AutoML picks an imputation method based on column type and content. You can override this with imputers (UI: the Impute with dropdown). The trade-off: if you specify a non-default imputation method, AutoML does not perform semantic type detection for that column.
| Parameter | Effect on features or the data split |
|---|---|
| exclude_cols | Columns listed here are ignored during AutoML calculations (default: []) |
| imputers | Per-column null handling: mean, median, most_frequent, or {"strategy": "constant", "fill_value": ...}; columns not listed get a default strategy based on type and content |
| time_col | Splits training, validation and test sets chronologically: earliest rows for training, latest for test |
| split_col | Splits rows by user-supplied values train, validate or test; the column is automatically excluded from training features |
AutoML also flags problem columns. With Databricks Runtime 10.1 ML and above, it shows warnings for potential dataset issues such as unsupported column types or high-cardinality columns, on the Warnings tab of the training page or experiment page. Use them to decide what to exclude on the next run. Databricks notes the warnings may not catch every issue.
Checkpoint 4 of 8· Check yourself
A data scientist sets imputers={"zip_code": "most_frequent"} for a numeric zip_code column that holds categorical IDs. What happens to semantic type detection for zip_code?
A custom imputation method switches off semantic type detection for that column. AutoML will not reinterpret zip_code as categorical automatically.
“AutoML does not perform semantic type detection for columns that have custom imputation methods specified.”Source: docs.databricks.com
Checkpoint 5 of 8· Exam question
A machine learning engineer is configuring a Databricks AutoML regression run through the Python API (`automl.regress`). The training dataframe holds only transaction IDs and timestamps, but the engineer wants AutoML to also train on precomputed customer aggregate features that already exist as tables in Feature Engineering in Unity Catalog, joined on `customer_id`. Which approach lets AutoML pull in those features directly as part of the run configuration?
Correct answer: C — Pass the `feature_store_lookups` parameter with the feature table names and the `customer_id` key so AutoML joins those tables into the training set.
- A. The `imputers` parameter configures how missing values are filled in columns that already exist in the training data; it does not join in new columns from separate feature tables. This would not bring the customer aggregates into the run at all.
- B. Semantic type detection only classifies the data type of columns already present in the training dataframe, such as recognizing dates or categorical codes; it does not discover or join external feature tables. Retrieving feature tables requires the explicit lookup parameter.
- C. The `feature_store_lookups` parameter is the documented way to join Feature Engineering in Unity Catalog tables into an AutoML run, by specifying the table names and the lookup key that matches a column in the training dataframe. This is exactly the mechanism needed to bring in the precomputed customer aggregates.
- D. The customer aggregate columns are not yet part of the training dataframe, so there is nothing for `exclude_cols` to skip; that parameter only removes existing columns from an already-assembled dataset. It cannot pull in new columns from external feature tables.
4.Adding features from feature tables
Feature selection can also mean adding columns, not only removing them. AutoML can augment the original input dataset with features from feature tables in Unity Catalog or in the legacy Workspace Feature Store. For classification and regression this requires Databricks Runtime 11.3 LTS ML and above. Forecasting requires 12.2 LTS ML and above.
In the UI, after configuring the experiment, click Join features (optional) and choose a feature table. For each primary key of the feature table, select the matching lookup key, which must be a column in your training dataset. For a time series feature table, you also select a timestamp lookup key, which must also be a column in the training dataset. Add more tables with Add another feature table.
In the API, the same join is a list of dictionaries passed to feature_store_lookups. Each dictionary needs table_name and lookup_key. The column names in lookup_key must be in the same order as the feature table's primary keys. timestamp_lookup_key is required only for time series feature tables, where it drives a point-in-time lookup.
Checkpoint 6 of 8· Fill the gap
Which AutoML run parameter takes this list of feature-table joins?
? = [
{
"table_name": "example.trip_pickup_features",
"lookup_key": ["pickup_zip", "rounded_pickup_datetime"],
},
{
"table_name": "example.trip_dropoff_features",
"lookup_key": ["dropoff_zip", "rounded_dropoff_datetime"],
}
]AutoML reads existing feature tables from the feature_store_lookups parameter. Each entry names a table and the lookup key columns from the training dataset.
Source: docs.databricks.com5.Reviewing the selection: generated notebooks and Shapley values
Because each trial's training code is generated, the model AutoML selects is never a black box. For classification and regression experiments, the data exploration notebook and the best trial's notebook are imported into your workspace automatically. Notebooks for the other trials are saved as MLflow artifacts on DBFS. For those trials, the notebook_path and notebook_url fields in the TrialInfo Python API are not set. You can import them through the AutoML experiment UI or the databricks.automl.import_notebook Python API. You can also open the MLflow run and find the notebook in the run page's Artifacts section, which you can download if your workspace admins allow artifact downloads. Forecasting is different: notebooks for all trials are imported.
When the experiment finishes, View notebook for best model opens the code behind the winner, and View data exploration notebook shows what AutoML saw in your data. Each run page also lists the trial's parameters, metrics and tags, plus its artifacts, including the model. The experiment's results, including the data exploration and training notebooks, are stored in a databricks_automl folder in the home folder of the user who ran it.
These notebooks also answer the feature question: which inputs actually drove the selected model? Regression and classification notebooks include code that computes Shapley values with the SHAP package. Shapley values come from game theory and estimate how much each feature contributes to a model's predictions. The calculation is memory-intensive, so it is off by default. To turn it on, go to the Feature importance section of a generated trial notebook, set shap_enabled = True, and re-run the notebook. On MLR 11.1 and below, SHAP plots are not generated if the dataset contains a datetime column.
Checkpoint 7 of 8· Match them up
Match each AutoML output to where it ends up or what it does.
Tap a term, then the definition that fits it.
Only the data exploration and best-trial notebooks are auto-imported. Other trial notebooks stay as MLflow artifacts. Shapley values are opt-in because they are memory-intensive. The top row of the runs table is the best model by the primary metric.
“Generated notebooks for other experiment trials are saved as MLflow artifacts on DBFS instead of auto-imported into your workspace.”Source: docs.databricks.com
Checkpoint 8 of 8· Exam question
After an AutoML classification run finishes, a data scientist wants to add a custom outlier-clipping step to the preprocessing pipeline of the best-performing trial before registering the model. AutoML has already produced its results. What is the most direct way to make and apply this change?
Correct answer: A — Open the auto-generated notebook for the best trial, which contains the full editable training source code, add the clipping step, and rerun the notebook.
- A. AutoML's glassbox approach generates a full, editable source code notebook for the best trial, so a data scientist can open it directly, insert the clipping step into the existing pipeline code, and rerun it to produce an updated model. This is the direct path from result to a customized, reproducible pipeline.
- B. The data exploration notebook is a separate artifact focused on summarizing and visualizing the dataset before training; it is not wired to automatically re-trigger or modify the best trial. Editing it would have no effect on the actual training pipeline.
- C. Logged parameters in the MLflow UI are a record of what already ran; editing them does not change the training code or re-execute preprocessing, and clipping logic is not something applied at inference time by an MLflow parameter field.
- D. The `imputers` parameter is scoped to missing-value handling for a specific column, not general preprocessing logic like outlier clipping, and launching an entirely new experiment discards the already-completed best trial rather than modifying it.
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.AutoML calculates SHAP feature importance automatically for every trial.Why is that wrong?
The generated notebooks contain the SHAP code, but it doesn't run by default because it is memory-intensive. You set shap_enabled = True in the Feature importance section and re-run the notebook.
Covered in Reviewing the selection: generated notebooks and Shapley values
2.Every AutoML trial notebook for classification and regression is imported into the workspace.Why is that wrong?
Only the data exploration notebook and the best trial's notebook are auto-imported. The others are MLflow artifacts that you import manually.
Covered in Reviewing the selection: generated notebooks and Shapley values
3.Setting a custom imputer for a column keeps AutoML's semantic type detection for that column.Why is that wrong?
A non-default imputation method turns off semantic type detection for that column.
Covered in Feature selection: which columns AutoML uses and how it interprets them
4.On current runtimes you control how many models AutoML tries by setting max_trials.Why is that wrong?
On Databricks Runtime 11.0 ML and above, max_trials is not supported and the trial count is not a stopping condition. Use timeout_minutes; early stopping also ends the search when the validation metric stops improving.
Covered in Model selection: candidate algorithms, one ranking metric, a time budget
5.lookup_key columns in feature_store_lookups can be listed in any order; AutoML matches them by name.Why is that wrong?
The API reference requires the lookup_key column names to be in the same order as the feature table's primary keys.
Covered in Adding features from feature tables
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Finds the best model using open source evaluation algorithms from scikit-learn, xgboost, LightGBM, Prophet, and ARIMA.”
↩︎ What AutoML automates, and what it hands back“In Databricks Runtime 18.0 ML or above, AutoML is not included as a built-in library.”
↩︎ What AutoML automates, and what it hands back“Shapley values are based in game theory and estimate the importance of each feature to a model's predictions.”
↩︎ Reviewing the selection: generated notebooks and Shapley values“Set shap_enabled = True.”
↩︎ Reviewing the selection: generated notebooks and Shapley values“manually import them into your workspace with the AutoML experiment UI or the databricks.automl.import_notebook Python API.”
↩︎ Reviewing the selection: generated notebooks and Shapley values“by automatically finding the best algorithm and hyperparameter configuration for you.”
↩︎ Key concept“Because these calculations are highly memory-intensive, the calculations are not performed by default.”
↩︎ Exam trap 1“Generated notebooks for other experiment trials are saved as MLflow artifacts on DBFS instead of auto-imported into your workspace.”
↩︎ Exam trap 2“AutoML also generates source code notebooks for each trial, allowing you to review, reproduce, and modify the code as needed.”
↩︎ Prediction“Orchestrates distributed model training and hyperparameter tuning across multiple algorithms.”
↩︎ Checkpoint“AutoML-generated notebooks for data exploration and the best trial in your experiment are automatically imported to your workspace.”
↩︎ Prediction“Generated notebooks for other experiment trials are saved as MLflow artifacts on DBFS instead of auto-imported into your workspace.”
↩︎ Checkpoint - 2.
“Metric used to evaluate and rank model performance.”
↩︎ Model selection: candidate algorithms, one ranking metric, a time budget“List of algorithm frameworks that AutoML should not consider as it develops models.”
↩︎ Model selection: candidate algorithms, one ranking metric, a time budget“List of columns to ignore during AutoML calculations.”
↩︎ Feature selection: which columns AutoML uses and how it interprets them“this column is automatically excluded from training features.”
↩︎ Feature selection: which columns AutoML uses and how it interprets them“Required if the specified table is a time series feature table.”
↩︎ Adding features from feature tables“If you specify a non-default imputation method, AutoML does not perform semantic type detection.”
↩︎ Exam trap 3“The order of the column names must match the order of the primary keys of the feature table.”
↩︎ Exam trap 5“Longer timeouts allow AutoML to run more trials and identify a model with better accuracy.”
↩︎ Checkpoint - 3.
“When a run completes, the top row shows the best model based on the primary metric.”
↩︎ Model selection: candidate algorithms, one ranking metric, a time budget“it stops training and tuning models if the validation metric is no longer improving.”
↩︎ Model selection: candidate algorithms, one ranking metric, a time budget“You cannot remove the column selected as the prediction target or the time column to split the data.”
↩︎ Feature selection: which columns AutoML uses and how it interprets them“AutoML displays warnings for potential issues with the dataset, such as unsupported column types or high cardinality columns.”
↩︎ Feature selection: which columns AutoML uses and how it interprets them“For Databricks Runtime 11.0 ML and above, the number of trials is not used as a stopping condition.”
↩︎ Exam trap 4 - 4.
“Numeric columns that contain categorical IDs are treated as a categorical feature.”
↩︎ Feature selection: which columns AutoML uses and how it interprets them“AutoML does not perform semantic type detection for columns that have custom imputation methods specified.”
↩︎ Checkpoint - 5.
“AutoML can augment the original input dataset with features from feature tables in Unity Catalog or in the legacy Workspace Feature Store.”
↩︎ Adding features from feature tables“The lookup key should be a column in the training dataset you provided for your AutoML experiment.”
↩︎ Adding features from feature tables