Databricks Certified Machine Learning Associate Lessons
48 lessons, one per exam-guide subdomain, in the order the guide teaches them. Every claim is cited to the official documentation.
Domain 1: Databricks Machine Learning
18 lessons · 37% of the exam
1.1MLOps Best Practices on Databricks: Environments and Deploy-Code vs Deploy-Models
Subdomain 1.1: Identify the best practices of an MLOps strategy
2 pages · 18 min read
1.2Databricks Runtime ML: Pre-Installed Libraries and When to Use It
Subdomain 1.2: Identify the advantages of using ML runtimes
2 pages · 19 min read
1.3AutoML Model and Feature Selection in Databricks
Subdomain 1.3: Identify how AutoML facilitates model/feature selection.
15 min read
1.4AutoML Advantages in Databricks Model Development
Subdomain 1.4: Identify the advantages AutoML brings to the model development process
14 min read
1.5Unity Catalog Feature Tables vs Workspace Feature Store: Scope and Access
Subdomain 1.5: Identify the benefits of creating feature store tables at the account level in Unity Catalog in Databricks vs at the workspace level
2 pages · 15 min read
1.6Create a Feature Table in Unity Catalog with SQL or FeatureEngineeringClient
Subdomain 1.6: Create a feature store table in Unity Catalog
2 pages · 22 min read
1.7Writing Data to Feature Tables with write_table
Subdomain 1.7: Write data to a feature store table
15 min read
1.8Build a Training Set with FeatureLookup and create_training_set
Subdomain 1.8: Train a model with features from a feature store table.
2 pages · 22 min read
1.9Batch Scoring with fe.score_batch and Feature Store Tables
Subdomain 1.9: Score a model using features from a feature store table.
2 pages · 18 min read
1.10Offline vs Online Feature Tables in Databricks
Subdomain 1.10: Describe the differences between online and offline feature tables
2 pages · 18 min read
1.11Find the Best MLflow Run with search_runs
Subdomain 1.11: Identify the best run using the MLflow Client API.
13 min read
1.12Manually Log Parameters, Metrics and Artifacts in an MLflow Run
Subdomain 1.12: Manually log metrics, artifacts, and models in an MLflow Run.
2 pages · 19 min read
1.13MLflow UI: Experiments, Run Pages, Parameters, Metrics, Tags and Artifacts
Subdomain 1.13: Identify information available in the MLFlow UI
2 pages · 17 min read
1.14Register Models in Unity Catalog with the MLflow Client API
Subdomain 1.14: Register a model using the MLflow Client API in the Unity Catalog registry
15 min read
1.15Unity Catalog Model Registry: Governance and Cross-Workspace Benefits
Subdomain 1.15: Identify benefits of registering models in the Unity Catalog registry over the workspace registry
2 pages · 16 min read
1.16Deploy Code vs Deploy Models: Why Databricks Promotes Code by Default
Subdomain 1.16: Identify scenarios where promoting code is preferred over promoting models and vice versa
2 pages · 17 min read
1.17Set and Remove Model Tags in Unity Catalog
Subdomain 1.17: Set or remove a tag for a model
15 min read
1.18Champion and Challenger Model Aliases in Unity Catalog
Subdomain 1.18: Promote a challenger model to a champion model using aliases
14 min read
Domain 2: Data Processing
9 lessons · 19% of the exam
2.1Spark DataFrame summary statistics with .summary() and dbutils.data.summarize
Subdomain 2.1: Compute summary statistics on a Spark DataFrame using .summary() or dbutils data summaries
15 min read
2.2Removing Outliers with Standard Deviation Bounds in Spark
Subdomain 2.2: Remove outliers from a Spark DataFrame based on standard deviation or IQR
2 pages · 18 min read
2.3Histograms, Bar Charts and Heatmaps for Feature Types
Subdomain 2.3: Create visualizations for categorical or continuous features
2 pages · 18 min read
2.4Comparing Two Features in Spark: Pearson Correlation vs Crosstab
Subdomain 2.4: Compare two categorical or two continuous features using the appropriate method
14 min read
2.5Imputing Missing Values: Mean vs Median vs Mode
Subdomain 2.5: Compare and contrast imputing missing values with the mean or median or mode value
16 min read
2.6Mean, Median, or Mode Imputation: Choosing the Fill Statistic
Subdomain 2.6: Impute missing values with the mode, mean, or median value
2 pages · 18 min read
2.7One-Hot Encoding Categorical Features with OneHotEncoder
Subdomain 2.7: Use one-hot encoding for categorical features
2 pages · 17 min read
2.8One-Hot Encoding: When to Use It and When Not To
Subdomain 2.8: Identify and explain the model types or data sets for which one-hot encoding is or is not appropriate.
14 min read
2.9Log Transformation: Recognizing Skewed and Multiplicative Data
Subdomain 2.9: Identify scenarios where log scale transformation is appropriate
2 pages · 17 min read
Domain 3: Model Development
15 lessons · 31% of the exam
3.1Algorithm Selection: Linear, Logistic and Tree-Based Models
Subdomain 3.1: Use ML foundations to select the appropriate algorithm for a given model scenario
12 min read
3.2Mitigating Class Imbalance: Resampling, SMOTE and Class Weights
Subdomain 3.2: Identify methods to mitigate data imbalance in training data
14 min read
3.3Estimators vs transformers: the fit, transform and predict contracts
Subdomain 3.3: Compare estimators and transformers
12 min read
3.4Building a Model Training Pipeline on Databricks: Structure, scikit-learn Pipelines, and Metrics
Subdomain 3.4: Develop a training pipeline
2 pages · 16 min read
3.5Hyperopt fmin: Objective Function, Search Space and Search Algorithm
Subdomain 3.5: Use Hyperopt's fmin operation to tune a model's hyperparameters
2 pages · 18 min read
3.6Grid, random and Bayesian hyperparameter search compared
Subdomain 3.6: Perform random or grid search or Bayesian search as a method for tuning hyperparameters.
2 pages · 18 min read
3.7Parallelize single-node hyperparameter tuning with Hyperopt SparkTrials
Subdomain 3.7: Parallelize single node models for hyperparameter tuning
17 min read
3.8Cross-Validation vs Train-Validation Split: Tradeoffs
Subdomain 3.8: Describe the benefits and downsides of using cross-validation over a train-validation split.
13 min read
3.9k-Fold Cross-Validation with cross_val_score
Subdomain 3.9: Perform cross-validation as a part of model fitting.
2 pages · 18 min read
3.10Counting Models in Grid Search with Cross-Validation
Subdomain 3.10: Identify the number of models being trained in conjunction with a grid-search and cross-validation process.
11 min read
3.11Classification Metrics: Precision, Recall, F1, ROC AUC and Log Loss
Subdomain 3.11: Use common classification metrics: F1, Log Loss, ROC/AUC, etc
20 min read
3.12Regression error metrics: MSE, RMSE, MAE and their variants
Subdomain 3.12: Use common regression metrics: RMSE, MAE, R-squared, etc.
2 pages · 19 min read
3.13Choosing the Right ML Metric for the Scenario Objective
Subdomain 3.13: Choose the most appropriate metric for a given scenario objective
14 min read
3.14Exponentiating Log-Transformed Targets Before Metrics and Interpretation
Subdomain 3.14: Identify the need to exponentiate log-transformed variables before calculating evaluation metrics or interpreting predictions
12 min read
3.15Bias-Variance Tradeoff and Model Complexity
Subdomain 3.15: Assess the impact of model complexity and the bias variance tradeoff on model performance
13 min read
Domain 4: Model Deployment
6 lessons · 12% of the exam
4.1Batch vs Streaming vs Real-Time Inference: Latency, Cost and Semantics
Subdomain 4.1: Identify the differences and advantages of model serving approaches: batch, realtime, and streaming
2 pages · 15 min read
4.2Deploy a Custom MLflow Model to a Databricks Model Serving Endpoint
Subdomain 4.2: Deploy a custom model to a model endpoint
2 pages · 21 min read
4.3Batch Inference with MLflow pyfunc Models and Spark
Subdomain 4.3: Use pandas to perform batch inference
2 pages · 23 min read
4.4Lakeflow Pipelines (formerly DLT) for Streaming Inference: Streaming Tables and Pipeline Modes
Subdomain 4.4: Identify how streaming inference is performed with Delta Live Tables
2 pages · 19 min read
4.5Deploy a Custom Model to a Databricks Model Serving Endpoint
Subdomain 4.5: Deploy and query a model for realtime inference
2 pages · 21 min read
4.6Traffic Splitting on Databricks Model Serving Endpoints
Subdomain 4.6: Split data between endpoints for realtime interference
17 min read