CertSafari

    Databricks Certified Machine Learning Associate Lessons

    48 lessons, one per exam-guide subdomain, in the order the guide teaches them. Every claim is cited to the official documentation.

    Domain 1: Databricks Machine Learning

    18 lessons · 37% of the exam

    1. 1.1MLOps Best Practices on Databricks: Environments and Deploy-Code vs Deploy-Models

      Subdomain 1.1: Identify the best practices of an MLOps strategy

      2 pages · 18 min read

    2. 1.2Databricks Runtime ML: Pre-Installed Libraries and When to Use It

      Subdomain 1.2: Identify the advantages of using ML runtimes

      2 pages · 19 min read

    3. 1.3AutoML Model and Feature Selection in Databricks

      Subdomain 1.3: Identify how AutoML facilitates model/feature selection.

      15 min read

    4. 1.4AutoML Advantages in Databricks Model Development

      Subdomain 1.4: Identify the advantages AutoML brings to the model development process

      14 min read

    5. 1.5Unity Catalog Feature Tables vs Workspace Feature Store: Scope and Access

      Subdomain 1.5: Identify the benefits of creating feature store tables at the account level in Unity Catalog in Databricks vs at the workspace level

      2 pages · 15 min read

    6. 1.6Create a Feature Table in Unity Catalog with SQL or FeatureEngineeringClient

      Subdomain 1.6: Create a feature store table in Unity Catalog

      2 pages · 22 min read

    7. 1.7Writing Data to Feature Tables with write_table

      Subdomain 1.7: Write data to a feature store table

      15 min read

    8. 1.8Build a Training Set with FeatureLookup and create_training_set

      Subdomain 1.8: Train a model with features from a feature store table.

      2 pages · 22 min read

    9. 1.9Batch Scoring with fe.score_batch and Feature Store Tables

      Subdomain 1.9: Score a model using features from a feature store table.

      2 pages · 18 min read

    10. 1.10Offline vs Online Feature Tables in Databricks

      Subdomain 1.10: Describe the differences between online and offline feature tables

      2 pages · 18 min read

    11. 1.11Find the Best MLflow Run with search_runs

      Subdomain 1.11: Identify the best run using the MLflow Client API.

      13 min read

    12. 1.12Manually Log Parameters, Metrics and Artifacts in an MLflow Run

      Subdomain 1.12: Manually log metrics, artifacts, and models in an MLflow Run.

      2 pages · 19 min read

    13. 1.13MLflow UI: Experiments, Run Pages, Parameters, Metrics, Tags and Artifacts

      Subdomain 1.13: Identify information available in the MLFlow UI

      2 pages · 17 min read

    14. 1.14Register Models in Unity Catalog with the MLflow Client API

      Subdomain 1.14: Register a model using the MLflow Client API in the Unity Catalog registry

      15 min read

    15. 1.15Unity Catalog Model Registry: Governance and Cross-Workspace Benefits

      Subdomain 1.15: Identify benefits of registering models in the Unity Catalog registry over the workspace registry

      2 pages · 16 min read

    16. 1.16Deploy Code vs Deploy Models: Why Databricks Promotes Code by Default

      Subdomain 1.16: Identify scenarios where promoting code is preferred over promoting models and vice versa

      2 pages · 17 min read

    17. 1.17Set and Remove Model Tags in Unity Catalog

      Subdomain 1.17: Set or remove a tag for a model

      15 min read

    18. 1.18Champion and Challenger Model Aliases in Unity Catalog

      Subdomain 1.18: Promote a challenger model to a champion model using aliases

      14 min read

    Domain 2: Data Processing

    9 lessons · 19% of the exam

    1. 2.1Spark DataFrame summary statistics with .summary() and dbutils.data.summarize

      Subdomain 2.1: Compute summary statistics on a Spark DataFrame using .summary() or dbutils data summaries

      15 min read

    2. 2.2Removing Outliers with Standard Deviation Bounds in Spark

      Subdomain 2.2: Remove outliers from a Spark DataFrame based on standard deviation or IQR

      2 pages · 18 min read

    3. 2.3Histograms, Bar Charts and Heatmaps for Feature Types

      Subdomain 2.3: Create visualizations for categorical or continuous features

      2 pages · 18 min read

    4. 2.4Comparing Two Features in Spark: Pearson Correlation vs Crosstab

      Subdomain 2.4: Compare two categorical or two continuous features using the appropriate method

      14 min read

    5. 2.5Imputing Missing Values: Mean vs Median vs Mode

      Subdomain 2.5: Compare and contrast imputing missing values with the mean or median or mode value

      16 min read

    6. 2.6Mean, Median, or Mode Imputation: Choosing the Fill Statistic

      Subdomain 2.6: Impute missing values with the mode, mean, or median value

      2 pages · 18 min read

    7. 2.7One-Hot Encoding Categorical Features with OneHotEncoder

      Subdomain 2.7: Use one-hot encoding for categorical features

      2 pages · 17 min read

    8. 2.8One-Hot Encoding: When to Use It and When Not To

      Subdomain 2.8: Identify and explain the model types or data sets for which one-hot encoding is or is not appropriate.

      14 min read

    9. 2.9Log Transformation: Recognizing Skewed and Multiplicative Data

      Subdomain 2.9: Identify scenarios where log scale transformation is appropriate

      2 pages · 17 min read

    Domain 3: Model Development

    15 lessons · 31% of the exam

    1. 3.1Algorithm Selection: Linear, Logistic and Tree-Based Models

      Subdomain 3.1: Use ML foundations to select the appropriate algorithm for a given model scenario

      12 min read

    2. 3.2Mitigating Class Imbalance: Resampling, SMOTE and Class Weights

      Subdomain 3.2: Identify methods to mitigate data imbalance in training data

      14 min read

    3. 3.3Estimators vs transformers: the fit, transform and predict contracts

      Subdomain 3.3: Compare estimators and transformers

      12 min read

    4. 3.4Building a Model Training Pipeline on Databricks: Structure, scikit-learn Pipelines, and Metrics

      Subdomain 3.4: Develop a training pipeline

      2 pages · 16 min read

    5. 3.5Hyperopt fmin: Objective Function, Search Space and Search Algorithm

      Subdomain 3.5: Use Hyperopt's fmin operation to tune a model's hyperparameters

      2 pages · 18 min read

    6. 3.6Grid, random and Bayesian hyperparameter search compared

      Subdomain 3.6: Perform random or grid search or Bayesian search as a method for tuning hyperparameters.

      2 pages · 18 min read

    7. 3.7Parallelize single-node hyperparameter tuning with Hyperopt SparkTrials

      Subdomain 3.7: Parallelize single node models for hyperparameter tuning

      17 min read

    8. 3.8Cross-Validation vs Train-Validation Split: Tradeoffs

      Subdomain 3.8: Describe the benefits and downsides of using cross-validation over a train-validation split.

      13 min read

    9. 3.9k-Fold Cross-Validation with cross_val_score

      Subdomain 3.9: Perform cross-validation as a part of model fitting.

      2 pages · 18 min read

    10. 3.10Counting Models in Grid Search with Cross-Validation

      Subdomain 3.10: Identify the number of models being trained in conjunction with a grid-search and cross-validation process.

      11 min read

    11. 3.11Classification Metrics: Precision, Recall, F1, ROC AUC and Log Loss

      Subdomain 3.11: Use common classification metrics: F1, Log Loss, ROC/AUC, etc

      20 min read

    12. 3.12Regression error metrics: MSE, RMSE, MAE and their variants

      Subdomain 3.12: Use common regression metrics: RMSE, MAE, R-squared, etc.

      2 pages · 19 min read

    13. 3.13Choosing the Right ML Metric for the Scenario Objective

      Subdomain 3.13: Choose the most appropriate metric for a given scenario objective

      14 min read

    14. 3.14Exponentiating Log-Transformed Targets Before Metrics and Interpretation

      Subdomain 3.14: Identify the need to exponentiate log-transformed variables before calculating evaluation metrics or interpreting predictions

      12 min read

    15. 3.15Bias-Variance Tradeoff and Model Complexity

      Subdomain 3.15: Assess the impact of model complexity and the bias variance tradeoff on model performance

      13 min read

    Domain 4: Model Deployment

    6 lessons · 12% of the exam

    1. 4.1Batch vs Streaming vs Real-Time Inference: Latency, Cost and Semantics

      Subdomain 4.1: Identify the differences and advantages of model serving approaches: batch, realtime, and streaming

      2 pages · 15 min read

    2. 4.2Deploy a Custom MLflow Model to a Databricks Model Serving Endpoint

      Subdomain 4.2: Deploy a custom model to a model endpoint

      2 pages · 21 min read

    3. 4.3Batch Inference with MLflow pyfunc Models and Spark

      Subdomain 4.3: Use pandas to perform batch inference

      2 pages · 23 min read

    4. 4.4Lakeflow Pipelines (formerly DLT) for Streaming Inference: Streaming Tables and Pipeline Modes

      Subdomain 4.4: Identify how streaming inference is performed with Delta Live Tables

      2 pages · 19 min read

    5. 4.5Deploy a Custom Model to a Databricks Model Serving Endpoint

      Subdomain 4.5: Deploy and query a model for realtime inference

      2 pages · 21 min read

    6. 4.6Traffic Splitting on Databricks Model Serving Endpoints

      Subdomain 4.6: Split data between endpoints for realtime interference

      17 min read