CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 3 · Lesson 38/48

    Classification Metrics: Precision, Recall, F1, ROC AUC and Log Loss

    Use common classification metrics: F1, Log Loss, ROC/AUC, etc

    20 min read
    2.08% of exam
    6 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Calculate precision and recall from confusion-matrix counts, and tell which one suffers when a model over-flags or misses positives
    • Explain F1 as a harmonic mean, and choose between macro, weighted and micro averaging for multiclass problems
    • Read a ROC curve as TPR against FPR across thresholds, and compute ROC AUC from predict_proba scores, not hard labels
    • Explain why log loss scores probabilities, and why two models with the same accuracy can differ on it
    • Pick a classification metric for a scenario and connect it to scikit-learn scoring, AutoML primary_metric, or a tuning objective

    Key concept

    Label metrics vs. score metrics — Some classification metrics (accuracy, precision, recall, F1) compare hard predicted labels with the truth, so they change when you move the decision threshold. Others (ROC AUC, log loss) take the model's continuous scores or probabilities, so they judge the ranking or calibration across every threshold at once.

    1.Precision and recall: counting the four outcomes

    Most classifiers produce a probability first and a label second. In the Databricks TabFM tutorial below, the model calls predict_proba, and the label is just that probability compared with 0.5. Accuracy is computed from the thresholded labels (clf_pred). ROC AUC is computed from the raw positive-class probabilities (clf_pred_proba[:, 1]). Keep that split in mind, because the rest of this lesson builds on it.

    One model, two kinds of metric: accuracy on thresholded labels, ROC AUC on positive-class probabilitiespython
    tabfm_clf.fit(X_train_clf, y_train_clf)
    clf_pred_proba = np.asarray(tabfm_clf.predict_proba(X_test_clf))
    clf_pred = (clf_pred_proba[:, 1] >= 0.5).astype(int)
    
    clf_results = pd.DataFrame({
        "actual": y_test_clf.reset_index(drop=True),
        "predicted": clf_pred.astype(int),
        "positive_class_probability": clf_pred_proba[:, 1],
    })
    
    accuracy = accuracy_score(y_test_clf, clf_pred)
    roc_auc = roc_auc_score(y_test_clf, clf_pred_proba[:, 1])

    Once you have labels, every prediction falls into one of four cells: true positive (tp), false positive (fp), true negative (tn) or false negative (fn). confusion_matrix counts these cells. multilabel_confusion_matrix gives one 2×2 table per class. Two ratios of these counts come up again and again.

    Precision is tp / (tp + fp). In scikit-learn's words, it is the classifier's ability not to label a negative sample as positive. Recall is tp / (tp + fn), the ability to find all the positive samples. Recall has two other names: true positive rate (TPR) and sensitivity. Two related rates come from the negatives. Specificity is tn / (tn + fp). *Fall-out* is fp / (fp + tn), better known as the false positive rate (FPR). TPR and FPR become the two axes of the ROC curve later in this lesson.

    Per-class recall (TPR) from a multilabel confusion matrixpython
    >>> mcm = multilabel_confusion_matrix(y_true, y_pred)
    >>> tn = mcm[:, 0, 0]
    >>> tp = mcm[:, 1, 1]
    >>> fn = mcm[:, 1, 0]
    >>> fp = mcm[:, 0, 1]
    >>> tp / (tp + fn)
    array([1. , 0.5, 0. ])

    Precision and recall are always measured relative to a positive class. For binary metrics, scikit-learn by default evaluates only the positive label and assumes it is labelled 1. You can change that with pos_label. Databricks AutoML's classify has its own pos_label parameter, described as useful for calculating metrics such as precision and recall. It should only be set for binary problems.

    Checkpoint 1 of 8· Check yourself

    A fraud model catches nearly every real fraud case but also flags many legitimate transactions. Which metric does this hurt most?

    Sources12

    2.F1: combining precision and recall, then averaging across classes

    In scikit-learn, f1_score computes the F1 score, "also known as balanced F-score or F-measure." The F-measure is a weighted harmonic mean of precision and recall. It is best at 1 and worst at 0. fbeta_score is the general form, and setting beta to 1 gives F1. A harmonic mean is pulled toward the smaller of its two inputs. So a model can't get a high F1 by maxing out recall while precision collapses, or the other way round.

    F1 matters on Databricks too. It is the default value of AutoML's primary_metric, the metric AutoML uses "to evaluate and rank model performance."

    Worked example, with F-beta. F1 is 2·P·R / (P + R). The scikit-learn example has y_true = [0, 1, 0, 1] and y_pred = [0, 1, 0, 0]. Precision is 1.0 and recall is 0.5, so F1 = 2·1.0·0.5 / 1.5 ≈ 0.66. Beta changes how much precision counts in the mean. On the same predictions, fbeta_score with beta=0.5 gives 0.83, and with beta=2 it gives 0.55. Because precision (1.0) is higher than recall (0.5) here, a beta below 1 (leaning toward precision) scores higher and a beta above 1 (leaning toward recall) scores lower. So pick beta above 1 when missing positives is the costlier mistake, and below 1 when false alarms are.

    F1 and ROC AUC are basically binary metrics. For multiclass or multilabel data, scikit-learn treats the problem as one binary problem per class and then averages the per-class scores. You choose how with the average parameter:

    - **macro: the plain mean of the per-class scores, with every class weighted equally. This highlights rare classes when they matter, but it can over-weight their usually poor performance when classes are not equally important. - weighted: each class's score is weighted by how often that class appears in the true labels. This accounts for class imbalance. - micro: adds up the numerators and denominators across all sample–class pairs, then divides once. Useful in multilabel settings, or in multiclass problems where you want to ignore a majority class. - samples: multilabel only. Computes the metric for each sample, then averages. - None**: returns one score per class.

    Databricks data profiling stores classification metrics the same way. For an InferenceLog classification profile, precision, recall and f1_score are each a struct with one_vs_all per-class values plus macro and weighted averages.

    Checkpoint 2 of 8· Match them up

    Match each average setting to what it does

    Tap a term, then the definition that fits it.

    Checkpoint 3 of 8· Exam question

    A data scientist at a payments company trains a classifier to flag fraudulent transactions in a dataset where only 2% of transactions are fraudulent. The model reports 98% accuracy, but the fraud team is skeptical that this reflects real detection ability. Which metric should the team report instead to better capture how well the model balances catching fraud against raising false alarms?

    Sources132

    3.ROC curve and AUC: scoring the ranking across all thresholds

    Precision, recall and F1 describe one threshold. The ROC (receiver operating characteristic) curve sweeps all of them. At each threshold it plots TPR (recall) against FPR (fall-out). scikit-learn's roc_curve and roc_auc_score both take y_score, not y_pred. roc_auc_score computes the area under that curve "from prediction scores" and summarizes the curve in one number. roc_curve is restricted to the binary case. roc_auc_score also supports multiclass.

    Databricks has a worked example: a Python UDTF that wraps roc_curve and roc_auc_score and runs them over 14 scored rows (6 positives, 8 negatives). A few of the ROC points it yields:

    ROC points from the Databricks UDTF example: lowering the threshold raises TPR and FPR together
    thresholdtrue_positive_ratefalse_positive_rate
    0.950.1670.0
    0.820.50.0
    0.520.8330.375
    0.311.00.625
    0.031.01.0

    Before the first threshold on the curve, nothing is called positive, so the curve starts at (FPR 0, TPR 0). At the lowest threshold everything is positive, giving (1, 1). Between them, each step down trades extra true positives for extra false positives. The exact value of the very first threshold depends on the scikit-learn version. The current roc_curve example shows inf there (array([ inf, 0.8 , 0.4 , 0.35, 0.1 ])), so don't memorize a particular first threshold.

    AUC is one number for the whole curve, not a per-threshold value. Be careful with the single auc value the Databricks page prints, 0.786. Check it against the 14 listed rows. The 6 positives (scores 0.95, 0.87, 0.82, 0.71, 0.52, 0.31) each outrank 8, 8, 8, 7, 5 and 3 of the 8 negatives respectively. That is 39 of 48 positive–negative pairs, and 39/48 = 0.8125. This matches the trapezoid area under the listed TPR/FPR points. The lesson: compute AUC yourself from the data rather than trusting a printed number. The UDTF collects every row before computing, a pattern the docs describe as useful for metrics that require the complete dataset.

    The Databricks getting-started notebook shows the usual scikit-learn pattern. Take column 1 of predict_proba (the positive-class probability), pass it to roc_auc_score, and log the result to MLflow. Autologging does not record test AUC for you.

    ROC AUC from positive-class probabilities, logged to MLflow (Databricks getting-started notebook)python
    with mlflow.start_run(run_name='gradient_boost') as run:
        model = sklearn.ensemble.GradientBoostingClassifier(random_state=0)
    
        # Models, parameters, and training metrics are tracked automatically
        model.fit(X_train, y_train)
    
        predicted_probs = model.predict_proba(X_test)
        roc_auc = sklearn.metrics.roc_auc_score(y_test, predicted_probs[:,1])
        roc_curve = sklearn.metrics.RocCurveDisplay.from_estimator(model, X_test, y_test)
    
        # Save the ROC curve plot to a file
        roc_curve.figure_.savefig("roc_curve.png")
    
        # The AUC score on test data is not automatically logged, so log it manually
        mlflow.log_metric("test_auc", roc_auc)

    Checkpoint 4 of 8· Check yourself

    A colleague computes ROC AUC for the same fitted model twice: once with a decision threshold of 0.5 and once with 0.7. They expect two different AUC values. What is correct?

    Sources452

    4.Log loss: scoring the probabilities themselves

    ROC AUC only cares about how scores rank positives against negatives. Log loss goes a step further and scores the probability values themselves. scikit-learn describes it as "also called logistic regression loss or cross-entropy loss." It is defined on probability estimates and "can be used to evaluate the probability outputs (predict_proba) of a classifier instead of its discrete predictions." For a binary label, the per-sample log loss is the negative log-likelihood of the classifier given the true label: −[y·ln(p) + (1−y)·ln(1−p)], where p is the predicted probability of class 1. A confident wrong probability is therefore penalized much more than a hesitant one. It is a loss, so lower is better, and it is never negative. The log_loss function takes y_proba, not y_pred. A related metric, d2_log_loss_score, reports the fraction of log loss explained.

    How a confident mistake blows up the loss. Take a sample whose true label is 1. If the model gives it p = 0.9, the loss is −ln(0.9) ≈ 0.11. If it gives p = 0.4 (wrong side, hesitant), the loss is −ln(0.4) ≈ 0.92. If it gives p = 0.01 (wrong and very confident), the loss is −ln(0.01) ≈ 4.6. The hard label is equally wrong for the last two, but the loss is five times larger for the confident one. The scikit-learn example below averages the per-sample losses: (0.105 + 0.223 + 0.357 + 0.010) / 4 ≈ 0.1738.

    scikit-learn log_loss on predict_proba-style probabilitiespython
    >>> from sklearn.metrics import log_loss
    >>> y_true = [0, 0, 1, 1]
    >>> y_proba = [[.9, .1], [.8, .2], [.3, .7], [.01, .99]]
    >>> log_loss(y_true, y_proba)
    0.1738

    scikit-learn's table of "strictly consistent" scoring functions makes the contrast precise. For classification, log loss and the Brier score both evaluate predict_proba outputs. Zero-one loss evaluates categorical predict outputs and is only consistent, not strictly consistent. It "is equivalent to one minus the accuracy score," so it ranks models the same way accuracy does. That answers the prediction above. Two models with identical labels have identical accuracy, but if their probabilities differ, log loss can still separate them.

    Because log loss is a loss, scikit-learn's scoring convention applies: scorers are higher-is-better, so loss metrics are exposed in negated form (like 'neg_mean_squared_error'). On Databricks, "log_loss" is one of AutoML's supported classification values for primary_metric, alongside f1, precision, accuracy and roc_auc.

    Checkpoint 5 of 8· Check yourself

    Which classifier output does log loss evaluate?

    Checkpoint 6 of 8· Exam question

    Two churn-prediction models score identically on accuracy, but the team wants to pick the one whose predicted probabilities are best calibrated, penalizing predictions that are both wrong and overconfident. Which metric should the team compute on the held-out predictions to make this comparison?

    Sources12

    5.Choosing a metric and getting its direction right

    The Databricks ML lifecycle guidance says to define evaluation metrics during development, based on your scoping requirements. These can be common metrics like accuracy or AUC, or domain-specific ones. The table below sums up the previous sections in terms of what each metric needs as input.

    What each classification metric consumes in scikit-learn
    Metricscikit-learn functionInput it needsThreshold-dependent?
    Accuracyaccuracy_score(y_true, y_pred)Hard labels from predictYes
    Precision / Recallprecision_score / recall_score(y_true, y_pred)Hard labels; defined relative to pos_labelYes
    F1f1_score(y_true, y_pred)Hard labels; average= for multiclassYes
    ROC AUCroc_auc_score(y_true, y_score)Scores, e.g. predict_proba[:,1]No: summarizes all thresholds
    Log losslog_loss(y_true, y_proba)Probabilities from predict_probaNo
    Balanced accuracybalanced_accuracy_score(y_true, y_pred)Hard labelsYes

    Imbalance. Plain accuracy can look good on an imbalanced test set just by predicting the majority class. balanced_accuracy_score "avoids inflated performance estimates on imbalanced datasets." It is the macro-average of per-class recall. If raw accuracy is above chance only because of the imbalance, balanced accuracy drops to 1/n_classes. Dummy estimators give you a baseline score for random predictions to compare against.

    Direction. Tools that cross-validate internally, such as GridSearchCV, take a scoring parameter. It can be None (the estimator's default score), a string name, a callable, or for some tools several metrics at once, e.g. ['accuracy', 'precision']. All scorer objects follow one rule: higher return values are better. Distance-style metrics are therefore exposed in negated form, such as 'neg_mean_squared_error'. Minimizing optimizers need the opposite flip. In the Databricks Optuna example, the objective returns -roc_auc because Optuna minimizes by default.

    Precision-recall curve and average precision. When the positive class is rare and you care about finding it, look at precision against recall rather than only at ROC. precision_recall_curve computes precision-recall pairs by varying the decision threshold (binary only). average_precision_score summarizes that curve from prediction scores as a value between 0 and 1, higher is better, and it also supports multiclass and multilabel via one-vs-rest. Its baseline depends on prevalence: "with random predictions, the AP is the fraction of positive samples." So when positives are rare, a random scorer gets a very low AP, and a good AP is a demanding result. scikit-learn also cites work noting that linear interpolation of points on the precision-recall curve (as auc does with the trapezoidal rule) gives an overly optimistic measure, so use average_precision_score rather than trapezoidal auc on this curve.

    precision_recall_curve and average_precision_score on scores (scikit-learn docs)python
    >>> from sklearn.metrics import precision_recall_curve
    >>> from sklearn.metrics import average_precision_score
    >>> y_true = np.array([0, 0, 1, 1])
    >>> y_scores = np.array([0.1, 0.4, 0.35, 0.8])
    >>> precision, recall, threshold = precision_recall_curve(y_true, y_scores)

    Checkpoint 7 of 8· Fill the gap

    This Optuna objective tunes on test AUC. Which method fills the blank so that roc_auc_score gets the scores it needs?

        predicted_probs = model_hp. ? (X_test)
        # Tune based on the test AUC
        # In production, you could use a separate validation set instead
        roc_auc = sklearn.metrics.roc_auc_score(y_test, predicted_probs[:,1])
        mlflow.log_metric('test_auc', roc_auc)
    
        # Negate the AUC because Optuna minimizes the objective by default
        return -roc_auc

    Checkpoint 8 of 8· Exam question

    A hospital builds a classifier to flag a rare condition that occurs in fewer than 1% of screened patients. Two candidate models have nearly identical ROC-AUC scores, but the clinical team wants a metric that will more clearly separate them given how rare the positive class is. Which evaluation should they add?

    Sources642

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.ROC AUC can be computed just as well from the 0/1 labels returned by predict().Why is that wrong?

      ROC AUC is computed from continuous prediction scores, usually predict_proba[:,1]. With hard labels it collapses to balanced accuracy at a single threshold.

      Covered in ROC curve and AUC: scoring the ranking across all thresholds

    2. 2.Macro-averaged F1 is always the neutral choice for an imbalanced multiclass problem (or, the opposite, weighted is always the safe one).Why is that wrong?

      Neither is universally safe. Macro gives every class equal weight, which highlights rare classes that matter but over-emphasizes their typically low performance when all classes are not equally important. Weighted averaging weights each class by how often it occurs, so it accounts for imbalance. Pick by what the scenario cares about.

      Covered in F1: combining precision and recall, then averaging across classes

    3. 3.A high accuracy on an imbalanced test set shows the model has learned the minority class.Why is that wrong?

      Accuracy can be inflated just by predicting the majority class. Balanced accuracy, the mean of per-class recall, drops to chance level when that happens.

      Covered in Choosing a metric and getting its direction right

    4. 4.Because a loss is lower-is-better, a scikit-learn scorer for a loss returns the raw loss and GridSearchCV picks the minimum.Why is that wrong?

      Every scikit-learn scorer is higher-is-better, so distance and loss metrics are exposed in negated form. Minimizing tuners such as Optuna need the opposite flip, for example returning -roc_auc.

      Covered in Choosing a metric and getting its direction right

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “This is useful for calculating metrics such as precision and recall.”
      ↩︎ Precision and recall: counting the four outcomes
      “Metric used to evaluate and rank model performance.”
      ↩︎ F1: combining precision and recall, then averaging across classes
      “Supported metrics for classification: “f1” (default), “log_loss”, “precision”, “accuracy”, “roc_auc””
      ↩︎ Log loss: scoring the probabilities themselves
    2. 2.
      “recall is the ability of the classifier to find all the positive samples.”
      ↩︎ Precision and recall: counting the four outcomes
      “Calculating recall (also called the true positive rate or the sensitivity) for each class:”
      ↩︎ Precision and recall: counting the four outcomes
      “assuming by default that the positive class is labelled 1”
      ↩︎ Precision and recall: counting the four outcomes
      “can be interpreted as a weighted harmonic mean of the precision and recall.”
      ↩︎ F1: combining precision and recall, then averaging across classes
      “Compute the F1 score, also known as balanced F-score or F-measure.”
      ↩︎ F1: combining precision and recall, then averaging across classes
      “each class’s score is weighted by its presence in the true data sample”
      ↩︎ F1: combining precision and recall, then averaging across classes
      “Micro-averaging may be preferred in multilabel settings, including multiclass classification where a majority class is to be ignored.”
      ↩︎ F1: combining precision and recall, then averaging across classes
      “In problems where infrequent classes are nonetheless important, macro-averaging may be a means of highlighting their performance.”
      ↩︎ F1: combining precision and recall, then averaging across classes
      “F-measure is the weighted harmonic mean of precision and recall”
      ↩︎ F1: combining precision and recall, then averaging across classes
      “Compute Area Under the Receiver Operating Characteristic Curve (ROC AUC) from prediction scores.”
      ↩︎ ROC curve and AUC: scoring the ranking across all thresholds
      “Calculating fall out (also called the false positive rate) for each class:”
      ↩︎ ROC curve and AUC: scoring the ranking across all thresholds
      “Log loss, also called logistic regression loss or cross-entropy loss, is defined on probability estimates.”
      ↩︎ Log loss: scoring the probabilities themselves
      “the log loss per sample is the negative log-likelihood of the classifier given the true label”
      ↩︎ Log loss: scoring the probabilities themselves
      “The zero-one loss is equivalent to one minus the accuracy score, meaning it gives different score values but the same ranking.”
      ↩︎ Log loss: scoring the probabilities themselves
      “All scorer objects follow the convention that higher return values are better than lower return values.”
      ↩︎ Choosing a metric and getting its direction right
      “Finally, Dummy estimators are useful to get a baseline value of those metrics for random predictions.”
      ↩︎ Choosing a metric and getting its direction right
      “With random predictions, the AP is the fraction of positive samples.”
      ↩︎ Choosing a metric and getting its direction right
      “Compute precision-recall pairs for different probability thresholds.”
      ↩︎ Choosing a metric and getting its direction right
      “Some metrics might require probability estimates of the positive class or non-thresholded decision values”
      ↩︎ Key concept
      “or the area under the ROC curve with binary predictions rather than scores”
      ↩︎ Exam trap 1
      “macro-averaging will over-emphasize the typically low performance on an infrequent class.”
      ↩︎ Exam trap 2
      “The balanced_accuracy_score function computes the balanced accuracy, which avoids inflated performance estimates on imbalanced datasets.”
      ↩︎ Exam trap 3
      “All scorer objects follow the convention that higher return values are better than lower return values.”
      ↩︎ Exam trap 4
      “Intuitively, precision is the ability of the classifier not to label as positive a sample that is negative”
      ↩︎ Checkpoint
      “simply calculates the mean of the binary metrics, giving equal weight to each class.”
      ↩︎ Checkpoint
      “or the area under the ROC curve with binary predictions rather than scores”
      ↩︎ Prediction
      “can be used to evaluate the probability outputs (predict_proba) of a classifier instead of its discrete predictions.”
      ↩︎ Checkpoint
    3. 4.
      “The AUC score on test data is not automatically logged, so log it manually”
      ↩︎ ROC curve and AUC: scoring the ranking across all thresholds
      “Negate the AUC because Optuna minimizes the objective by default”
      ↩︎ Choosing a metric and getting its direction right
    4. 5.
      “This pattern is useful for metrics that require the complete dataset for calculation.”
      ↩︎ ROC curve and AUC: scoring the ranking across all thresholds
    5. 6.
      “Your base metrics might be common ML metrics like accuracy, AUC, RMSE or domain-specific metrics.”
      ↩︎ Choosing a metric and getting its direction right

    Ready to test yourself?

    Practise Databricks Certified Machine Learning Associate in quiz mode.

    Spotted a mistake, or was something unclear? Tell us.