What you will be able to do
- Calculate precision and recall from confusion-matrix counts, and tell which one suffers when a model over-flags or misses positives
- Explain F1 as a harmonic mean, and choose between macro, weighted and micro averaging for multiclass problems
- Read a ROC curve as TPR against FPR across thresholds, and compute ROC AUC from predict_proba scores, not hard labels
- Explain why log loss scores probabilities, and why two models with the same accuracy can differ on it
- Pick a classification metric for a scenario and connect it to scikit-learn scoring, AutoML primary_metric, or a tuning objective
Key concept
Label metrics vs. score metrics — Some classification metrics (accuracy, precision, recall, F1) compare hard predicted labels with the truth, so they change when you move the decision threshold. Others (ROC AUC, log loss) take the model's continuous scores or probabilities, so they judge the ranking or calibration across every threshold at once.
1.Precision and recall: counting the four outcomes
Most classifiers produce a probability first and a label second. In the Databricks TabFM tutorial below, the model calls predict_proba, and the label is just that probability compared with 0.5. Accuracy is computed from the thresholded labels (clf_pred). ROC AUC is computed from the raw positive-class probabilities (clf_pred_proba[:, 1]). Keep that split in mind, because the rest of this lesson builds on it.
tabfm_clf.fit(X_train_clf, y_train_clf)
clf_pred_proba = np.asarray(tabfm_clf.predict_proba(X_test_clf))
clf_pred = (clf_pred_proba[:, 1] >= 0.5).astype(int)
clf_results = pd.DataFrame({
"actual": y_test_clf.reset_index(drop=True),
"predicted": clf_pred.astype(int),
"positive_class_probability": clf_pred_proba[:, 1],
})
accuracy = accuracy_score(y_test_clf, clf_pred)
roc_auc = roc_auc_score(y_test_clf, clf_pred_proba[:, 1])Once you have labels, every prediction falls into one of four cells: true positive (tp), false positive (fp), true negative (tn) or false negative (fn). confusion_matrix counts these cells. multilabel_confusion_matrix gives one 2×2 table per class. Two ratios of these counts come up again and again.
Precision is tp / (tp + fp). In scikit-learn's words, it is the classifier's ability not to label a negative sample as positive. Recall is tp / (tp + fn), the ability to find all the positive samples. Recall has two other names: true positive rate (TPR) and sensitivity. Two related rates come from the negatives. Specificity is tn / (tn + fp). *Fall-out* is fp / (fp + tn), better known as the false positive rate (FPR). TPR and FPR become the two axes of the ROC curve later in this lesson.
>>> mcm = multilabel_confusion_matrix(y_true, y_pred)
>>> tn = mcm[:, 0, 0]
>>> tp = mcm[:, 1, 1]
>>> fn = mcm[:, 1, 0]
>>> fp = mcm[:, 0, 1]
>>> tp / (tp + fn)
array([1. , 0.5, 0. ])Precision and recall are always measured relative to a positive class. For binary metrics, scikit-learn by default evaluates only the positive label and assumes it is labelled 1. You can change that with pos_label. Databricks AutoML's classify has its own pos_label parameter, described as useful for calculating metrics such as precision and recall. It should only be set for binary problems.
tp = 60, fp = 20, fn = 40. Precision = 60 / 80 = 0.75. Recall = 60 / (60 + 40) = 0.60.
Checkpoint 1 of 8· Check yourself
A fraud model catches nearly every real fraud case but also flags many legitimate transactions. Which metric does this hurt most?
Flagging legitimate transactions creates false positives, which lowers precision. Catching nearly all fraud means few false negatives, so recall stays high.
“Intuitively, precision is the ability of the classifier not to label as positive a sample that is negative”Source: scikit-learn.org
2.F1: combining precision and recall, then averaging across classes
In scikit-learn, f1_score computes the F1 score, "also known as balanced F-score or F-measure." The F-measure is a weighted harmonic mean of precision and recall. It is best at 1 and worst at 0. fbeta_score is the general form, and setting beta to 1 gives F1. A harmonic mean is pulled toward the smaller of its two inputs. So a model can't get a high F1 by maxing out recall while precision collapses, or the other way round.
F1 matters on Databricks too. It is the default value of AutoML's primary_metric, the metric AutoML uses "to evaluate and rank model performance."
Worked example, with F-beta. F1 is 2·P·R / (P + R). The scikit-learn example has y_true = [0, 1, 0, 1] and y_pred = [0, 1, 0, 0]. Precision is 1.0 and recall is 0.5, so F1 = 2·1.0·0.5 / 1.5 ≈ 0.66. Beta changes how much precision counts in the mean. On the same predictions, fbeta_score with beta=0.5 gives 0.83, and with beta=2 it gives 0.55. Because precision (1.0) is higher than recall (0.5) here, a beta below 1 (leaning toward precision) scores higher and a beta above 1 (leaning toward recall) scores lower. So pick beta above 1 when missing positives is the costlier mistake, and below 1 when false alarms are.
F1 and ROC AUC are basically binary metrics. For multiclass or multilabel data, scikit-learn treats the problem as one binary problem per class and then averages the per-class scores. You choose how with the average parameter:
- **macro: the plain mean of the per-class scores, with every class weighted equally. This highlights rare classes when they matter, but it can over-weight their usually poor performance when classes are not equally important.
- weighted: each class's score is weighted by how often that class appears in the true labels. This accounts for class imbalance.
- micro: adds up the numerators and denominators across all sample–class pairs, then divides once. Useful in multilabel settings, or in multiclass problems where you want to ignore a majority class.
- samples: multilabel only. Computes the metric for each sample, then averages.
- None**: returns one score per class.
Databricks data profiling stores classification metrics the same way. For an InferenceLog classification profile, precision, recall and f1_score are each a struct with one_vs_all per-class values plus macro and weighted averages.
Checkpoint 2 of 8· Match them up
Match each average setting to what it does
Tap a term, then the definition that fits it.
Macro weights every class equally. Weighted uses class frequency. Micro pools the counts before dividing. None skips averaging and returns per-class scores.
“simply calculates the mean of the binary metrics, giving equal weight to each class.”Source: scikit-learn.org
Checkpoint 3 of 8· Exam question
A data scientist at a payments company trains a classifier to flag fraudulent transactions in a dataset where only 2% of transactions are fraudulent. The model reports 98% accuracy, but the fraud team is skeptical that this reflects real detection ability. Which metric should the team report instead to better capture how well the model balances catching fraud against raising false alarms?
Correct answer: A — F1 score is the right choice because it combines precision and recall into a single harmonic-mean value that only rises when the model both catches real fraud and avoids flagging too many legitimate transactions.
- A. The F1 score is the harmonic mean of precision and recall, so it stays low unless the model both catches a meaningful share of true fraud cases and keeps false alarms in check. That combination makes it far more informative than a raw correctness percentage on a dataset this skewed.
- B. Accuracy does not account for class imbalance at all; a model that predicts every transaction as legitimate would still score 98% accuracy on this dataset while catching zero fraud. It is misleading here precisely because it rewards ignoring the minority class.
- C. Root mean squared error is a regression metric for continuous numeric targets, not the standard tool for evaluating a fraud/not-fraud classification decision. It does not summarize precision or recall tradeoffs the fraud team cares about.
- D. R-squared measures explained variance for continuous outcomes and is not a standard way to evaluate a binary classifier's ability to separate fraud from legitimate transactions. It gives no direct read on false positives or false negatives.
3.ROC curve and AUC: scoring the ranking across all thresholds
Precision, recall and F1 describe one threshold. The ROC (receiver operating characteristic) curve sweeps all of them. At each threshold it plots TPR (recall) against FPR (fall-out). scikit-learn's roc_curve and roc_auc_score both take y_score, not y_pred. roc_auc_score computes the area under that curve "from prediction scores" and summarizes the curve in one number. roc_curve is restricted to the binary case. roc_auc_score also supports multiclass.
Databricks has a worked example: a Python UDTF that wraps roc_curve and roc_auc_score and runs them over 14 scored rows (6 positives, 8 negatives). A few of the ROC points it yields:
| threshold | true_positive_rate | false_positive_rate |
|---|---|---|
| 0.95 | 0.167 | 0.0 |
| 0.82 | 0.5 | 0.0 |
| 0.52 | 0.833 | 0.375 |
| 0.31 | 1.0 | 0.625 |
| 0.03 | 1.0 | 1.0 |
Before the first threshold on the curve, nothing is called positive, so the curve starts at (FPR 0, TPR 0). At the lowest threshold everything is positive, giving (1, 1). Between them, each step down trades extra true positives for extra false positives. The exact value of the very first threshold depends on the scikit-learn version. The current roc_curve example shows inf there (array([ inf, 0.8 , 0.4 , 0.35, 0.1 ])), so don't memorize a particular first threshold.
AUC is one number for the whole curve, not a per-threshold value. Be careful with the single auc value the Databricks page prints, 0.786. Check it against the 14 listed rows. The 6 positives (scores 0.95, 0.87, 0.82, 0.71, 0.52, 0.31) each outrank 8, 8, 8, 7, 5 and 3 of the 8 negatives respectively. That is 39 of 48 positive–negative pairs, and 39/48 = 0.8125. This matches the trapezoid area under the listed TPR/FPR points. The lesson: compute AUC yourself from the data rather than trusting a printed number. The UDTF collects every row before computing, a pattern the docs describe as useful for metrics that require the complete dataset.
The Databricks getting-started notebook shows the usual scikit-learn pattern. Take column 1 of predict_proba (the positive-class probability), pass it to roc_auc_score, and log the result to MLflow. Autologging does not record test AUC for you.
with mlflow.start_run(run_name='gradient_boost') as run:
model = sklearn.ensemble.GradientBoostingClassifier(random_state=0)
# Models, parameters, and training metrics are tracked automatically
model.fit(X_train, y_train)
predicted_probs = model.predict_proba(X_test)
roc_auc = sklearn.metrics.roc_auc_score(y_test, predicted_probs[:,1])
roc_curve = sklearn.metrics.RocCurveDisplay.from_estimator(model, X_test, y_test)
# Save the ROC curve plot to a file
roc_curve.figure_.savefig("roc_curve.png")
# The AUC score on test data is not automatically logged, so log it manually
mlflow.log_metric("test_auc", roc_auc)Checkpoint 4 of 8· Check yourself
A colleague computes ROC AUC for the same fitted model twice: once with a decision threshold of 0.5 and once with 0.7. They expect two different AUC values. What is correct?
ROC AUC takes continuous scores and summarizes TPR against FPR over every threshold. Precision, recall and F1 change with the threshold; AUC does not.
“Compute Area Under the Receiver Operating Characteristic Curve (ROC AUC) from prediction scores.”Source: scikit-learn.org
4.Log loss: scoring the probabilities themselves
ROC AUC only cares about how scores rank positives against negatives. Log loss goes a step further and scores the probability values themselves. scikit-learn describes it as "also called logistic regression loss or cross-entropy loss." It is defined on probability estimates and "can be used to evaluate the probability outputs (predict_proba) of a classifier instead of its discrete predictions." For a binary label, the per-sample log loss is the negative log-likelihood of the classifier given the true label: −[y·ln(p) + (1−y)·ln(1−p)], where p is the predicted probability of class 1. A confident wrong probability is therefore penalized much more than a hesitant one. It is a loss, so lower is better, and it is never negative. The log_loss function takes y_proba, not y_pred. A related metric, d2_log_loss_score, reports the fraction of log loss explained.
How a confident mistake blows up the loss. Take a sample whose true label is 1. If the model gives it p = 0.9, the loss is −ln(0.9) ≈ 0.11. If it gives p = 0.4 (wrong side, hesitant), the loss is −ln(0.4) ≈ 0.92. If it gives p = 0.01 (wrong and very confident), the loss is −ln(0.01) ≈ 4.6. The hard label is equally wrong for the last two, but the loss is five times larger for the confident one. The scikit-learn example below averages the per-sample losses: (0.105 + 0.223 + 0.357 + 0.010) / 4 ≈ 0.1738.
>>> from sklearn.metrics import log_loss
>>> y_true = [0, 0, 1, 1]
>>> y_proba = [[.9, .1], [.8, .2], [.3, .7], [.01, .99]]
>>> log_loss(y_true, y_proba)
0.1738scikit-learn's table of "strictly consistent" scoring functions makes the contrast precise. For classification, log loss and the Brier score both evaluate predict_proba outputs. Zero-one loss evaluates categorical predict outputs and is only consistent, not strictly consistent. It "is equivalent to one minus the accuracy score," so it ranks models the same way accuracy does. That answers the prediction above. Two models with identical labels have identical accuracy, but if their probabilities differ, log loss can still separate them.
Because log loss is a loss, scikit-learn's scoring convention applies: scorers are higher-is-better, so loss metrics are exposed in negated form (like 'neg_mean_squared_error'). On Databricks, "log_loss" is one of AutoML's supported classification values for primary_metric, alongside f1, precision, accuracy and roc_auc.
Checkpoint 5 of 8· Check yourself
Which classifier output does log loss evaluate?
Log loss is defined on probability estimates. Being scored on discrete labels is what zero-one loss and accuracy do.
“can be used to evaluate the probability outputs (predict_proba) of a classifier instead of its discrete predictions.”Source: scikit-learn.org
Checkpoint 6 of 8· Exam question
Two churn-prediction models score identically on accuracy, but the team wants to pick the one whose predicted probabilities are best calibrated, penalizing predictions that are both wrong and overconfident. Which metric should the team compute on the held-out predictions to make this comparison?
Correct answer: B — Log loss should be computed, because it evaluates the predicted probability assigned to the true class and grows sharply for confident predictions that turn out to be wrong, not just for wrong predictions in general.
- A. A confusion matrix built from hard labels at one threshold throws away the underlying probability values entirely, so it cannot distinguish a confident wrong prediction from a barely-wrong one. That makes it the wrong tool for comparing probability calibration.
- B. Log loss takes the predicted probability assigned to the true class and applies a logarithmic penalty, so a confident but incorrect prediction is penalized far more heavily than a cautious, uncertain one. This directly rewards models whose probabilities are well calibrated rather than just their hard classifications.
- C. Precision at a fixed threshold still only looks at hard classification outcomes after the cutoff is applied, so it says nothing about how confident or overconfident the underlying probability estimates were. Two models can share identical precision while differing sharply in calibration.
- D. Cohen's kappa compares hard-label agreement against chance-level agreement, which is a useful check for classifier consistency but still collapses each prediction to a category before scoring. It does not capture how confidently the model made a given prediction.
5.Choosing a metric and getting its direction right
The Databricks ML lifecycle guidance says to define evaluation metrics during development, based on your scoping requirements. These can be common metrics like accuracy or AUC, or domain-specific ones. The table below sums up the previous sections in terms of what each metric needs as input.
| Metric | scikit-learn function | Input it needs | Threshold-dependent? |
|---|---|---|---|
| Accuracy | accuracy_score(y_true, y_pred) | Hard labels from predict | Yes |
| Precision / Recall | precision_score / recall_score(y_true, y_pred) | Hard labels; defined relative to pos_label | Yes |
| F1 | f1_score(y_true, y_pred) | Hard labels; average= for multiclass | Yes |
| ROC AUC | roc_auc_score(y_true, y_score) | Scores, e.g. predict_proba[:,1] | No: summarizes all thresholds |
| Log loss | log_loss(y_true, y_proba) | Probabilities from predict_proba | No |
| Balanced accuracy | balanced_accuracy_score(y_true, y_pred) | Hard labels | Yes |
Imbalance. Plain accuracy can look good on an imbalanced test set just by predicting the majority class. balanced_accuracy_score "avoids inflated performance estimates on imbalanced datasets." It is the macro-average of per-class recall. If raw accuracy is above chance only because of the imbalance, balanced accuracy drops to 1/n_classes. Dummy estimators give you a baseline score for random predictions to compare against.
Direction. Tools that cross-validate internally, such as GridSearchCV, take a scoring parameter. It can be None (the estimator's default score), a string name, a callable, or for some tools several metrics at once, e.g. ['accuracy', 'precision']. All scorer objects follow one rule: higher return values are better. Distance-style metrics are therefore exposed in negated form, such as 'neg_mean_squared_error'. Minimizing optimizers need the opposite flip. In the Databricks Optuna example, the objective returns -roc_auc because Optuna minimizes by default.
Precision-recall curve and average precision. When the positive class is rare and you care about finding it, look at precision against recall rather than only at ROC. precision_recall_curve computes precision-recall pairs by varying the decision threshold (binary only). average_precision_score summarizes that curve from prediction scores as a value between 0 and 1, higher is better, and it also supports multiclass and multilabel via one-vs-rest. Its baseline depends on prevalence: "with random predictions, the AP is the fraction of positive samples." So when positives are rare, a random scorer gets a very low AP, and a good AP is a demanding result. scikit-learn also cites work noting that linear interpolation of points on the precision-recall curve (as auc does with the trapezoidal rule) gives an overly optimistic measure, so use average_precision_score rather than trapezoidal auc on this curve.
>>> from sklearn.metrics import precision_recall_curve
>>> from sklearn.metrics import average_precision_score
>>> y_true = np.array([0, 0, 1, 1])
>>> y_scores = np.array([0.1, 0.4, 0.35, 0.8])
>>> precision, recall, threshold = precision_recall_curve(y_true, y_scores)Checkpoint 7 of 8· Fill the gap
This Optuna objective tunes on test AUC. Which method fills the blank so that roc_auc_score gets the scores it needs?
predicted_probs = model_hp. ? (X_test)
# Tune based on the test AUC
# In production, you could use a separate validation set instead
roc_auc = sklearn.metrics.roc_auc_score(y_test, predicted_probs[:,1])
mlflow.log_metric('test_auc', roc_auc)
# Negate the AUC because Optuna minimizes the objective by default
return -roc_aucpredicted_probs[:,1] slices the positive-class column, which only exists in the per-class probability array from predict_proba. predict returns hard labels, and ROC AUC on hard labels collapses to balanced accuracy.
Checkpoint 8 of 8· Exam question
A hospital builds a classifier to flag a rare condition that occurs in fewer than 1% of screened patients. Two candidate models have nearly identical ROC-AUC scores, but the clinical team wants a metric that will more clearly separate them given how rare the positive class is. Which evaluation should they add?
Correct answer: D — They should compute the precision-recall AUC, because with the positive class this rare, ROC-AUC can look similarly high despite very different false-positive counts, while precision and recall track that imbalance far more sharply.
- A. Overall accuracy is dominated by the huge majority of patients who do not have the condition, so a model that rarely flags anyone can still score very high accuracy while missing most true cases. It would not help distinguish two models on how well they handle the rare positive class.
- B. Specificity only describes correctness among patients without the condition and ignores how many true positive cases each model actually catches. Two models with very different sensitivity to the rare condition could still post similar specificity.
- C. Changing the train-test split addresses sampling variability rather than choosing a metric better suited to a rare positive class, and it does not change what ROC-AUC itself is measuring. The underlying insensitivity of ROC-AUC to extreme imbalance would remain.
- D. Because negatives vastly outnumber positives here, the false positive rate used by ROC-AUC can stay small even when the raw number of false positives is large relative to the few true positives, making ROC-AUC look deceptively similar across models. Precision and recall are computed directly against the small positive class, so the precision-recall curve reacts much more sharply to that difference.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.ROC AUC can be computed just as well from the 0/1 labels returned by predict().Why is that wrong?
ROC AUC is computed from continuous prediction scores, usually predict_proba[:,1]. With hard labels it collapses to balanced accuracy at a single threshold.
Covered in ROC curve and AUC: scoring the ranking across all thresholds
2.Macro-averaged F1 is always the neutral choice for an imbalanced multiclass problem (or, the opposite, weighted is always the safe one).Why is that wrong?
Neither is universally safe. Macro gives every class equal weight, which highlights rare classes that matter but over-emphasizes their typically low performance when all classes are not equally important. Weighted averaging weights each class by how often it occurs, so it accounts for imbalance. Pick by what the scenario cares about.
Covered in F1: combining precision and recall, then averaging across classes
3.A high accuracy on an imbalanced test set shows the model has learned the minority class.Why is that wrong?
Accuracy can be inflated just by predicting the majority class. Balanced accuracy, the mean of per-class recall, drops to chance level when that happens.
Covered in Choosing a metric and getting its direction right
4.Because a loss is lower-is-better, a scikit-learn scorer for a loss returns the raw loss and GridSearchCV picks the minimum.Why is that wrong?
Every scikit-learn scorer is higher-is-better, so distance and loss metrics are exposed in negated form. Minimizing tuners such as Optuna need the opposite flip, for example returning -roc_auc.
Covered in Choosing a metric and getting its direction right
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“This is useful for calculating metrics such as precision and recall.”
↩︎ Precision and recall: counting the four outcomes“primary_metric: str = "f1",”
↩︎ F1: combining precision and recall, then averaging across classes“Metric used to evaluate and rank model performance.”
↩︎ F1: combining precision and recall, then averaging across classes“Supported metrics for classification: “f1” (default), “log_loss”, “precision”, “accuracy”, “roc_auc””
↩︎ Log loss: scoring the probabilities themselves - 2.https://scikit-learn.org/stable/modules/model_evaluation.htmlSecondary source
“recall is the ability of the classifier to find all the positive samples.”
↩︎ Precision and recall: counting the four outcomes“Calculating recall (also called the true positive rate or the sensitivity) for each class:”
↩︎ Precision and recall: counting the four outcomes“assuming by default that the positive class is labelled 1”
↩︎ Precision and recall: counting the four outcomes“can be interpreted as a weighted harmonic mean of the precision and recall.”
↩︎ F1: combining precision and recall, then averaging across classes“Compute the F1 score, also known as balanced F-score or F-measure.”
↩︎ F1: combining precision and recall, then averaging across classes“each class’s score is weighted by its presence in the true data sample”
↩︎ F1: combining precision and recall, then averaging across classes“Micro-averaging may be preferred in multilabel settings, including multiclass classification where a majority class is to be ignored.”
↩︎ F1: combining precision and recall, then averaging across classes“In problems where infrequent classes are nonetheless important, macro-averaging may be a means of highlighting their performance.”
↩︎ F1: combining precision and recall, then averaging across classes“F-measure is the weighted harmonic mean of precision and recall”
↩︎ F1: combining precision and recall, then averaging across classes“Compute Area Under the Receiver Operating Characteristic Curve (ROC AUC) from prediction scores.”
↩︎ ROC curve and AUC: scoring the ranking across all thresholds“Calculating fall out (also called the false positive rate) for each class:”
↩︎ ROC curve and AUC: scoring the ranking across all thresholds“array([ inf, 0.8 , 0.4 , 0.35, 0.1 ])”
↩︎ ROC curve and AUC: scoring the ranking across all thresholds“Log loss, also called logistic regression loss or cross-entropy loss, is defined on probability estimates.”
↩︎ Log loss: scoring the probabilities themselves“the log loss per sample is the negative log-likelihood of the classifier given the true label”
↩︎ Log loss: scoring the probabilities themselves“The log loss is non-negative.”
↩︎ Log loss: scoring the probabilities themselves“The zero-one loss is equivalent to one minus the accuracy score, meaning it gives different score values but the same ranking.”
↩︎ Log loss: scoring the probabilities themselves“All scorer objects follow the convention that higher return values are better than lower return values.”
↩︎ Choosing a metric and getting its direction right“Finally, Dummy estimators are useful to get a baseline value of those metrics for random predictions.”
↩︎ Choosing a metric and getting its direction right“With random predictions, the AP is the fraction of positive samples.”
↩︎ Choosing a metric and getting its direction right“Compute precision-recall pairs for different probability thresholds.”
↩︎ Choosing a metric and getting its direction right“Some metrics might require probability estimates of the positive class or non-thresholded decision values”
↩︎ Key concept“or the area under the ROC curve with binary predictions rather than scores”
↩︎ Exam trap 1“macro-averaging will over-emphasize the typically low performance on an infrequent class.”
↩︎ Exam trap 2“The balanced_accuracy_score function computes the balanced accuracy, which avoids inflated performance estimates on imbalanced datasets.”
↩︎ Exam trap 3“All scorer objects follow the convention that higher return values are better than lower return values.”
↩︎ Exam trap 4“Intuitively, precision is the ability of the classifier not to label as positive a sample that is negative”
↩︎ Checkpoint“simply calculates the mean of the binary metrics, giving equal weight to each class.”
↩︎ Checkpoint“or the area under the ROC curve with binary predictions rather than scores”
↩︎ Prediction“can be used to evaluate the probability outputs (predict_proba) of a classifier instead of its discrete predictions.”
↩︎ Checkpoint - 3.
“Format of struct for confusion_matrix, precision, recall, f1_score, and roc_auc_score:”
↩︎ F1: combining precision and recall, then averaging across classes - 4.
“The AUC score on test data is not automatically logged, so log it manually”
↩︎ ROC curve and AUC: scoring the ranking across all thresholds“Negate the AUC because Optuna minimizes the objective by default”
↩︎ Choosing a metric and getting its direction right - 5.
“This pattern is useful for metrics that require the complete dataset for calculation.”
↩︎ ROC curve and AUC: scoring the ranking across all thresholds - 6.
“Your base metrics might be common ML metrics like accuracy, AUC, RMSE or domain-specific metrics.”
↩︎ Choosing a metric and getting its direction right