Skip to content
EvalSuite
Documentation menu

Calibration

Curve and ECE since v0.1.0; slope, intercept, MCE and Hosmer–Lemeshow since v0.2.0

A model is calibrated when, among cases given a predicted risk of 0.2, about 20% have the outcome. Discrimination metrics such as ROC AUC do not measure this.

Usage

Pythonv0.1.0
import evalsuite as es

es.expected_calibration_error(y_true, y_prob, n_bins=10, strategy="quantile")
prob_true, prob_pred, bin_weight = es.calibration_curve(y_true, y_prob, n_bins=10)
es.brier_score(y_true, y_prob)
es.plot.calibration(y_true, y_prob)   # reliability diagram with ECE and Brier score
Pythonv0.2.0
report = es.calibration_report(y_true, y_prob)   # Brier, ECE, MCE, intercept, slope, Hosmer–Lemeshow
print(report)

es.calibration_slope(y_true, y_prob)       # ideal 1; below 1 means predictions are too extreme
es.calibration_intercept(y_true, y_prob)   # ideal 0 (calibration-in-the-large)
es.maximum_calibration_error(y_true, y_prob)
hl = es.hosmer_lemeshow(y_true, y_prob, n_groups=10)
hl.statistic, hl.p_value, hl.params["df"]

The slope and intercept come from the logistic recalibration model logit P(y = 1) = a + b · logit(p̂) (Cox, 1958), fitted by maximum likelihood; the intercept is estimated with the slope fixed at 1. Both match statsmodels' logistic GLM in the test suite.

Interpreting calibration

No single number proves that a model is calibrated. Binned metrics depend on the number of bins and the binning strategy. A non-significant Hosmer–Lemeshow result does not show that calibration is good, particularly in small samples. Report the calibration curve with slope and intercept alongside any summary statistic.

Metrics

  • Brier scoreImplemented

    Mean squared difference between predicted probabilities and binary outcomes.

    es.brier_score

  • Weighted average gap between predicted confidence and observed frequency across probability bins.

    es.expected_calibration_error

  • Largest gap between confidence and observed frequency over all bins.

    es.maximum_calibration_error

  • Coefficients of a logistic regression of the outcome on the logit of the predicted probability.

    es.calibration_slope

  • Goodness-of-fit test comparing observed and expected event counts across risk groups.

    es.hosmer_lemeshow