Calibration
Curve and ECE since v0.1.0; slope, intercept, MCE and Hosmer–Lemeshow since v0.2.0A model is calibrated when, among cases given a predicted risk of 0.2, about 20% have the outcome. Discrimination metrics such as ROC AUC do not measure this.
Usage
import evalsuite as es
es.expected_calibration_error(y_true, y_prob, n_bins=10, strategy="quantile")
prob_true, prob_pred, bin_weight = es.calibration_curve(y_true, y_prob, n_bins=10)
es.brier_score(y_true, y_prob)
es.plot.calibration(y_true, y_prob) # reliability diagram with ECE and Brier scorereport = es.calibration_report(y_true, y_prob) # Brier, ECE, MCE, intercept, slope, Hosmer–Lemeshow
print(report)
es.calibration_slope(y_true, y_prob) # ideal 1; below 1 means predictions are too extreme
es.calibration_intercept(y_true, y_prob) # ideal 0 (calibration-in-the-large)
es.maximum_calibration_error(y_true, y_prob)
hl = es.hosmer_lemeshow(y_true, y_prob, n_groups=10)
hl.statistic, hl.p_value, hl.params["df"]The slope and intercept come from the logistic recalibration model logit P(y = 1) = a + b · logit(p̂) (Cox, 1958), fitted by maximum likelihood; the intercept is estimated with the slope fixed at 1. Both match statsmodels' logistic GLM in the test suite.
Interpreting calibration
No single number proves that a model is calibrated. Binned metrics depend on the number of bins and the binning strategy. A non-significant Hosmer–Lemeshow result does not show that calibration is good, particularly in small samples. Report the calibration curve with slope and intercept alongside any summary statistic.
Metrics
| Metric | Description | Status | API |
|---|---|---|---|
| Brier score | Mean squared difference between predicted probabilities and binary outcomes. | Implemented | es.brier_score |
| Expected calibration error | Weighted average gap between predicted confidence and observed frequency across probability bins. | Implemented | es.expected_calibration_error |
| Maximum calibration error | Largest gap between confidence and observed frequency over all bins. | Implemented | es.maximum_calibration_error |
| Calibration slope and intercept | Coefficients of a logistic regression of the outcome on the logit of the predicted probability. | Implemented | es.calibration_slope |
| Hosmer–Lemeshow test | Goodness-of-fit test comparing observed and expected event counts across risk groups. | Implemented | es.hosmer_lemeshow |
- Brier scoreImplemented
Mean squared difference between predicted probabilities and binary outcomes.
es.brier_score
- Expected calibration errorImplemented
Weighted average gap between predicted confidence and observed frequency across probability bins.
es.expected_calibration_error
- Maximum calibration errorImplemented
Largest gap between confidence and observed frequency over all bins.
es.maximum_calibration_error
- Calibration slope and interceptImplemented
Coefficients of a logistic regression of the outcome on the logit of the predicted probability.
es.calibration_slope
- Hosmer–Lemeshow testImplemented
Goodness-of-fit test comparing observed and expected event counts across risk groups.
es.hosmer_lemeshow