Model comparison
Available in v0.1.0es.compare evaluates several models on the same test set, attaches a confidence interval to every estimate, runs a paired test for every pair of models (or every model against a baseline you name), corrects the p-values for multiple testing and reports the best model per metric.
import evalsuite as es
comparison = es.compare(
y_true,
{"baseline": pred_base, "candidate": pred_new},
probabilities={"baseline": prob_base, "candidate": prob_new},
metrics=["accuracy", "f1", "roc_auc"],
random_state=0,
)
print(comparison)
comparison.to_latex() # best value per metric in bold
comparison.save("comparison.html")Which test is used
| Metric | Paired test |
|---|---|
| Accuracy (label predictions) | McNemar |
| ROC AUC | DeLong |
| Any other metric | paired bootstrap |
The tests are also available on their own: es.mcnemar_test, es.delong_test and es.paired_bootstrap_test. Holm correction is the default (correction="holm"; also bonferroni, bh, by).
Segmentation and detection
es.compare(y_true_masks, {"unet": masks_a, "deeplab": masks_b}) # Dice and IoU
es.compare(y_true_boxes, {"yolo": preds_a, "detr": preds_b}) # mAP@[.50:.95]For masks and detections the resampling unit is the image, not the pixel or box, so intervals and paired tests reflect variation between images.
A significant difference on one test set does not establish that a model is superior in general. Report effect sizes, intervals and the evaluation design alongside any test.