Skip to content
EvalSuite
Documentation menu

Model comparison

Available in v0.1.0

es.compare evaluates several models on the same test set, attaches a confidence interval to every estimate, runs a paired test for every pair of models (or every model against a baseline you name), corrects the p-values for multiple testing and reports the best model per metric.

Pythonv0.1.0
import evalsuite as es

comparison = es.compare(
  y_true,
  {"baseline": pred_base, "candidate": pred_new},
  probabilities={"baseline": prob_base, "candidate": prob_new},
  metrics=["accuracy", "f1", "roc_auc"],
  random_state=0,
)
print(comparison)
comparison.to_latex()          # best value per metric in bold
comparison.save("comparison.html")

Which test is used

MetricPaired test
Accuracy (label predictions)McNemar
ROC AUCDeLong
Any other metricpaired bootstrap

The tests are also available on their own: es.mcnemar_test, es.delong_test and es.paired_bootstrap_test. Holm correction is the default (correction="holm"; also bonferroni, bh, by).

Segmentation and detection

Pythonv0.3.0
es.compare(y_true_masks, {"unet": masks_a, "deeplab": masks_b})   # Dice and IoU
es.compare(y_true_boxes, {"yolo": preds_a, "detr": preds_b})      # mAP@[.50:.95]

For masks and detections the resampling unit is the image, not the pixel or box, so intervals and paired tests reflect variation between images.

A significant difference on one test set does not establish that a model is superior in general. Report effect sizes, intervals and the evaluation design alongside any test.