Statistical tests
Model comparison
Implemented
statistics.compareDefinition
Compare several models on the same examples: every metric with a bootstrap interval, pairwise paired tests and multiple-testing correction, in one report.
Formula
per metric: estimate, bootstrap CI; per pair: paired test p-value, adjusted
Range: report
Inputs and outputs
- y_true
- predictions: mapping model name → predictions
- metrics: metric names
- n_resamples, random_state
Returns: result object
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.compare(y_true, {"a": pred_a, "b": pred_b}, metrics=["accuracy"], n_resamples=200, random_state=0)References
- Efron B, Tibshirani RJ. An Introduction to the Bootstrap. Chapman & Hall; 1993.