Statistical tests
Paired bootstrap test
Implemented
statistics.paired_bootstrap_testDefinition
Test whether two models differ on any metric by resampling the same examples for both and counting how often the difference changes sign.
Formula
p = 2 · min(P*(Δ ≤ 0), P*(Δ ≥ 0)) over paired resamples
Range: p in [0, 1]
Inputs and outputs
- metric: any EvalSuite metric
- y_true
- y_pred_a, y_pred_b: the two models' predictions
- n_resamples, random_state
Returns: result object
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.paired_bootstrap_test(es.accuracy, y_true, pred_a, pred_b, n_resamples=500, random_state=0)References
- Koehn P. Statistical significance tests for machine translation evaluation. EMNLP. 2004:388-395.