Skip to content
EvalSuite
Documentation menu

Statistical tests

Paired bootstrap test

Implementedstatistics.paired_bootstrap_test

Definition

Test whether two models differ on any metric by resampling the same examples for both and counting how often the difference changes sign.

Formula

p = 2 · min(P*(Δ ≤ 0), P*(Δ ≥ 0)) over paired resamples

Range: p in [0, 1]

Inputs and outputs

  • metric: any EvalSuite metric
  • y_true
  • y_pred_a, y_pred_b: the two models' predictions
  • n_resamples, random_state

Returns: result object

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.1.0
import evalsuite as es

es.paired_bootstrap_test(es.accuracy, y_true, pred_a, pred_b, n_resamples=500, random_state=0)

References

  1. Koehn P. Statistical significance tests for machine translation evaluation. EMNLP. 2004:388-395.

Implementation status