Skip to content
EvalSuite
Documentation menu

Bootstrap and confidence intervals

Available in v0.1.0

A metric estimated on a finite test set is uncertain. EvalSuite attaches an interval to estimates using either an analytical method suited to the parameter or a reproducible bootstrap.

Analytical intervals

For proportions such as accuracy, sensitivity, or specificity, the Wilson score interval is the default because it keeps good coverage for small samples and proportions near 0 or 1. Clopper–Pearson is available when guaranteed (conservative) coverage is preferred.

Pythonv0.1.0
import evalsuite as es

es.accuracy_ci(y_true, y_pred, method="wilson", level=0.95)
es.proportion_ci(8, 10, method="clopper-pearson")
es.roc_auc_ci(y_true, y_prob)          # DeLong

Bootstrap engine

Pythonv0.1.0
ci = es.bootstrap_ci(
  "f1", y_true, y_pred,
  n_resamples=2000,
  level=0.95,
  method="bca",          # or "percentile", "basic"
  random_state=42,
)
ci.estimate, ci.low, ci.high

The engine is deterministic for a fixed random_state, and never touches global random state. Resampling is stratified by class for classification metrics.

Resampling needs at least two observations: bootstrap_ci, paired_bootstrap_test and compare raise StatisticalTestError for a single one, because an interval from one value would have zero width and no meaning (since v0.3.1). Every resampling function is reproducible when random_state is set.

Methods

  • Confidence interval for a binomial proportion with good coverage for small samples and extreme proportions.

    es.proportion_ci

  • Exact binomial interval obtained by inverting two one-sided binomial tests.

    es.proportion_ci

  • Interval from the empirical quantiles of a statistic recomputed on resampled data.

    es.bootstrap_ci

  • BCa bootstrapImplemented

    Bias-corrected and accelerated bootstrap interval.

    es.bootstrap_ci