Skip to content
EvalSuite
Documentation menu

Statistical analysis

Available since v0.2.0 (paired tests since v0.1.0)

Statistical tests in EvalSuite return a structured result containing the statistic, p-value, degrees of freedom where applicable, an effect size, and the method used. They never return a bare p-value.

Usage

Pythonv0.1.0
import evalsuite as es

res = es.mcnemar_test(y_true, pred_a, pred_b)
res.statistic, res.p_value

es.delong_test(y_true, prob_a, prob_b)
es.paired_bootstrap_test("f1", y_true, pred_a, pred_b, random_state=0)

es.cohens_d(scores_a, scores_b, paired=True)
es.cliffs_delta(scores_a, scores_b)
es.adjust_pvalues(p_values, method="holm")   # bonferroni, holm, hochberg, bh, by

es.t_test(scores_a, scores_b)             # Welch; mean difference with CI and Cohen's d
es.paired_t_test(fold_a, fold_b)          # with Cohen's d_z
es.wilcoxon_test(fold_a, fold_b)          # with matched-pairs rank-biserial r
es.mann_whitney_test(a, b)                # with rank-biserial r (= Cliff's delta)
es.friedman_test(model_a, model_b, model_c)   # several models on the same datasets; Kendall's W
es.kruskal_wallis_test(g1, g2, g3)        # epsilon-squared
es.shapiro_wilk_test(residuals)
es.chi_square_test(table)                 # Cramér's V
es.fisher_exact_test([[8, 2], [1, 5]])    # odds ratio
es.cramers_v(table, bias_correction=True)

Available tests

FamilyTestsEffect size reported
ParametricWelch and Student t-test, paired t-testmean difference with CI, Cohen's d / d_z
Non-parametricMann–Whitney U, Wilcoxon signed-rank, Kruskal–Wallis, Friedmanrank-biserial r, epsilon-squared, Kendall's W
Categoricalχ² test of independence, Fisher's exact, McNemarCramér's V, odds ratio
Model comparisonDeLong, paired bootstrapAUC difference with CI
AssumptionsShapiro–Wilk

Statistics and p-values come from SciPy and are checked against SciPy and statsmodels in the test suite. ANOVA and further assumption tests (Levene, Bartlett, Kolmogorov–Smirnov) are not included yet.

Tests in the registry

  • Tests whether two independent groups have equal means. Welch's version does not assume equal variances.

    es.t_test

  • Paired t-testImplemented

    Tests whether the mean of paired differences is zero, for example per-fold scores of two models.

    es.paired_t_test

  • Rank-based test for whether one of two independent samples tends to have larger values.

    es.mann_whitney_test

  • Rank-based test for paired samples.

    es.wilcoxon_test

  • Rank-based test comparing three or more independent groups.

    es.kruskal_wallis_test

  • Friedman testImplemented

    Rank-based test for three or more related samples, such as several models evaluated on the same datasets.

    es.friedman_test

  • McNemar's testImplemented

    Compares two classifiers on the same observations using their discordant predictions.

    es.mcnemar_test

  • Tests the null hypothesis that a sample comes from a normal distribution.

    es.shapiro_wilk_test

  • Compares the ROC AUCs of two models evaluated on the same observations.

    es.delong_test

Effect sizes

  • Cohen's dImplemented

    Standardised mean difference using the pooled standard deviation.

    es.cohens_d

  • Hedges' gImplemented

    Cohen's d with a small-sample bias correction.

    es.hedges_g

  • Cramér's VImplemented

    Strength of association between two categorical variables.

    es.cramers_v

Multiple-testing corrections

  • Controls the family-wise error rate by multiplying each p-value by the number of tests.

    es.adjust_pvalues

  • Holm step-downImplemented

    Sequentially rejective procedure that controls the family-wise error rate and is uniformly more powerful than Bonferroni.

    es.adjust_pvalues

  • Hochberg step-upImplemented

    Step-up procedure controlling the family-wise error rate under independence or certain positive dependence.

    es.adjust_pvalues

  • Controls the expected proportion of false discoveries among rejected hypotheses.

    es.adjust_pvalues