Statistical analysis
Available since v0.2.0 (paired tests since v0.1.0)Statistical tests in EvalSuite return a structured result containing the statistic, p-value, degrees of freedom where applicable, an effect size, and the method used. They never return a bare p-value.
Usage
import evalsuite as es
res = es.mcnemar_test(y_true, pred_a, pred_b)
res.statistic, res.p_value
es.delong_test(y_true, prob_a, prob_b)
es.paired_bootstrap_test("f1", y_true, pred_a, pred_b, random_state=0)
es.cohens_d(scores_a, scores_b, paired=True)
es.cliffs_delta(scores_a, scores_b)
es.adjust_pvalues(p_values, method="holm") # bonferroni, holm, hochberg, bh, by
es.t_test(scores_a, scores_b) # Welch; mean difference with CI and Cohen's d
es.paired_t_test(fold_a, fold_b) # with Cohen's d_z
es.wilcoxon_test(fold_a, fold_b) # with matched-pairs rank-biserial r
es.mann_whitney_test(a, b) # with rank-biserial r (= Cliff's delta)
es.friedman_test(model_a, model_b, model_c) # several models on the same datasets; Kendall's W
es.kruskal_wallis_test(g1, g2, g3) # epsilon-squared
es.shapiro_wilk_test(residuals)
es.chi_square_test(table) # Cramér's V
es.fisher_exact_test([[8, 2], [1, 5]]) # odds ratio
es.cramers_v(table, bias_correction=True)Available tests
| Family | Tests | Effect size reported |
|---|---|---|
| Parametric | Welch and Student t-test, paired t-test | mean difference with CI, Cohen's d / d_z |
| Non-parametric | Mann–Whitney U, Wilcoxon signed-rank, Kruskal–Wallis, Friedman | rank-biserial r, epsilon-squared, Kendall's W |
| Categorical | χ² test of independence, Fisher's exact, McNemar | Cramér's V, odds ratio |
| Model comparison | DeLong, paired bootstrap | AUC difference with CI |
| Assumptions | Shapiro–Wilk |
Statistics and p-values come from SciPy and are checked against SciPy and statsmodels in the test suite. ANOVA and further assumption tests (Levene, Bartlett, Kolmogorov–Smirnov) are not included yet.
Tests in the registry
| Metric | Description | Status | API |
|---|---|---|---|
| Independent-samples t-test | Tests whether two independent groups have equal means. Welch's version does not assume equal variances. | Implemented | es.t_test |
| Paired t-test | Tests whether the mean of paired differences is zero, for example per-fold scores of two models. | Implemented | es.paired_t_test |
| Mann–Whitney U test | Rank-based test for whether one of two independent samples tends to have larger values. | Implemented | es.mann_whitney_test |
| Wilcoxon signed-rank test | Rank-based test for paired samples. | Implemented | es.wilcoxon_test |
| Kruskal–Wallis test | Rank-based test comparing three or more independent groups. | Implemented | es.kruskal_wallis_test |
| Friedman test | Rank-based test for three or more related samples, such as several models evaluated on the same datasets. | Implemented | es.friedman_test |
| McNemar's test | Compares two classifiers on the same observations using their discordant predictions. | Implemented | es.mcnemar_test |
| Shapiro–Wilk test | Tests the null hypothesis that a sample comes from a normal distribution. | Implemented | es.shapiro_wilk_test |
| DeLong test for correlated ROC AUCs | Compares the ROC AUCs of two models evaluated on the same observations. | Implemented | es.delong_test |
- Independent-samples t-testImplemented
Tests whether two independent groups have equal means. Welch's version does not assume equal variances.
es.t_test
- Paired t-testImplemented
Tests whether the mean of paired differences is zero, for example per-fold scores of two models.
es.paired_t_test
- Mann–Whitney U testImplemented
Rank-based test for whether one of two independent samples tends to have larger values.
es.mann_whitney_test
- Wilcoxon signed-rank testImplemented
Rank-based test for paired samples.
es.wilcoxon_test
- Kruskal–Wallis testImplemented
Rank-based test comparing three or more independent groups.
es.kruskal_wallis_test
- Friedman testImplemented
Rank-based test for three or more related samples, such as several models evaluated on the same datasets.
es.friedman_test
- McNemar's testImplemented
Compares two classifiers on the same observations using their discordant predictions.
es.mcnemar_test
- Shapiro–Wilk testImplemented
Tests the null hypothesis that a sample comes from a normal distribution.
es.shapiro_wilk_test
- DeLong test for correlated ROC AUCsImplemented
Compares the ROC AUCs of two models evaluated on the same observations.
es.delong_test
Effect sizes
| Metric | Description | Status | API |
|---|---|---|---|
| Cohen's d | Standardised mean difference using the pooled standard deviation. | Implemented | es.cohens_d |
| Hedges' g | Cohen's d with a small-sample bias correction. | Implemented | es.hedges_g |
| Cramér's V | Strength of association between two categorical variables. | Implemented | es.cramers_v |
- Cohen's dImplemented
Standardised mean difference using the pooled standard deviation.
es.cohens_d
- Hedges' gImplemented
Cohen's d with a small-sample bias correction.
es.hedges_g
- Cramér's VImplemented
Strength of association between two categorical variables.
es.cramers_v
Multiple-testing corrections
| Metric | Description | Status | API |
|---|---|---|---|
| Bonferroni correction | Controls the family-wise error rate by multiplying each p-value by the number of tests. | Implemented | es.adjust_pvalues |
| Holm step-down | Sequentially rejective procedure that controls the family-wise error rate and is uniformly more powerful than Bonferroni. | Implemented | es.adjust_pvalues |
| Hochberg step-up | Step-up procedure controlling the family-wise error rate under independence or certain positive dependence. | Implemented | es.adjust_pvalues |
| Benjamini–Hochberg (FDR) | Controls the expected proportion of false discoveries among rejected hypotheses. | Implemented | es.adjust_pvalues |
- Bonferroni correctionImplemented
Controls the family-wise error rate by multiplying each p-value by the number of tests.
es.adjust_pvalues
- Holm step-downImplemented
Sequentially rejective procedure that controls the family-wise error rate and is uniformly more powerful than Bonferroni.
es.adjust_pvalues
- Hochberg step-upImplemented
Step-up procedure controlling the family-wise error rate under independence or certain positive dependence.
es.adjust_pvalues
- Benjamini–Hochberg (FDR)Implemented
Controls the expected proportion of false discoveries among rejected hypotheses.
es.adjust_pvalues