Benchmarks
| Suite | Cases | Measurements | Match reference | EvalSuite faster | Geometric-mean speed-up | Range |
|---|---|---|---|---|---|---|
| Classification and regression | 4 | 12 | 12/12 | 10/12 | 5.88× | 0.47×–49.07× |
| Clinical, calibration and statistics | 8 | 24 | 24/24 | 17/24 | 1.71× | 0.25×–11.13× |
| LLM evaluation | 6 | 18 | 18/18 | 10/18 | 1.48× | 0.83×–7.72× |
| Segmentation and object detection | 3 | 9 | 9/9 | 5/9 | 1.97× | 0.64×–10.39× |
| All suites | 21 | 63 | 63/63 | 42/63 | 2.12× | 0.25×–49.07× |
| Case | Release | Reference | Speed-up, n = 1,000 | Speed-up, n = 100,000 | Speed-up, n = 1,000,000 | Max |difference| |
|---|---|---|---|---|---|---|
| 10 classes: macro F1 | v0.1 | scikit-learn | 13.25× | 5.08× | 5.69× | 0 |
| Agreement: Krippendorff's alpha, interval (4 raters × 10 items) | v0.4 | krippendorff | 0.92× | – | – | 0 |
| Agreement: Krippendorff's alpha, interval (4 raters × 1000 items) | v0.4 | krippendorff | – | 0.94× | – | 3.3e-15 |
| Agreement: Krippendorff's alpha, interval (4 raters × 10000 items) | v0.4 | krippendorff | – | – | 0.96× | 1.8e-14 |
| Binary: 8 label metrics via evaluate() | v0.1 | scikit-learn | 35.87× | 49.07× | 35.98× | 0 |
| Binary: ROC AUC | v0.1 | scikit-learn | 8.73× | 1.57× | 1.75× | 1.1e-16 |
| Calibration: slope and intercept | v0.2 | statsmodels | 4.61× | 7.32× | 5.71× | 3.9e-16 |
| Clinical: diagnostic report (7 CIs) | v0.2 | statsmodels | 2.23× | 0.82× | 0.25× | 1.1e-14 |
| Clinical: sensitivity, specificity, LR+, LR− | v0.2 | scikit-learn | 11.13× | 6.54× | 5.47× | 4.4e-16 |
| Decision curve: 99 thresholds | v0.2 | NumPy loop | 5.36× | 1.40× | 1.22× | 5.6e-17 |
| Detection: COCO evaluation | v0.3 | pycocotools | 1.65× | 0.92× | 1.01× | 0 |
| Multiple testing: Hochberg (n p-values) | v0.2 | statsmodels | 1.71× | 0.96× | 1.09× | 0 |
| Regression: MAE, MSE, RMSE, R² via evaluate() | v0.1 | scikit-learn | 6.85× | 0.92× | 0.47× | 0 |
| Retrieval: MRR, MAP@20, NDCG@10 (10 queries) | v0.4 | ranx | 7.72× | – | – | 0 |
| Retrieval: MRR, MAP@20, NDCG@10 (1000 queries) | v0.4 | ranx | – | 2.58× | – | 0 |
| Retrieval: MRR, MAP@20, NDCG@10 (10000 queries) | v0.4 | ranx | – | – | 5.05× | 6.9e-18 |
| Segmentation: Dice and IoU per class (n = pixels) | v0.3 | scikit-learn | 9.19× | 10.39× | 7.47× | 0 |
| Segmentation: Hausdorff distance | v0.3 | SciPy | 0.77× | 0.64× | 0.82× | 0 |
| Statistics: Cramér's V (5×5 table) | v0.2 | SciPy | 1.58× | 1.07× | 1.00× | 0 |
| Statistics: Mann–Whitney U | v0.2 | SciPy | 0.88× | 1.14× | 0.96× | 0 |
| Statistics: Welch t-test | v0.2 | SciPy | 1.20× | 0.66× | 0.52× | 0 |
| Structured: JSON Schema compliance (10 documents) | v0.4 | jsonschema | 2.67× | – | – | 0 |
| Structured: JSON Schema compliance (1000 documents) | v0.4 | jsonschema | – | 2.36× | – | 0 |
| Structured: JSON Schema compliance (10000 documents) | v0.4 | jsonschema | – | – | 2.29× | 0 |
| Text: corpus BLEU and chrF (10 sentences) | v0.4 | sacreBLEU | 0.89× | – | – | 1.4e-14 |
| Text: corpus BLEU and chrF (1000 sentences) | v0.4 | sacreBLEU | – | 0.98× | – | 1.4e-14 |
| Text: corpus BLEU and chrF (10000 sentences) | v0.4 | sacreBLEU | – | – | 1.09× | 0 |
| Text: METEOR, exact and stem matches (10 sentences) | v0.4 | NLTK | 1.09× | – | – | 0 |
| Text: METEOR, exact and stem matches (1000 sentences) | v0.4 | NLTK | – | 1.13× | – | 0 |
| Text: METEOR, exact and stem matches (10000 sentences) | v0.4 | NLTK | – | – | 1.28× | 0 |
| Text: ROUGE-1, ROUGE-2, ROUGE-L (10 sentences) | v0.4 | rouge-score | 0.83× | – | – | 0 |
| Text: ROUGE-1, ROUGE-2, ROUGE-L (1000 sentences) | v0.4 | rouge-score | – | 0.93× | – | 0 |
| Text: ROUGE-1, ROUGE-2, ROUGE-L (10000 sentences) | v0.4 | rouge-score | – | – | 0.85× | 0 |
| Case | n | Reference | EvalSuite (ms) | Reference (ms) | Speed-up | EvalSuite peak (MiB) | Reference peak (MiB) | Max |difference| |
|---|---|---|---|---|---|---|---|---|
| 10 classes: macro F1 | 1,000 | scikit-learn | 0.101 | 1.342 | 13.25× | 0.05 | 0.03 | 0 |
| 10 classes: macro F1 | 100,000 | scikit-learn | 2.995 | 15.218 | 5.08× | 3.05 | 2.18 | 0 |
| 10 classes: macro F1 | 1,000,000 | scikit-learn | 23.809 | 135.515 | 5.69× | 30.52 | 21.79 | 0 |
| Binary: 8 label metrics via evaluate() | 1,000 | scikit-learn | 0.256 | 9.178 | 35.87× | 0.05 | 0.05 | 0 |
| Binary: 8 label metrics via evaluate() | 100,000 | scikit-learn | 2.350 | 115.325 | 49.07× | 3.21 | 3.07 | 0 |
| Binary: 8 label metrics via evaluate() | 1,000,000 | scikit-learn | 28.459 | 1023.820 | 35.98× | 31.54 | 30.53 | 0 |
| Binary: ROC AUC | 1,000 | scikit-learn | 0.167 | 1.460 | 8.73× | 0.09 | 0.08 | 1.1e-16 |
| Binary: ROC AUC | 100,000 | scikit-learn | 17.642 | 27.624 | 1.57× | 9.16 | 7.64 | 0 |
| Binary: ROC AUC | 1,000,000 | scikit-learn | 187.361 | 327.279 | 1.75× | 91.56 | 76.30 | 0 |
| Regression: MAE, MSE, RMSE, R² via evaluate() | 1,000 | scikit-learn | 0.095 | 0.650 | 6.85× | 0.03 | 0.02 | 0 |
| Regression: MAE, MSE, RMSE, R² via evaluate() | 100,000 | scikit-learn | 1.912 | 1.763 | 0.92× | 2.29 | 1.53 | 0 |
| Regression: MAE, MSE, RMSE, R² via evaluate() | 1,000,000 | scikit-learn | 22.854 | 10.633 | 0.47× | 22.89 | 15.26 | 0 |
| Case | n | Reference | EvalSuite (ms) | Reference (ms) | Speed-up | EvalSuite peak (MiB) | Reference peak (MiB) | Max |difference| |
|---|---|---|---|---|---|---|---|---|
| Calibration: slope and intercept | 1,000 | statsmodels | 0.559 | 2.578 | 4.61× | 0.10 | 0.60 | 2.8e-16 |
| Calibration: slope and intercept | 100,000 | statsmodels | 15.730 | 115.193 | 7.32× | 8.46 | 58.00 | 5.6e-17 |
| Calibration: slope and intercept | 1,000,000 | statsmodels | 205.659 | 1173.684 | 5.71× | 83.99 | 579.85 | 3.9e-16 |
| Clinical: diagnostic report (7 CIs) | 1,000 | statsmodels | 0.235 | 0.523 | 2.23× | 0.05 | 0.01 | 1.1e-14 |
| Clinical: diagnostic report (7 CIs) | 100,000 | statsmodels | 1.951 | 1.602 | 0.82× | 3.05 | 0.29 | 7.1e-15 |
| Clinical: diagnostic report (7 CIs) | 1,000,000 | statsmodels | 19.570 | 4.878 | 0.25× | 30.52 | 1.91 | 7.1e-15 |
| Clinical: sensitivity, specificity, LR+, LR− | 1,000 | scikit-learn | 0.309 | 3.440 | 11.13× | 0.05 | 0.03 | 4.4e-16 |
| Clinical: sensitivity, specificity, LR+, LR− | 100,000 | scikit-learn | 6.632 | 43.403 | 6.54× | 3.05 | 2.24 | 5.6e-17 |
| Clinical: sensitivity, specificity, LR+, LR− | 1,000,000 | scikit-learn | 75.166 | 411.527 | 5.47× | 30.52 | 22.32 | 4.4e-16 |
| Decision curve: 99 thresholds | 1,000 | NumPy loop | 0.268 | 1.437 | 5.36× | 0.07 | 0.01 | 5.6e-17 |
| Decision curve: 99 thresholds | 100,000 | NumPy loop | 11.350 | 15.891 | 1.40× | 6.87 | 0.29 | 5.6e-17 |
| Decision curve: 99 thresholds | 1,000,000 | NumPy loop | 165.010 | 201.852 | 1.22× | 68.67 | 1.97 | 5.6e-17 |
| Multiple testing: Hochberg (n p-values) | 1,000 | statsmodels | 0.045 | 0.077 | 1.71× | 0.05 | 0.05 | 0 |
| Multiple testing: Hochberg (n p-values) | 100,000 | statsmodels | 4.867 | 4.675 | 0.96× | 4.58 | 3.97 | 0 |
| Multiple testing: Hochberg (n p-values) | 1,000,000 | statsmodels | 72.148 | 78.964 | 1.09× | 45.78 | 39.17 | 0 |
| Statistics: Cramér's V (5×5 table) | 1,000 | SciPy | 0.259 | 0.409 | 1.58× | 0.00 | 0.00 | 0 |
| Statistics: Cramér's V (5×5 table) | 100,000 | SciPy | 0.259 | 0.275 | 1.07× | 0.00 | 0.00 | 0 |
| Statistics: Cramér's V (5×5 table) | 1,000,000 | SciPy | 0.231 | 0.232 | 1.00× | 0.00 | 0.00 | 0 |
| Statistics: Mann–Whitney U | 1,000 | SciPy | 0.640 | 0.564 | 0.88× | 0.16 | 0.14 | 0 |
| Statistics: Mann–Whitney U | 100,000 | SciPy | 33.722 | 38.490 | 1.14× | 15.45 | 13.93 | 0 |
| Statistics: Mann–Whitney U | 1,000,000 | SciPy | 353.520 | 340.573 | 0.96× | 154.50 | 139.24 | 0 |
| Statistics: Welch t-test | 1,000 | SciPy | 0.821 | 0.987 | 1.20× | 0.04 | 0.02 | 0 |
| Statistics: Welch t-test | 100,000 | SciPy | 2.097 | 1.382 | 0.66× | 3.06 | 1.53 | 0 |
| Statistics: Welch t-test | 1,000,000 | SciPy | 16.124 | 8.349 | 0.52× | 30.52 | 15.26 | 0 |
| Case | n | Reference | EvalSuite (ms) | Reference (ms) | Speed-up | EvalSuite peak (MiB) | Reference peak (MiB) | Max |difference| |
|---|---|---|---|---|---|---|---|---|
| Detection: COCO evaluation (10 images) | 1,000 | pycocotools | 13.097 | 21.589 | 1.65× | 0.54 | 1.33 | 0 |
| Detection: COCO evaluation (100 images) | 100,000 | pycocotools | 102.270 | 93.783 | 0.92× | 1.15 | 4.29 | 0 |
| Detection: COCO evaluation (1000 images) | 1,000,000 | pycocotools | 838.714 | 848.551 | 1.01× | 6.92 | 34.02 | 0 |
| Segmentation: Dice and IoU per class (n = pixels) | 1,000 | scikit-learn | 0.314 | 2.882 | 9.19× | 0.16 | 0.10 | 0 |
| Segmentation: Dice and IoU per class (n = pixels) | 100,000 | scikit-learn | 3.864 | 40.162 | 10.39× | 0.17 | 2.32 | 0 |
| Segmentation: Dice and IoU per class (n = pixels) | 1,000,000 | scikit-learn | 34.283 | 256.019 | 7.47× | 0.25 | 23.48 | 0 |
| Segmentation: Hausdorff distance (1 image) | 1,000 | SciPy | 0.515 | 0.397 | 0.77× | 0.15 | 0.03 | 0 |
| Segmentation: Hausdorff distance (24 images) | 100,000 | SciPy | 21.556 | 13.741 | 0.64× | 0.15 | 0.03 | 0 |
| Segmentation: Hausdorff distance (50 images) | 1,000,000 | SciPy | 24.165 | 19.832 | 0.82× | 0.15 | 0.03 | 0 |
| Case | n | Reference | EvalSuite (ms) | Reference (ms) | Speed-up | EvalSuite peak (MiB) | Reference peak (MiB) | Max |difference| |
|---|---|---|---|---|---|---|---|---|
| Agreement: Krippendorff's alpha, interval (4 raters × 10 items) | 1,000 | krippendorff | 0.049 | 0.045 | 0.92× | 0.01 | 0.01 | 0 |
| Agreement: Krippendorff's alpha, interval (4 raters × 1000 items) | 100,000 | krippendorff | 0.412 | 0.387 | 0.94× | 0.26 | 0.69 | 3.3e-15 |
| Agreement: Krippendorff's alpha, interval (4 raters × 10000 items) | 1,000,000 | krippendorff | 3.511 | 3.372 | 0.96× | 2.26 | 6.32 | 1.8e-14 |
| Retrieval: MRR, MAP@20, NDCG@10 (10 queries) | 1,000 | ranx | 0.182 | 1.407 | 7.72× | 0.01 | 0.04 | 0 |
| Retrieval: MRR, MAP@20, NDCG@10 (1000 queries) | 100,000 | ranx | 29.821 | 77.001 | 2.58× | 0.49 | 3.12 | 0 |
| Retrieval: MRR, MAP@20, NDCG@10 (10000 queries) | 1,000,000 | ranx | 170.388 | 861.009 | 5.05× | 5.36 | 31.01 | 6.9e-18 |
| Structured: JSON Schema compliance (10 documents) | 1,000 | jsonschema | 0.157 | 0.421 | 2.67× | 0.00 | 0.00 | 0 |
| Structured: JSON Schema compliance (1000 documents) | 100,000 | jsonschema | 14.393 | 33.995 | 2.36× | 0.05 | 0.02 | 0 |
| Structured: JSON Schema compliance (10000 documents) | 1,000,000 | jsonschema | 144.622 | 330.462 | 2.29× | 0.50 | 0.16 | 0 |
| Text: corpus BLEU and chrF (10 sentences) | 1,000 | sacreBLEU | 2.206 | 1.960 | 0.89× | 0.10 | 0.22 | 1.4e-14 |
| Text: corpus BLEU and chrF (1000 sentences) | 100,000 | sacreBLEU | 255.898 | 250.494 | 0.98× | 0.74 | 21.91 | 1.4e-14 |
| Text: corpus BLEU and chrF (10000 sentences) | 1,000,000 | sacreBLEU | 2789.790 | 3051.755 | 1.09× | 7.26 | 210.17 | 0 |
| Text: METEOR, exact and stem matches (10 sentences) | 1,000 | NLTK | 0.459 | 0.501 | 1.09× | 0.01 | 0.01 | 0 |
| Text: METEOR, exact and stem matches (1000 sentences) | 100,000 | NLTK | 52.275 | 58.965 | 1.13× | 0.12 | 0.04 | 0 |
| Text: METEOR, exact and stem matches (10000 sentences) | 1,000,000 | NLTK | 488.126 | 626.693 | 1.28× | 1.15 | 0.39 | 0 |
| Text: ROUGE-1, ROUGE-2, ROUGE-L (10 sentences) | 1,000 | rouge-score | 1.145 | 0.947 | 0.83× | 0.01 | 0.01 | 0 |
| Text: ROUGE-1, ROUGE-2, ROUGE-L (1000 sentences) | 100,000 | rouge-score | 120.724 | 111.884 | 0.93× | 0.12 | 0.64 | 0 |
| Text: ROUGE-1, ROUGE-2, ROUGE-L (10000 sentences) | 1,000,000 | rouge-score | 1290.112 | 1095.609 | 0.85× | 1.16 | 6.55 | 0 |
Reading the results
Many metrics at once is where EvalSuite is fastest. evaluate() validates the inputs once and builds the confusion matrix once, then derives all eight label metrics from it: 36–49× faster than eight separate scikit-learn calls. Sensitivity, specificity and both likelihood ratios together are 5–11× faster.
Calibration slope and intercept are 5–6× faster than statsmodels' GLM and use far less memory (84 MiB against 580 MiB at a million samples), because EvalSuite fits the two small logistic models with a dedicated Newton–Raphson solver.
Object detection matches pycocotools exactly on all twelve COCO numbers and is 0.9–1.6× as fast, using a fifth of its memory. Per-class Dice and IoU are 7–10× faster than building scikit-learn's confusion matrix. Hausdorff distance (0.6–0.8×) extracts surfaces and computes HD95 and ASSD alongside the maximum that SciPy returns.
LLM metrics give the reference libraries' numbers and run at about their speed: corpus BLEU and chrF 0.89–1.09× sacreBLEU, ROUGE 0.8–0.9× rouge-score, METEOR 1.09–1.28× NLTK, ranking metrics 2.6–7.7× ranx, JSON Schema compliance 2.3–2.7× jsonschema and Krippendorff's alpha 0.9–1.0× the krippendorff package.
Hypothesis tests call SciPy for the statistic and p-value, so they match its speed at best. The Welch t-test is about half SciPy's speed at large n because EvalSuite also validates the inputs and computes the confidence interval and Cohen's d.
Some rows are slower, and we show them. The diagnostic report (0.2–2.2× across sizes) validates labels and reports ten measures, where the reference computes seven intervals from counts it takes directly. Regression on a million values (0.5×) spends most of its ~20 ms checking every value for NaN, infinity, shape and dtype.
Reproduce on your machine
pip install evalsuite-python scikit-learn statsmodels pycocotools sacrebleu rouge-score nltk ranx krippendorff jsonschema
evalsuite benchmark # every case at 1,000 / 100,000 / 1,000,000 samples
evalsuite benchmark --suite clinical # only the v0.2 clinical, calibration and statistics cases
evalsuite benchmark --suite vision # only the v0.3 segmentation and detection cases
evalsuite benchmark --suite llm # only the v0.4 LLM cases
evalsuite benchmark --quick # small sizes onlyThe full table and notes are in BENCHMARKS.md in the package repository. Results on your hardware will differ in absolute time; the ratios are what to compare.