Skip to content
EvalSuite
v0.4.0 results

Benchmarks

Speed and peak memory of EvalSuite against reference implementations (scikit-learn, statsmodels, SciPy, pycocotools, sacreBLEU, rouge-score, NLTK, ranx, krippendorff, jsonschema), computing the same quantities on the same data. Every result agrees with the reference to floating-point rounding.
Overall benchmarkper suite and in total, across 1,000 / 100,000 / 1,000,000 samples
SuiteCasesMeasurementsMatch referenceEvalSuite fasterGeometric-mean speed-upRange
Classification and regression41212/1210/125.88×0.47×–49.07×
Clinical, calibration and statistics82424/2417/241.71×0.25×–11.13×
LLM evaluation61818/1810/181.48×0.83×–7.72×
Segmentation and object detection399/95/91.97×0.64×–10.39×
All suites216363/6342/632.12×0.25×–49.07×
Overall: every case, all releases33 cases, 45 measurements; all agree with the reference; EvalSuite is faster in 32 of 45
CaseReleaseReferenceSpeed-up, n = 1,000Speed-up, n = 100,000Speed-up, n = 1,000,000Max |difference|
10 classes: macro F1v0.1scikit-learn13.25×5.08×5.69×0
Agreement: Krippendorff's alpha, interval (4 raters × 10 items)v0.4krippendorff0.92×––0
Agreement: Krippendorff's alpha, interval (4 raters × 1000 items)v0.4krippendorff–0.94×–3.3e-15
Agreement: Krippendorff's alpha, interval (4 raters × 10000 items)v0.4krippendorff––0.96×1.8e-14
Binary: 8 label metrics via evaluate()v0.1scikit-learn35.87×49.07×35.98×0
Binary: ROC AUCv0.1scikit-learn8.73×1.57×1.75×1.1e-16
Calibration: slope and interceptv0.2statsmodels4.61×7.32×5.71×3.9e-16
Clinical: diagnostic report (7 CIs)v0.2statsmodels2.23×0.82×0.25×1.1e-14
Clinical: sensitivity, specificity, LR+, LR−v0.2scikit-learn11.13×6.54×5.47×4.4e-16
Decision curve: 99 thresholdsv0.2NumPy loop5.36×1.40×1.22×5.6e-17
Detection: COCO evaluationv0.3pycocotools1.65×0.92×1.01×0
Multiple testing: Hochberg (n p-values)v0.2statsmodels1.71×0.96×1.09×0
Regression: MAE, MSE, RMSE, R² via evaluate()v0.1scikit-learn6.85×0.92×0.47×0
Retrieval: MRR, MAP@20, NDCG@10 (10 queries)v0.4ranx7.72×––0
Retrieval: MRR, MAP@20, NDCG@10 (1000 queries)v0.4ranx–2.58×–0
Retrieval: MRR, MAP@20, NDCG@10 (10000 queries)v0.4ranx––5.05×6.9e-18
Segmentation: Dice and IoU per class (n = pixels)v0.3scikit-learn9.19×10.39×7.47×0
Segmentation: Hausdorff distancev0.3SciPy0.77×0.64×0.82×0
Statistics: Cramér's V (5×5 table)v0.2SciPy1.58×1.07×1.00×0
Statistics: Mann–Whitney Uv0.2SciPy0.88×1.14×0.96×0
Statistics: Welch t-testv0.2SciPy1.20×0.66×0.52×0
Structured: JSON Schema compliance (10 documents)v0.4jsonschema2.67×––0
Structured: JSON Schema compliance (1000 documents)v0.4jsonschema–2.36×–0
Structured: JSON Schema compliance (10000 documents)v0.4jsonschema––2.29×0
Text: corpus BLEU and chrF (10 sentences)v0.4sacreBLEU0.89×––1.4e-14
Text: corpus BLEU and chrF (1000 sentences)v0.4sacreBLEU–0.98×–1.4e-14
Text: corpus BLEU and chrF (10000 sentences)v0.4sacreBLEU––1.09×0
Text: METEOR, exact and stem matches (10 sentences)v0.4NLTK1.09×––0
Text: METEOR, exact and stem matches (1000 sentences)v0.4NLTK–1.13×–0
Text: METEOR, exact and stem matches (10000 sentences)v0.4NLTK––1.28×0
Text: ROUGE-1, ROUGE-2, ROUGE-L (10 sentences)v0.4rouge-score0.83×––0
Text: ROUGE-1, ROUGE-2, ROUGE-L (1000 sentences)v0.4rouge-score–0.93×–0
Text: ROUGE-1, ROUGE-2, ROUGE-L (10000 sentences)v0.4rouge-score––0.85×0
Classification and regression (v0.1) against scikit-learn
CasenReferenceEvalSuite (ms)Reference (ms)Speed-upEvalSuite peak (MiB)Reference peak (MiB)Max |difference|
10 classes: macro F11,000scikit-learn0.1011.34213.25×0.050.030
10 classes: macro F1100,000scikit-learn2.99515.2185.08×3.052.180
10 classes: macro F11,000,000scikit-learn23.809135.5155.69×30.5221.790
Binary: 8 label metrics via evaluate()1,000scikit-learn0.2569.17835.87×0.050.050
Binary: 8 label metrics via evaluate()100,000scikit-learn2.350115.32549.07×3.213.070
Binary: 8 label metrics via evaluate()1,000,000scikit-learn28.4591023.82035.98×31.5430.530
Binary: ROC AUC1,000scikit-learn0.1671.4608.73×0.090.081.1e-16
Binary: ROC AUC100,000scikit-learn17.64227.6241.57×9.167.640
Binary: ROC AUC1,000,000scikit-learn187.361327.2791.75×91.5676.300
Regression: MAE, MSE, RMSE, R² via evaluate()1,000scikit-learn0.0950.6506.85×0.030.020
Regression: MAE, MSE, RMSE, R² via evaluate()100,000scikit-learn1.9121.7630.92×2.291.530
Regression: MAE, MSE, RMSE, R² via evaluate()1,000,000scikit-learn22.85410.6330.47×22.8915.260
Clinical, calibration and statistics (v0.2) against scikit-learn, statsmodels, SciPy
CasenReferenceEvalSuite (ms)Reference (ms)Speed-upEvalSuite peak (MiB)Reference peak (MiB)Max |difference|
Calibration: slope and intercept1,000statsmodels0.5592.5784.61×0.100.602.8e-16
Calibration: slope and intercept100,000statsmodels15.730115.1937.32×8.4658.005.6e-17
Calibration: slope and intercept1,000,000statsmodels205.6591173.6845.71×83.99579.853.9e-16
Clinical: diagnostic report (7 CIs)1,000statsmodels0.2350.5232.23×0.050.011.1e-14
Clinical: diagnostic report (7 CIs)100,000statsmodels1.9511.6020.82×3.050.297.1e-15
Clinical: diagnostic report (7 CIs)1,000,000statsmodels19.5704.8780.25×30.521.917.1e-15
Clinical: sensitivity, specificity, LR+, LR−1,000scikit-learn0.3093.44011.13×0.050.034.4e-16
Clinical: sensitivity, specificity, LR+, LR−100,000scikit-learn6.63243.4036.54×3.052.245.6e-17
Clinical: sensitivity, specificity, LR+, LR−1,000,000scikit-learn75.166411.5275.47×30.5222.324.4e-16
Decision curve: 99 thresholds1,000NumPy loop0.2681.4375.36×0.070.015.6e-17
Decision curve: 99 thresholds100,000NumPy loop11.35015.8911.40×6.870.295.6e-17
Decision curve: 99 thresholds1,000,000NumPy loop165.010201.8521.22×68.671.975.6e-17
Multiple testing: Hochberg (n p-values)1,000statsmodels0.0450.0771.71×0.050.050
Multiple testing: Hochberg (n p-values)100,000statsmodels4.8674.6750.96×4.583.970
Multiple testing: Hochberg (n p-values)1,000,000statsmodels72.14878.9641.09×45.7839.170
Statistics: Cramér's V (5×5 table)1,000SciPy0.2590.4091.58×0.000.000
Statistics: Cramér's V (5×5 table)100,000SciPy0.2590.2751.07×0.000.000
Statistics: Cramér's V (5×5 table)1,000,000SciPy0.2310.2321.00×0.000.000
Statistics: Mann–Whitney U1,000SciPy0.6400.5640.88×0.160.140
Statistics: Mann–Whitney U100,000SciPy33.72238.4901.14×15.4513.930
Statistics: Mann–Whitney U1,000,000SciPy353.520340.5730.96×154.50139.240
Statistics: Welch t-test1,000SciPy0.8210.9871.20×0.040.020
Statistics: Welch t-test100,000SciPy2.0971.3820.66×3.061.530
Statistics: Welch t-test1,000,000SciPy16.1248.3490.52×30.5215.260
Segmentation and object detection (v0.3) against scikit-learn, SciPy, pycocotools
CasenReferenceEvalSuite (ms)Reference (ms)Speed-upEvalSuite peak (MiB)Reference peak (MiB)Max |difference|
Detection: COCO evaluation (10 images)1,000pycocotools13.09721.5891.65×0.541.330
Detection: COCO evaluation (100 images)100,000pycocotools102.27093.7830.92×1.154.290
Detection: COCO evaluation (1000 images)1,000,000pycocotools838.714848.5511.01×6.9234.020
Segmentation: Dice and IoU per class (n = pixels)1,000scikit-learn0.3142.8829.19×0.160.100
Segmentation: Dice and IoU per class (n = pixels)100,000scikit-learn3.86440.16210.39×0.172.320
Segmentation: Dice and IoU per class (n = pixels)1,000,000scikit-learn34.283256.0197.47×0.2523.480
Segmentation: Hausdorff distance (1 image)1,000SciPy0.5150.3970.77×0.150.030
Segmentation: Hausdorff distance (24 images)100,000SciPy21.55613.7410.64×0.150.030
Segmentation: Hausdorff distance (50 images)1,000,000SciPy24.16519.8320.82×0.150.030
LLM evaluation (v0.4) against sacreBLEU, rouge-score, NLTK, ranx, krippendorff, jsonschema
CasenReferenceEvalSuite (ms)Reference (ms)Speed-upEvalSuite peak (MiB)Reference peak (MiB)Max |difference|
Agreement: Krippendorff's alpha, interval (4 raters × 10 items)1,000krippendorff0.0490.0450.92×0.010.010
Agreement: Krippendorff's alpha, interval (4 raters × 1000 items)100,000krippendorff0.4120.3870.94×0.260.693.3e-15
Agreement: Krippendorff's alpha, interval (4 raters × 10000 items)1,000,000krippendorff3.5113.3720.96×2.266.321.8e-14
Retrieval: MRR, MAP@20, NDCG@10 (10 queries)1,000ranx0.1821.4077.72×0.010.040
Retrieval: MRR, MAP@20, NDCG@10 (1000 queries)100,000ranx29.82177.0012.58×0.493.120
Retrieval: MRR, MAP@20, NDCG@10 (10000 queries)1,000,000ranx170.388861.0095.05×5.3631.016.9e-18
Structured: JSON Schema compliance (10 documents)1,000jsonschema0.1570.4212.67×0.000.000
Structured: JSON Schema compliance (1000 documents)100,000jsonschema14.39333.9952.36×0.050.020
Structured: JSON Schema compliance (10000 documents)1,000,000jsonschema144.622330.4622.29×0.500.160
Text: corpus BLEU and chrF (10 sentences)1,000sacreBLEU2.2061.9600.89×0.100.221.4e-14
Text: corpus BLEU and chrF (1000 sentences)100,000sacreBLEU255.898250.4940.98×0.7421.911.4e-14
Text: corpus BLEU and chrF (10000 sentences)1,000,000sacreBLEU2789.7903051.7551.09×7.26210.170
Text: METEOR, exact and stem matches (10 sentences)1,000NLTK0.4590.5011.09×0.010.010
Text: METEOR, exact and stem matches (1000 sentences)100,000NLTK52.27558.9651.13×0.120.040
Text: METEOR, exact and stem matches (10000 sentences)1,000,000NLTK488.126626.6931.28×1.150.390
Text: ROUGE-1, ROUGE-2, ROUGE-L (10 sentences)1,000rouge-score1.1450.9470.83×0.010.010
Text: ROUGE-1, ROUGE-2, ROUGE-L (1000 sentences)100,000rouge-score120.724111.8840.93×0.120.640
Text: ROUGE-1, ROUGE-2, ROUGE-L (10000 sentences)1,000,000rouge-score1290.1121095.6090.85×1.166.550

Reading the results

Many metrics at once is where EvalSuite is fastest. evaluate() validates the inputs once and builds the confusion matrix once, then derives all eight label metrics from it: 36–49× faster than eight separate scikit-learn calls. Sensitivity, specificity and both likelihood ratios together are 5–11× faster.

Calibration slope and intercept are 5–6× faster than statsmodels' GLM and use far less memory (84 MiB against 580 MiB at a million samples), because EvalSuite fits the two small logistic models with a dedicated Newton–Raphson solver.

Object detection matches pycocotools exactly on all twelve COCO numbers and is 0.9–1.6× as fast, using a fifth of its memory. Per-class Dice and IoU are 7–10× faster than building scikit-learn's confusion matrix. Hausdorff distance (0.6–0.8×) extracts surfaces and computes HD95 and ASSD alongside the maximum that SciPy returns.

LLM metrics give the reference libraries' numbers and run at about their speed: corpus BLEU and chrF 0.89–1.09× sacreBLEU, ROUGE 0.8–0.9× rouge-score, METEOR 1.09–1.28× NLTK, ranking metrics 2.6–7.7× ranx, JSON Schema compliance 2.3–2.7× jsonschema and Krippendorff's alpha 0.9–1.0× the krippendorff package.

Hypothesis tests call SciPy for the statistic and p-value, so they match its speed at best. The Welch t-test is about half SciPy's speed at large n because EvalSuite also validates the inputs and computes the confidence interval and Cohen's d.

Some rows are slower, and we show them. The diagnostic report (0.2–2.2× across sizes) validates labels and reports ten measures, where the reference computes seven intervals from counts it takes directly. Regression on a million values (0.5×) spends most of its ~20 ms checking every value for NaN, infinity, shape and dtype.

Reproduce on your machine

pip install evalsuite-python scikit-learn statsmodels pycocotools sacrebleu rouge-score nltk ranx krippendorff jsonschema
evalsuite benchmark                    # every case at 1,000 / 100,000 / 1,000,000 samples
evalsuite benchmark --suite clinical   # only the v0.2 clinical, calibration and statistics cases
evalsuite benchmark --suite vision     # only the v0.3 segmentation and detection cases
evalsuite benchmark --suite llm        # only the v0.4 LLM cases
evalsuite benchmark --quick            # small sizes only

The full table and notes are in BENCHMARKS.md in the package repository. Results on your hardware will differ in absolute time; the ratios are what to compare.