Research and methodology
Motivation
Model evaluation is where many research conclusions are decided, yet it is usually assembled from separate libraries and project-specific scripts. Differences in averaging, zero-division handling, empty-mask conventions or interval methods can change reported numbers without anyone noticing.
Problem
Each established library is sound within its scope. The difficulty is the seams between them: inputs validated several times in different ways, conventions that differ silently, and uncertainty estimates added by hand, if at all.
Architecture
EvalSuite is designed around a single path: input, validation, evaluation context, metrics, statistical analysis, uncertainty, visualization and reporting. Each layer has one responsibility and communicates through typed result objects.
Unified evaluation
The same parameter conventions and result types apply across classification, regression, clinical, statistical, segmentation and detection metrics, at both the function level and through es.evaluate.
Shared computation
Intermediates such as confusion matrices and class counts are computed once per evaluation and reused. This makes evaluate() 26–34× faster than separate scikit-learn calls for the same metrics, with identical results (see the benchmarks page).
Numerical validation
Each metric is tested against analytically derived cases and, where conventions match, against scikit-learn, SciPy and statsmodels (1,300+ tests on Python 3.9 to 3.14, Linux, Windows and macOS). Intentional differences in convention are documented rather than hidden.
Reproducibility
Every stochastic procedure accepts a seed, results record the parameters that produced them, and exports carry the software version.
Clinical and statistical evaluation
Clinical metrics are applied only when their assumptions hold and are documented with their dependence on prevalence. Statistical tests report effect sizes and intervals alongside p-values, and are not presented as decision rules.
Benchmarking
Performance claims come only from reproducible runs that list hardware, software versions and workloads; anyone can rerun them with evalsuite benchmark. See the benchmark results.
Intended contributions
A single, documented evaluation layer spanning several research domains; a metric registry that keeps documentation in sync with code; and publication-ready reporting with uncertainty included by default. v0.1.0 delivers these for classification, regression and model comparison.
Limitations
EvalSuite v0.4.0 covers classification, regression, clinical evaluation, calibration, model comparison, statistical testing, segmentation, object detection and LLM evaluation. It does not replace domain expertise, clinical validation or study design. A unified interface cannot remove the need to choose metrics that fit the question being asked.