Skip to content
EvalSuite

Research and methodology

Why EvalSuite is being built, how it is designed, and what it does and does not claim.

Motivation

Model evaluation is where many research conclusions are decided, yet it is usually assembled from separate libraries and project-specific scripts. Differences in averaging, zero-division handling, empty-mask conventions or interval methods can change reported numbers without anyone noticing.

Problem

Each established library is sound within its scope. The difficulty is the seams between them: inputs validated several times in different ways, conventions that differ silently, and uncertainty estimates added by hand, if at all.

Architecture

EvalSuite is designed around a single path: input, validation, evaluation context, metrics, statistical analysis, uncertainty, visualization and reporting. Each layer has one responsibility and communicates through typed result objects.

Unified evaluation

The same parameter conventions and result types apply across classification, regression, clinical, statistical, segmentation and detection metrics, at both the function level and through es.evaluate.

Shared computation

Intermediates such as confusion matrices and class counts are computed once per evaluation and reused. This makes evaluate() 26–34× faster than separate scikit-learn calls for the same metrics, with identical results (see the benchmarks page).

Numerical validation

Each metric is tested against analytically derived cases and, where conventions match, against scikit-learn, SciPy and statsmodels (1,300+ tests on Python 3.9 to 3.14, Linux, Windows and macOS). Intentional differences in convention are documented rather than hidden.

Reproducibility

Every stochastic procedure accepts a seed, results record the parameters that produced them, and exports carry the software version.

Clinical and statistical evaluation

Clinical metrics are applied only when their assumptions hold and are documented with their dependence on prevalence. Statistical tests report effect sizes and intervals alongside p-values, and are not presented as decision rules.

Benchmarking

Performance claims come only from reproducible runs that list hardware, software versions and workloads; anyone can rerun them with evalsuite benchmark. See the benchmark results.

Intended contributions

A single, documented evaluation layer spanning several research domains; a metric registry that keeps documentation in sync with code; and publication-ready reporting with uncertainty included by default. v0.1.0 delivers these for classification, regression and model comparison.

Limitations

EvalSuite v0.4.0 covers classification, regression, clinical evaluation, calibration, model comparison, statistical testing, segmentation, object detection and LLM evaluation. It does not replace domain expertise, clinical validation or study design. A unified interface cannot remove the need to choose metrics that fit the question being asked.