Unified evaluation for modern machine learning and research.
EvalSuite brings machine learning, clinical, statistical, segmentation, and object-detection evaluation into one consistent evaluation framework.
$ pip install evalsuite-pythonView on PyPIPython 3.9 or newer. Optional extras: [plot], [all].
Counting visitors
- Classification
- Regression
- Clinical
- Statistics
- Segmentation
- Detection
import evalsuite as es
result = es.evaluate(y_true, y_pred, y_prob=y_prob)
print(result.summary())
for m in ["accuracy", "precision", "recall", "f1", "mcc"]:
print(es.bootstrap_ci(m, y_true, y_pred, random_state=0))
print(es.bootstrap_ci("roc_auc", y_true, y_prob=y_prob, random_state=0))
result.save("results.tex")result.summary()
example, n = 200
| Metric | Estimate | 95% CI |
|---|---|---|
| Accuracy | 0.845 | 0.790–0.890 |
| Precision | 0.885 | 0.817–0.935 |
| Recall | 0.810 | 0.724–0.876 |
| F1 score | 0.846 | 0.786–0.891 |
| MCC | 0.693 | 0.583–0.783 |
| ROC AUC | 0.926 | 0.886–0.953 |
One study. Six toolchains.
Research evaluation is usually stitched together from excellent but separate libraries, each with its own conventions. EvalSuite is designed as one consistent layer on top, with those libraries used as references for numerical validation.
scikit-learn
Classification and regression metrics
SciPy
Hypothesis tests and distributions
statsmodels
Diagnostics and multiple-testing corrections
TorchMetrics
Metrics inside training loops
pycocotools
COCO-style detection evaluation
Custom scripts
Bootstrap, DCA, report tables
What you gain with EvalSuite
Evaluation that is easier to get right, easier to report and easier for others to check.
Fewer silent errors
Explicit averaging, zero-division and empty-mask rules replace defaults that differ between libraries.
Uncertainty by default
Confidence intervals, effect sizes and paired tests sit next to every estimate, as reviewers expect.
Less glue code
One validated input, one result object and one report, instead of scripts stitching five libraries together.
Reproducible numbers
Seeded resampling, recorded parameters and versioned exports make results easy to rerun and check.
Who benefits
ML researchers
- Classification and regression metrics with stated conventions
- Paired model comparison with effect sizes
- LaTeX tables straight into the manuscript
Clinical AI teams
- Sensitivity, specificity, PPV, NPV and likelihood ratios
- Calibration, Hosmer–Lemeshow and decision curve analysis
- Metrics applied only when their assumptions hold
Computer vision
- Dice, IoU, boundary and surface metrics
- AP and mAP with a documented matching convention
- Per-class results with explicit empty-mask handling
Reviewers and statisticians
- Every metric documented with formula, assumptions and references
- Test results that report statistics, not just p-values
- Methodology recorded alongside the numbers
Students and educators
- A reference that explains when each metric is appropriate
- An interactive playground to explore metrics on sample data
- Limitations stated, not hidden
Research software teams
- A typed HTTP API with personal keys
- A metric registry that keeps code and docs in sync
- Open source under the MIT License
v0.1.0 delivers these for classification, regression and model comparison, backed by 1,300+ tests and published benchmarks; later releases extend them to more tasks.
From predictions to paper.
Every evaluation follows the same path. Each layer has one job, so new metrics and exporters plug in without touching the rest.
- 1
Input
Arrays, DataFrames, masks or boxes, converted to NumPy once.
- 2
Validation
Shapes, labels, probability ranges, NaN and Inf checked with explicit errors.
- 3
Evaluation context
Validated data plus shared intermediates such as the confusion matrix.
- 4
Metrics
Task-appropriate metrics computed from the shared context.
- 5
Statistical analysis
Hypothesis tests, effect sizes and multiple-testing correction.
- 6
Uncertainty
Analytical intervals and seeded bootstrap estimates.
- 7
Visualization
ROC, PR, calibration and decision curves as Matplotlib figures.
- 8
Reporting
One structured result exported to JSON, CSV, Markdown, HTML or LaTeX.
Compute the confusion matrix once. Reuse it everywhere.
The EvaluationContext holds validated inputs and expensive intermediates. Precision, recall, specificity, F1, PPV and NPV all read the same four counts instead of each recounting the data.
This makes evaluate() 26–34× faster than separate scikit-learn calls for the same metrics, with identical results. See the benchmarks.
Five domains, one framework.
What each release covers. Statuses move to implemented only when a version is published.
Machine learning
- ClassificationSince v0.1.0
- RegressionSince v0.1.0
- Model comparisonSince v0.1.0
Clinical
- Diagnostic metricsSince v0.2.0
- Calibration curve and ECESince v0.1.0
- Hosmer–LemeshowSince v0.2.0
- Decision curve analysisSince v0.2.0
Statistics
- Paired tests (McNemar, DeLong, bootstrap)Since v0.1.0
- Further hypothesis testsSince v0.2.0
- Effect sizesSince v0.1.0
- Confidence intervalsSince v0.1.0
- BootstrapSince v0.1.0
- Multiple testingSince v0.1.0
Computer vision
- Semantic segmentationSince v0.3.0
- Boundary metricsSince v0.3.0
- Surface metricsSince v0.3.0
- Object detectionSince v0.3.0
- mAPSince v0.3.0
LLM evaluation
- Text generation (BLEU, ROUGE, METEOR, chrF, TER, CIDEr)Since v0.4.0
- Semantic (BERTScore, MoverScore, MAUVE)Since v0.4.0
- Factuality and citationsSince v0.4.0
- LLM-as-a-judge and preferencesSince v0.4.0
- Reasoning (pass@k, GSM8K, MATH)Since v0.4.0
- Retrieval and RAGSince v0.4.0
- Structured output and tool callsSince v0.4.0
Research output
- VisualizationSince v0.1.0
- Reports (HTML, Markdown)Since v0.1.0
- LaTeX tablesSince v0.1.0
- JSON and CSV exportSince v0.1.0
- Reproducibility controlsSince v0.1.0
Built for results you can defend.
The design commitments behind the package. Each one is verified by tests as its release lands.
Unified API
Low-level functions such as es.classification.f1 and a high-level es.evaluate share one parameter convention and one result type.
Shared computation
Inputs validated once; intermediates cached in an evaluation context.
Numerical reliability
Zero denominators and empty masks raise explicit warnings, never silent zeros.
Reproducibility
Every stochastic step takes random_state and never touches global RNG state.
Statistical rigour
Tests return statistics, degrees of freedom, effect sizes and intervals, not just p-values.
Clinical evaluation
Diagnostic metrics, calibration and decision curves, used only when assumptions hold.
Computer vision
Segmentation and detection with documented empty-mask and matching rules.
Publication-ready output
LaTeX tables, HTML, Markdown, JSON and CSV from one structured result.
Extensible registry
Every metric is described in a registry that powers code and docs alike.
Research documentation
Each metric documents its formula, assumptions, edge cases, limitations and primary references.
Release plan
Three releases, each shipped with tests, documentation and a PyPI build.
v0.1.0
ImplementedCore
Released 8 October 2026. The foundation every later module plugs into: result types, validation, the metric registry, and the first two task families.
v0.2.0
ImplementedClinical and statistics
Released 8 October 2026. Diagnostic and calibration metrics, uncertainty quantification, and statistical testing.
v0.3.0
ImplementedComputer vision
Released 9 October 2026. Segmentation and detection evaluation, with comparison, plotting and reporting extended to every task.
v0.4.0
ImplementedLLM evaluation
Released 9 October 2026. Text generation, semantic similarity, factuality and hallucination, LLM-as-a-judge, reasoning benchmarks, retrieval-augmented generation and structured output, with the same intervals, comparison, plots and reports as every other task.
v0.5.0
PlannedLLM systems: safety, agents and operations
Planned. Safety and responsible AI, robustness, calibration and uncertainty for LLMs, agent and tool use, multilingual, code generation, long-context and summarization, and inference efficiency and cost.
8 areas, 85 metrics planned
Try the evaluation workflow now.
The playground runs a labelled demo engine in your browser. No data leaves the page.