Documentation
EvalSuite is a Python package that brings machine learning, clinical, statistical, computer-vision and language-model evaluation into one consistent framework. These docs describe evalsuite-python v0.4.0, installed with pip install evalsuite-python.
Where to start
- Getting started explains installation and the two API levels.
- Core concepts covers validation, the evaluation context, results, and conventions.
- Metric reference lists every metric with its formula, assumptions, and references, and whether it is implemented.
Evaluation guides
| Guide | Status |
|---|---|
| Classification and Regression | Available since v0.1.0 |
| Model comparison and Reporting | Available since v0.1.0 |
| Bootstrap and confidence intervals | Available since v0.1.0 |
| Statistical analysis | Available since v0.2.0 (paired tests since v0.1.0) |
| Calibration | Available since v0.2.0 (curve and ECE since v0.1.0) |
| Clinical evaluation | Available since v0.2.0 |
| Segmentation and Object detection | Available since v0.3.0 |
| LLM evaluation: text generation, factuality, LLM-as-a-judge, RAG, structured output | Available since v0.4.0 |
| CLI reference and API reference | Available since v0.1.0 |
How these docs are maintained
Every code example marked with a version was run against the released package before publishing. The package keeps its own metric registry (es.list_metrics(), es.metric_info(...)), and the metric reference marks exactly the metrics that registry provides as implemented.