Skip to content
EvalSuite

About EvalSuite

An open-source Python package for consistent, well-documented model evaluation across research domains. Current version: v0.4.0.

What it is

EvalSuite brings machine learning, clinical, statistical, segmentation and object-detection evaluation into one framework with a shared API, structured results and publication-ready reporting.

Why it exists

To make evaluation easier to do correctly: fewer silent convention mismatches, uncertainty reported by default, and every metric documented with its assumptions and limitations.

Who it is for

Machine learning and computer vision researchers, clinical AI researchers, statisticians, data scientists, students, educators and reviewers who need to check how numbers were produced.

Philosophy

Scientific correctness comes before speed. Numerical edge cases are surfaced, not hidden. Claims about performance or validity are made only with evidence.

Open source

The package is released under the MIT License, with public issue tracking and a changelog. Source code: github.com/mkcs28/evalsuite-python.

Credits

Authors and maintainers: Manoj Kumar C S and Nikhil D Bharadwaj.

Roadmap

Core metrics and infrastructure shipped in v0.1.0, clinical and statistical evaluation in v0.2.0, computer vision in v0.3.0 and LLM evaluation in v0.4.0. See the roadmap.