Skip to content
EvalSuite
v0.4.0 releasedStable on PyPI · LLM evaluation in v0.4.0

Unified evaluation for modern machine learning and research.

EvalSuite brings machine learning, clinical, statistical, segmentation, and object-detection evaluation into one consistent evaluation framework.

$ pip install evalsuite-python

View on PyPIPython 3.9 or newer. Optional extras: [plot], [all].

Counting visitors

  • Classification
  • Regression
  • Clinical
  • Statistics
  • Segmentation
  • Detection
evaluate.py
Pythonv0.1.0
import evalsuite as es

result = es.evaluate(y_true, y_pred, y_prob=y_prob)
print(result.summary())

for m in ["accuracy", "precision", "recall", "f1", "mcc"]:
    print(es.bootstrap_ci(m, y_true, y_pred, random_state=0))
print(es.bootstrap_ci("roc_auc", y_true, y_prob=y_prob, random_state=0))

result.save("results.tex")

result.summary()

example, n = 200

Example output: estimates with 95% bootstrap confidence intervals on a 200-sample test set.
MetricEstimate95% CI
Accuracy0.8450.790–0.890
Precision0.8850.817–0.935
Recall0.8100.724–0.876
F1 score0.8460.786–0.891
MCC0.6930.583–0.783
ROC AUC0.9260.886–0.953

One study. Six toolchains.

Research evaluation is usually stitched together from excellent but separate libraries, each with its own conventions. EvalSuite is designed as one consistent layer on top, with those libraries used as references for numerical validation.

Delivered in v0.1.0, extended through v0.4.0

What you gain with EvalSuite

Evaluation that is easier to get right, easier to report and easier for others to check.

  • Fewer silent errors

    Explicit averaging, zero-division and empty-mask rules replace defaults that differ between libraries.

  • Uncertainty by default

    Confidence intervals, effect sizes and paired tests sit next to every estimate, as reviewers expect.

  • Less glue code

    One validated input, one result object and one report, instead of scripts stitching five libraries together.

  • Reproducible numbers

    Seeded resampling, recorded parameters and versioned exports make results easy to rerun and check.

Who benefits

  • ML researchers

    • Classification and regression metrics with stated conventions
    • Paired model comparison with effect sizes
    • LaTeX tables straight into the manuscript
  • Clinical AI teams

    • Sensitivity, specificity, PPV, NPV and likelihood ratios
    • Calibration, Hosmer–Lemeshow and decision curve analysis
    • Metrics applied only when their assumptions hold
  • Computer vision

    • Dice, IoU, boundary and surface metrics
    • AP and mAP with a documented matching convention
    • Per-class results with explicit empty-mask handling
  • Reviewers and statisticians

    • Every metric documented with formula, assumptions and references
    • Test results that report statistics, not just p-values
    • Methodology recorded alongside the numbers
  • Students and educators

    • A reference that explains when each metric is appropriate
    • An interactive playground to explore metrics on sample data
    • Limitations stated, not hidden
  • Research software teams

    • A typed HTTP API with personal keys
    • A metric registry that keeps code and docs in sync
    • Open source under the MIT License

v0.1.0 delivers these for classification, regression and model comparison, backed by 1,300+ tests and published benchmarks; later releases extend them to more tasks.

From predictions to paper.

Every evaluation follows the same path. Each layer has one job, so new metrics and exporters plug in without touching the rest.

  1. 1

    Input

    Arrays, DataFrames, masks or boxes, converted to NumPy once.

  2. 2

    Validation

    Shapes, labels, probability ranges, NaN and Inf checked with explicit errors.

  3. 3

    Evaluation context

    Validated data plus shared intermediates such as the confusion matrix.

  4. 4

    Metrics

    Task-appropriate metrics computed from the shared context.

  5. 5

    Statistical analysis

    Hypothesis tests, effect sizes and multiple-testing correction.

  6. 6

    Uncertainty

    Analytical intervals and seeded bootstrap estimates.

  7. 7

    Visualization

    ROC, PR, calibration and decision curves as Matplotlib figures.

  8. 8

    Reporting

    One structured result exported to JSON, CSV, Markdown, HTML or LaTeX.

Compute the confusion matrix once. Reuse it everywhere.

The EvaluationContext holds validated inputs and expensive intermediates. Precision, recall, specificity, F1, PPV and NPV all read the same four counts instead of each recounting the data.

This makes evaluate() 26–34× faster than separate scikit-learn calls for the same metrics, with identical results. See the benchmarks.

Shared confusion matrixA two by two confusion matrix with cells TP, FP, FN and TN connects to six derived metrics: precision, recall, specificity, F1, PPV and NPV.TPFPFNTNcomputed oncePrecisionTP / (TP+FP)RecallTP / (TP+FN)SpecificityTN / (TN+FP)F12TP / (2TP+FP+FN)PPVTP / (TP+FP)NPVTN / (TN+FN)

Five domains, one framework.

What each release covers. Statuses move to implemented only when a version is published.

Machine learning

  • ClassificationSince v0.1.0
  • RegressionSince v0.1.0
  • Model comparisonSince v0.1.0

Clinical

  • Diagnostic metricsSince v0.2.0
  • Calibration curve and ECESince v0.1.0
  • Hosmer–LemeshowSince v0.2.0
  • Decision curve analysisSince v0.2.0

Statistics

  • Paired tests (McNemar, DeLong, bootstrap)Since v0.1.0
  • Further hypothesis testsSince v0.2.0
  • Effect sizesSince v0.1.0
  • Confidence intervalsSince v0.1.0
  • BootstrapSince v0.1.0
  • Multiple testingSince v0.1.0

Computer vision

  • Semantic segmentationSince v0.3.0
  • Boundary metricsSince v0.3.0
  • Surface metricsSince v0.3.0
  • Object detectionSince v0.3.0
  • mAPSince v0.3.0

LLM evaluation

  • Text generation (BLEU, ROUGE, METEOR, chrF, TER, CIDEr)Since v0.4.0
  • Semantic (BERTScore, MoverScore, MAUVE)Since v0.4.0
  • Factuality and citationsSince v0.4.0
  • LLM-as-a-judge and preferencesSince v0.4.0
  • Reasoning (pass@k, GSM8K, MATH)Since v0.4.0
  • Retrieval and RAGSince v0.4.0
  • Structured output and tool callsSince v0.4.0

Research output

  • VisualizationSince v0.1.0
  • Reports (HTML, Markdown)Since v0.1.0
  • LaTeX tablesSince v0.1.0
  • JSON and CSV exportSince v0.1.0
  • Reproducibility controlsSince v0.1.0

Built for results you can defend.

The design commitments behind the package. Each one is verified by tests as its release lands.

Release plan

Three releases, each shipped with tests, documentation and a PyPI build.

View full roadmap
  1. v0.1.0

    Implemented

    Core

    Released 8 October 2026. The foundation every later module plugs into: result types, validation, the metric registry, and the first two task families.

  2. v0.2.0

    Implemented

    Clinical and statistics

    Released 8 October 2026. Diagnostic and calibration metrics, uncertainty quantification, and statistical testing.

  3. v0.3.0

    Implemented

    Computer vision

    Released 9 October 2026. Segmentation and detection evaluation, with comparison, plotting and reporting extended to every task.

  4. v0.4.0

    Implemented

    LLM evaluation

    Released 9 October 2026. Text generation, semantic similarity, factuality and hallucination, LLM-as-a-judge, reasoning benchmarks, retrieval-augmented generation and structured output, with the same intervals, comparison, plots and reports as every other task.

  5. v0.5.0

    Planned

    LLM systems: safety, agents and operations

    Planned. Safety and responsible AI, robustness, calibration and uncertainty for LLMs, agent and tool use, multilingual, code generation, long-context and summarization, and inference efficiency and cost.

    8 areas, 85 metrics planned

Try the evaluation workflow now.

The playground runs a labelled demo engine in your browser. No data leaves the page.