Skip to content
EvalSuite
Documentation menu

Core concepts

These concepts shipped in v0.1.0 and underpin every later module.

Validation happens once

Inputs are converted to NumPy and checked a single time per evaluation: lengths, shapes, label sets, probability ranges, NaN and infinity. Failures raise a specific exception that says what failed, why, and how to fix it.

Textv0.1.0
InputValidationError: y_true and y_pred must contain the same number of observations.
Received 776 and 775.

Labels must use one type throughout: mixing numbers and strings ([0, "a"], or string y_true with integer y_pred) raises InputValidationError rather than silently comparing them as text (since v0.3.1).

Exception types are EvalSuiteError, InputValidationError, MetricInputError, UnsupportedTaskError, OptionalDependencyError, and StatisticalTestError.

The evaluation context

EvaluationContext holds the validated inputs together with expensive intermediates such as the confusion matrix and class counts. Metrics read from the context instead of recomputing those intermediates. This removes redundant computation: evaluate() is 26–34× faster than separate scikit-learn calls for the same eight label metrics (see benchmarks).

Results are structured and portable

Every function returns a typed result object rather than a bare float.

ResultContains
MetricResultvalue (or per-class values), metric id, parameters used
EvaluationResultall metric results for one evaluation, plus metadata
ConfidenceIntervalestimate, lower, upper, level, method
TestResultstatistic, p-value, estimate, method
ComparisonResultper-model estimates with intervals, paired tests, adjusted p-values, best model
ClassificationReportper-class precision, recall, F1, specificity and support

Results export with to_dict(), to_dataframe(), to_json(), to_markdown(), to_latex(), to_html() and to_csv(). JSON uses standard serialization; pickle is never used for portable storage.

Explicit conventions

  • Averaging for multiclass and multilabel inputs must be stated: binary, micro, macro, weighted, samples, or None. Multiclass data is never silently treated as binary.
  • Undefined values such as a zero denominator follow zero_division="warn", 0, or np.nan. They are never silently reported as zero.
  • Randomness is controlled by random_state. Global random state is never modified.
  • No side effects: metric functions do not modify inputs, print, create files, or change plotting state.