Core concepts
These concepts shipped in v0.1.0 and underpin every later module.
Validation happens once
Inputs are converted to NumPy and checked a single time per evaluation: lengths, shapes, label sets, probability ranges, NaN and infinity. Failures raise a specific exception that says what failed, why, and how to fix it.
InputValidationError: y_true and y_pred must contain the same number of observations.
Received 776 and 775.Labels must use one type throughout: mixing numbers and strings ([0, "a"], or string y_true with integer y_pred) raises InputValidationError rather than silently comparing them as text (since v0.3.1).
Exception types are EvalSuiteError, InputValidationError, MetricInputError, UnsupportedTaskError, OptionalDependencyError, and StatisticalTestError.
The evaluation context
EvaluationContext holds the validated inputs together with expensive intermediates such as the confusion matrix and class counts. Metrics read from the context instead of recomputing those intermediates. This removes redundant computation: evaluate() is 26–34× faster than separate scikit-learn calls for the same eight label metrics (see benchmarks).
Results are structured and portable
Every function returns a typed result object rather than a bare float.
| Result | Contains |
|---|---|
MetricResult | value (or per-class values), metric id, parameters used |
EvaluationResult | all metric results for one evaluation, plus metadata |
ConfidenceInterval | estimate, lower, upper, level, method |
TestResult | statistic, p-value, estimate, method |
ComparisonResult | per-model estimates with intervals, paired tests, adjusted p-values, best model |
ClassificationReport | per-class precision, recall, F1, specificity and support |
Results export with to_dict(), to_dataframe(), to_json(), to_markdown(), to_latex(), to_html() and to_csv(). JSON uses standard serialization; pickle is never used for portable storage.
Explicit conventions
- Averaging for multiclass and multilabel inputs must be stated:
binary,micro,macro,weighted,samples, orNone. Multiclass data is never silently treated as binary. - Undefined values such as a zero denominator follow
zero_division="warn",0, ornp.nan. They are never silently reported as zero. - Randomness is controlled by
random_state. Global random state is never modified. - No side effects: metric functions do not modify inputs, print, create files, or change plotting state.