Skip to content
EvalSuite
Documentation menu

Classification

Available in v0.1.0

Binary, multiclass and multilabel classification, with sample weights throughout.

Usage

Pythonv0.1.0
import evalsuite as es

es.accuracy(y_true, y_pred)
es.f1(y_true, y_pred, average="macro", zero_division="warn")
es.roc_auc(y_true, y_prob)

cm = es.confusion_matrix(y_true, y_pred, labels=[0, 1, 2])

report = es.classification_report(y_true, y_pred)   # precision, recall, F1, specificity, support
print(report.summary())

Averaging

For more than two classes, average controls how per-class values are combined. macro treats every class equally, weighted weights by support, micro pools counts across classes, samples averages per observation for multilabel data, and None returns per-class values. The default "auto" resolves to binary for binary targets and macro otherwise; the resolved value is stored in result.params["average"].

Metrics

  • AccuracyImplemented

    Fraction of observations whose predicted label equals the true label.

    es.accuracy

  • PrecisionImplemented

    Of the observations predicted positive, the fraction that are truly positive.

    es.precision

  • RecallImplemented

    Of the truly positive observations, the fraction predicted positive. Equal to sensitivity in the binary case.

    es.recall

  • F1 scoreImplemented

    Harmonic mean of precision and recall.

    es.f1

  • Mean of per-class recall. Reduces the influence of class imbalance compared with accuracy.

    es.balanced_accuracy

  • Correlation between observed and predicted binary labels that uses all four confusion-matrix cells.

    es.mcc

  • Cohen's kappaImplemented

    Agreement between predicted and true labels corrected for agreement expected by chance.

    es.cohen_kappa

  • ROC AUCImplemented

    Area under the receiver operating characteristic curve; the probability that a random positive is ranked above a random negative.

    es.roc_auc

  • Summary of the precision-recall curve, computed as average precision over recall steps.

    es.average_precision

  • Log lossImplemented

    Negative mean log-likelihood of the true labels under the predicted probabilities.

    es.log_loss

  • Confusion matrixImplemented

    Counts of true versus predicted labels. Computed once per evaluation and shared by every count-based metric.

    es.confusion_matrix

Choosing metrics

Accuracy is easy to read but misleading under class imbalance. Balanced accuracy, MCC and PR AUC are more informative when classes are imbalanced. ROC AUC measures ranking, not calibration; use the calibration guide to assess probabilities.