Classification
Available in v0.1.0Binary, multiclass and multilabel classification, with sample weights throughout.
Usage
import evalsuite as es
es.accuracy(y_true, y_pred)
es.f1(y_true, y_pred, average="macro", zero_division="warn")
es.roc_auc(y_true, y_prob)
cm = es.confusion_matrix(y_true, y_pred, labels=[0, 1, 2])
report = es.classification_report(y_true, y_pred) # precision, recall, F1, specificity, support
print(report.summary())Averaging
For more than two classes, average controls how per-class values are combined. macro treats every class equally, weighted weights by support, micro pools counts across classes, samples averages per observation for multilabel data, and None returns per-class values. The default "auto" resolves to binary for binary targets and macro otherwise; the resolved value is stored in result.params["average"].
Metrics
| Metric | Description | Status | API |
|---|---|---|---|
| Accuracy | Fraction of observations whose predicted label equals the true label. | Implemented | es.accuracy |
| Precision | Of the observations predicted positive, the fraction that are truly positive. | Implemented | es.precision |
| Recall | Of the truly positive observations, the fraction predicted positive. Equal to sensitivity in the binary case. | Implemented | es.recall |
| F1 score | Harmonic mean of precision and recall. | Implemented | es.f1 |
| Balanced accuracy | Mean of per-class recall. Reduces the influence of class imbalance compared with accuracy. | Implemented | es.balanced_accuracy |
| Matthews correlation coefficient | Correlation between observed and predicted binary labels that uses all four confusion-matrix cells. | Implemented | es.mcc |
| Cohen's kappa | Agreement between predicted and true labels corrected for agreement expected by chance. | Implemented | es.cohen_kappa |
| ROC AUC | Area under the receiver operating characteristic curve; the probability that a random positive is ranked above a random negative. | Implemented | es.roc_auc |
| PR AUC (average precision) | Summary of the precision-recall curve, computed as average precision over recall steps. | Implemented | es.average_precision |
| Log loss | Negative mean log-likelihood of the true labels under the predicted probabilities. | Implemented | es.log_loss |
| Confusion matrix | Counts of true versus predicted labels. Computed once per evaluation and shared by every count-based metric. | Implemented | es.confusion_matrix |
- AccuracyImplemented
Fraction of observations whose predicted label equals the true label.
es.accuracy
- PrecisionImplemented
Of the observations predicted positive, the fraction that are truly positive.
es.precision
- RecallImplemented
Of the truly positive observations, the fraction predicted positive. Equal to sensitivity in the binary case.
es.recall
- F1 scoreImplemented
Harmonic mean of precision and recall.
es.f1
- Balanced accuracyImplemented
Mean of per-class recall. Reduces the influence of class imbalance compared with accuracy.
es.balanced_accuracy
- Matthews correlation coefficientImplemented
Correlation between observed and predicted binary labels that uses all four confusion-matrix cells.
es.mcc
- Cohen's kappaImplemented
Agreement between predicted and true labels corrected for agreement expected by chance.
es.cohen_kappa
- ROC AUCImplemented
Area under the receiver operating characteristic curve; the probability that a random positive is ranked above a random negative.
es.roc_auc
- PR AUC (average precision)Implemented
Summary of the precision-recall curve, computed as average precision over recall steps.
es.average_precision
- Log lossImplemented
Negative mean log-likelihood of the true labels under the predicted probabilities.
es.log_loss
- Confusion matrixImplemented
Counts of true versus predicted labels. Computed once per evaluation and shared by every count-based metric.
es.confusion_matrix
Choosing metrics
Accuracy is easy to read but misleading under class imbalance. Balanced accuracy, MCC and PR AUC are more informative when classes are imbalanced. ROC AUC measures ranking, not calibration; use the calibration guide to assess probabilities.