Skip to content
EvalSuite
Documentation menu

Clinical evaluation

Available since v0.2.0

Clinical metrics describe how a binary test or risk model would behave when used for decisions about people. They carry assumptions that general machine-learning metrics do not, so EvalSuite applies them only when they are requested explicitly.

Assumptions

  • The reference standard is treated as ground truth.
  • Diagnostic metrics are defined for binary outcomes.
  • Predictive values (PPV, NPV) depend on prevalence and do not transfer between populations with different prevalence.

Usage

Pythonv0.2.0
import evalsuite as es

report = es.diagnostic_report(y_true, y_pred)   # binary test vs reference standard
print(report)
report["lr_positive"]          # LR+ with its 95% CI (log method, Simel 1991)
report.save("diagnostic.html") # also .csv .md .tex .json

es.sensitivity(y_true, y_pred)
es.lr_positive(y_true, y_pred)
es.diagnostic_odds_ratio(y_true, y_pred, correction=0.5)   # Haldane–Anscombe when a cell is 0
es.youden_j(y_true, y_pred)

dca = es.decision_curve(y_true, {"model": y_prob})   # thresholds 0.01 … 0.99
dca.to_dataframe()      # threshold, net_benefit[model], treat_all, treat_none
dca.useful_range()      # thresholds where the model beats both references
es.plot.decision_curve(y_true, {"model": y_prob})

Confidence intervals in the diagnostic report

MeasureInterval
Sensitivity, specificity, PPV, NPV, accuracy, prevalenceWilson score (or Clopper–Pearson with proportion_method="clopper-pearson")
LR+ and LR−log method (Simel, Samsa and Matchar, 1991)
Diagnostic odds ratioWoolf's log method; 0.5 is added to every cell when one is zero
Youden's JWald, from the independent variances of sensitivity and specificity, clipped to [−1, 1]

Ratios that divide by zero (for example LR+ when specificity is 1) are reported as inf or NaN with a warning, never as 0.

Decision curve analysis

Net benefit at threshold probability pt is TP/N − (FP/N) · pt / (1 − pt). It is compared with treating everyone and treating no one. The threshold expresses how a clinician weighs the harm of a false positive against a false negative (Vickers and Elkin, 2006).

Metrics

  • SensitivityImplemented

    Proportion of people with the condition whom the test correctly identifies as positive.

    es.recall

  • SpecificityImplemented

    Proportion of people without the condition whom the test correctly identifies as negative.

    es.specificity

  • Probability that a person with a positive result has the condition.

    es.precision

  • Probability that a person with a negative result does not have the condition.

    es.npv

  • How much a positive result increases the odds of the condition.

    es.lr_positive

  • How much a negative result decreases the odds of the condition.

    es.lr_negative

  • Ratio of the odds of a positive result in people with the condition to the odds in people without it.

    es.diagnostic_odds_ratio

  • Youden's JImplemented

    Single summary of sensitivity and specificity at one threshold.

    es.youden_j

  • Clinical utility of a model across threshold probabilities, compared with treat-all and treat-none strategies.

    es.decision_curve