Skip to content
EvalSuite
Documentation menu

Calibration and uncertainty

Confidence–accuracy correlation

Implementeduncertainty.confidence_accuracy_correlation

Definition

How well confidence discriminates correct from incorrect answers: the AUROC of confidence for correctness (P(true) discrimination, Kadavath et al.), with the point-biserial and Spearman correlations.

Formula

AUROC = P(conf_correct > conf_wrong) + ½ P(tie)

Range: [0, 1] (0.5 = no signal)

Inputs and outputs

  • correct: see the signature of es.confidence_accuracy_correlation
  • confidence: see the signature of es.confidence_accuracy_correlation

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.confidence_accuracy_correlation(correct, confidence)

References

  1. Kadavath S, Conerly T, Askell A, et al. Language models (mostly) know what they know. arXiv:2207.05221. 2022.

Implementation status