Calibration and uncertainty
Confidence–accuracy correlation
Implemented
uncertainty.confidence_accuracy_correlationDefinition
How well confidence discriminates correct from incorrect answers: the AUROC of confidence for correctness (P(true) discrimination, Kadavath et al.), with the point-biserial and Spearman correlations.
Formula
AUROC = P(conf_correct > conf_wrong) + ½ P(tie)
Range: [0, 1] (0.5 = no signal)
Inputs and outputs
- correct: see the signature of es.confidence_accuracy_correlation
- confidence: see the signature of es.confidence_accuracy_correlation
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.confidence_accuracy_correlation(correct, confidence)References
- Kadavath S, Conerly T, Askell A, et al. Language models (mostly) know what they know. arXiv:2207.05221. 2022.