Factuality and QA
Abstention accuracy
Implemented
factuality.abstention_accuracyDefinition
Whether a system abstains exactly when it should (the evidence is insufficient or it would otherwise be wrong): accuracy of the abstain / answer decision; abstention precision, recall and the accuracy on answered questions are reported alongside.
Formula
mean[abstained_i = should_abstain_i]
Range: [0, 1]
Inputs and outputs
- should_abstain: whether each question should be declined
- abstained: whether the system declined
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.abstention_accuracy(should_abstain, abstained, correct=correct)References
- Feng S, Shi W, Wang Y, Ding W, Balachandran V, Tsvetkov Y. Don't hallucinate, abstain: identifying LLM knowledge gaps via multi-LLM collaboration. ACL. 2024:14664-14690.