Factuality and QA
Faithfulness
Implemented
factuality.faithfulnessDefinition
Share of an answer's claims that the evidence supports (FActScore factual precision; RAGAS faithfulness when the evidence is the retrieved context).
Formula
supported claims / claims
Range: [0, 1]
Inputs and outputs
- claim_verdicts: per answer, the verdict of each claim (supported / contradicted / unsupported)
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.faithfulness(claim_verdicts)References
- Min S, Krishna K, Lyu X, et al. FActScore: fine-grained atomic evaluation of factual precision in long form text generation. EMNLP. 2023:12076-12100.
- Es S, James J, Espinosa-Anke L, Schockaert S. RAGAS: automated evaluation of retrieval augmented generation. EACL (demos). 2024:150-158.