Factuality and QA
Groundedness
Implemented
factuality.groundednessDefinition
Average support score of an answer's content given the supplied context, from graded scores in [0, 1] (e.g. NLI entailment probabilities or judge ratings) rather than binary verdicts.
Formula
mean over claims (or sentences) of support score
Range: [0, 1]
Inputs and outputs
- support_scores: per answer, support scores in [0, 1]
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.groundedness(support_scores)References
- Es S, James J, Espinosa-Anke L, Schockaert S. RAGAS: automated evaluation of retrieval augmented generation. EACL (demos). 2024:150-158.