Factuality and QA
Answer correctness
Implemented
factuality.answer_correctnessDefinition
Agreement of an answer with the reference at the level of statements: F1 over true-positive (in both), false-positive (only in the answer) and false-negative (only in the reference) statements, optionally blended with a semantic similarity score (RAGAS answer correctness).
Formula
w_f · TP/(TP + ½(FP + FN)) + w_s · similarity
Range: [0, 1]
Inputs and outputs
- tp, fp, fn: statement counts per answer
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.answer_correctness(tp, fp, fn, similarity=similarity)References
- Es S, James J, Espinosa-Anke L, Schockaert S. RAGAS: automated evaluation of retrieval augmented generation. EACL (demos). 2024:150-158.