Skip to content
EvalSuite
Documentation menu

Factuality and QA

Answer correctness

Implementedfactuality.answer_correctness

Definition

Agreement of an answer with the reference at the level of statements: F1 over true-positive (in both), false-positive (only in the answer) and false-negative (only in the reference) statements, optionally blended with a semantic similarity score (RAGAS answer correctness).

Formula

w_f · TP/(TP + ½(FP + FN)) + w_s · similarity

Range: [0, 1]

Inputs and outputs

  • tp, fp, fn: statement counts per answer

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.answer_correctness(tp, fp, fn, similarity=similarity)

References

  1. Es S, James J, Espinosa-Anke L, Schockaert S. RAGAS: automated evaluation of retrieval augmented generation. EACL (demos). 2024:150-158.

Implementation status