Skip to content
EvalSuite
Documentation menu

Factuality, question answering and reasoning

Available since v0.4.0

Factuality metrics score verdicts produced by a verification step: annotators, an NLI model, a retrieval check or an LLM judge decide whether each claim of an answer is supported. EvalSuite aggregates them consistently (per answer or pooled), so verification pipelines can be compared with intervals. Report how claims were extracted and judged next to the numbers.

Python
import evalsuite as es

claim_verdicts = [["supported", "supported", "contradicted"], [True, False]]
es.faithfulness(claim_verdicts)            # FActScore / RAGAS faithfulness
es.hallucination_rate(claim_verdicts)      # unsupported-claim rate
es.groundedness([[0.9, 0.7], [0.2]])       # graded support scores

citations = [[{"supported": True, "citations": [True, False]}, {"supported": False, "citations": []}]]
es.citation_recall(citations)              # ALCE
es.citation_precision(citations)
es.abstention_accuracy([True, False, False], [True, False, True], correct=[False, True, False])

Question answering

Python
references = [["Paris", "Paris, France"], ["1990"]]
predictions = ["paris", "in 1991"]
es.exact_match(references, predictions)    # SQuAD normalisation: lowercase, no punctuation or articles
es.token_f1(references, predictions)
es.answer_correctness(tp=[3, 2], fp=[1, 0], fn=[0, 1], similarity=[0.9, 0.8])   # RAGAS

Reasoning benchmarks

Python
es.pass_at_k(n_samples=[10, 10], n_correct=[3, 0], k=5)              # unbiased estimator (Chen et al. 2021)
es.majority_vote_accuracy(["42", "7"], [["42", "41", "42"], ["7", "8", "8"]])
es.benchmark_accuracy(["... #### 18", "#### 42"], ["so the answer is 18", "41"], style="gsm8k")
es.extract_answer(r"thus \boxed{\frac{1}{2}}", style="boxed")       # MATH convention

Metrics

  • Whether a system abstains exactly when it should (the evidence is insufficient or it would otherwise be wrong): accuracy of the abstain / answer decision; abstention precision, recall and the accuracy on answered questions are reported alongside.

    es.abstention_accuracy

  • Agreement of an answer with the reference at the level of statements: F1 over true-positive (in both), false-positive (only in the answer) and false-negative (only in the reference) statements, optionally blended with a semantic similarity score (RAGAS answer correctness).

    es.answer_correctness

  • Answer relevanceImplemented

    How directly an answer addresses the question: mean cosine similarity between the question's embedding and embeddings of questions regenerated from the answer (RAGAS answer relevancy).

    es.answer_relevance

  • Share of citations that are relevant to the statement they are attached to (ALCE: a citation counts if it supports, or is needed to support, the statement).

    es.citation_precision

  • Citation recallImplemented

    Share of statements whose cited passages, taken together, support them (statements without citations count as unsupported); ALCE citation recall.

    es.citation_recall

  • Accuracy of supported / contradicted / not-enough-info verdicts against gold labels (FEVER label accuracy); macro F1 is reported alongside because the classes are often imbalanced.

    es.claim_verification_accuracy

  • Exact matchImplemented

    Share of predictions identical to a reference after normalization (SQuAD: lowercase, no punctuation or articles, single spaces); with several references, any match counts.

    es.exact_match

  • FaithfulnessImplemented

    Share of an answer's claims that the evidence supports (FActScore factual precision; RAGAS faithfulness when the evidence is the retrieved context).

    es.faithfulness

  • GroundednessImplemented

    Average support score of an answer's content given the supplied context, from graded scores in [0, 1] (e.g. NLI entailment probabilities or judge ratings) rather than binary verdicts.

    es.groundedness

  • Share of an answer's claims that the evidence does not support, i.e. contradicted or unverifiable claims, under the stated verification protocol.

    es.hallucination_rate

  • Share of facts extracted from the outputs (e.g. subject–relation–object triples) that are present in a specified trusted knowledge source.

    es.knowledge_consistency

  • Token F1Implemented

    Harmonic mean of token precision and recall between the normalized prediction and reference (bag of words, SQuAD); best reference per example, averaged over examples.

    es.token_f1

  • Share of benchmark questions answered correctly after extracting the final answer with the benchmark's convention (GSM8K '####' or last number, MATH \boxed{}, multiple-choice letter) and comparing with the gold answer.

    es.benchmark_accuracy

  • Accuracy of the most frequent answer among several samples per question (self-consistency); ties go to the answer seen first.

    es.majority_vote_accuracy

  • pass@kImplemented

    Probability that at least one of k samples drawn without replacement from the n generated for a problem passes its tests, estimated without bias from the c passing samples and averaged over problems.

    es.pass_at_k