Factuality, question answering and reasoning
Available since v0.4.0Factuality metrics score verdicts produced by a verification step: annotators, an NLI model, a retrieval check or an LLM judge decide whether each claim of an answer is supported. EvalSuite aggregates them consistently (per answer or pooled), so verification pipelines can be compared with intervals. Report how claims were extracted and judged next to the numbers.
import evalsuite as es
claim_verdicts = [["supported", "supported", "contradicted"], [True, False]]
es.faithfulness(claim_verdicts) # FActScore / RAGAS faithfulness
es.hallucination_rate(claim_verdicts) # unsupported-claim rate
es.groundedness([[0.9, 0.7], [0.2]]) # graded support scores
citations = [[{"supported": True, "citations": [True, False]}, {"supported": False, "citations": []}]]
es.citation_recall(citations) # ALCE
es.citation_precision(citations)
es.abstention_accuracy([True, False, False], [True, False, True], correct=[False, True, False])Question answering
references = [["Paris", "Paris, France"], ["1990"]]
predictions = ["paris", "in 1991"]
es.exact_match(references, predictions) # SQuAD normalisation: lowercase, no punctuation or articles
es.token_f1(references, predictions)
es.answer_correctness(tp=[3, 2], fp=[1, 0], fn=[0, 1], similarity=[0.9, 0.8]) # RAGASReasoning benchmarks
es.pass_at_k(n_samples=[10, 10], n_correct=[3, 0], k=5) # unbiased estimator (Chen et al. 2021)
es.majority_vote_accuracy(["42", "7"], [["42", "41", "42"], ["7", "8", "8"]])
es.benchmark_accuracy(["... #### 18", "#### 42"], ["so the answer is 18", "41"], style="gsm8k")
es.extract_answer(r"thus \boxed{\frac{1}{2}}", style="boxed") # MATH conventionMetrics
| Metric | Description | Status | API |
|---|---|---|---|
| Abstention accuracy | Whether a system abstains exactly when it should (the evidence is insufficient or it would otherwise be wrong): accuracy of the abstain / answer decision; abstention precision, recall and the accuracy on answered questions are reported alongside. | Implemented | es.abstention_accuracy |
| Answer correctness | Agreement of an answer with the reference at the level of statements: F1 over true-positive (in both), false-positive (only in the answer) and false-negative (only in the reference) statements, optionally blended with a semantic similarity score (RAGAS answer correctness). | Implemented | es.answer_correctness |
| Answer relevance | How directly an answer addresses the question: mean cosine similarity between the question's embedding and embeddings of questions regenerated from the answer (RAGAS answer relevancy). | Implemented | es.answer_relevance |
| Citation precision | Share of citations that are relevant to the statement they are attached to (ALCE: a citation counts if it supports, or is needed to support, the statement). | Implemented | es.citation_precision |
| Citation recall | Share of statements whose cited passages, taken together, support them (statements without citations count as unsupported); ALCE citation recall. | Implemented | es.citation_recall |
| Claim verification accuracy | Accuracy of supported / contradicted / not-enough-info verdicts against gold labels (FEVER label accuracy); macro F1 is reported alongside because the classes are often imbalanced. | Implemented | es.claim_verification_accuracy |
| Exact match | Share of predictions identical to a reference after normalization (SQuAD: lowercase, no punctuation or articles, single spaces); with several references, any match counts. | Implemented | es.exact_match |
| Faithfulness | Share of an answer's claims that the evidence supports (FActScore factual precision; RAGAS faithfulness when the evidence is the retrieved context). | Implemented | es.faithfulness |
| Groundedness | Average support score of an answer's content given the supplied context, from graded scores in [0, 1] (e.g. NLI entailment probabilities or judge ratings) rather than binary verdicts. | Implemented | es.groundedness |
| Hallucination rate (unsupported-claim rate) | Share of an answer's claims that the evidence does not support, i.e. contradicted or unverifiable claims, under the stated verification protocol. | Implemented | es.hallucination_rate |
| Knowledge consistency | Share of facts extracted from the outputs (e.g. subject–relation–object triples) that are present in a specified trusted knowledge source. | Implemented | es.knowledge_consistency |
| Token F1 | Harmonic mean of token precision and recall between the normalized prediction and reference (bag of words, SQuAD); best reference per example, averaged over examples. | Implemented | es.token_f1 |
- Abstention accuracyImplemented
Whether a system abstains exactly when it should (the evidence is insufficient or it would otherwise be wrong): accuracy of the abstain / answer decision; abstention precision, recall and the accuracy on answered questions are reported alongside.
es.abstention_accuracy
- Answer correctnessImplemented
Agreement of an answer with the reference at the level of statements: F1 over true-positive (in both), false-positive (only in the answer) and false-negative (only in the reference) statements, optionally blended with a semantic similarity score (RAGAS answer correctness).
es.answer_correctness
- Answer relevanceImplemented
How directly an answer addresses the question: mean cosine similarity between the question's embedding and embeddings of questions regenerated from the answer (RAGAS answer relevancy).
es.answer_relevance
- Citation precisionImplemented
Share of citations that are relevant to the statement they are attached to (ALCE: a citation counts if it supports, or is needed to support, the statement).
es.citation_precision
- Citation recallImplemented
Share of statements whose cited passages, taken together, support them (statements without citations count as unsupported); ALCE citation recall.
es.citation_recall
- Claim verification accuracyImplemented
Accuracy of supported / contradicted / not-enough-info verdicts against gold labels (FEVER label accuracy); macro F1 is reported alongside because the classes are often imbalanced.
es.claim_verification_accuracy
- Exact matchImplemented
Share of predictions identical to a reference after normalization (SQuAD: lowercase, no punctuation or articles, single spaces); with several references, any match counts.
es.exact_match
- FaithfulnessImplemented
Share of an answer's claims that the evidence supports (FActScore factual precision; RAGAS faithfulness when the evidence is the retrieved context).
es.faithfulness
- GroundednessImplemented
Average support score of an answer's content given the supplied context, from graded scores in [0, 1] (e.g. NLI entailment probabilities or judge ratings) rather than binary verdicts.
es.groundedness
- Hallucination rate (unsupported-claim rate)Implemented
Share of an answer's claims that the evidence does not support, i.e. contradicted or unverifiable claims, under the stated verification protocol.
es.hallucination_rate
- Knowledge consistencyImplemented
Share of facts extracted from the outputs (e.g. subject–relation–object triples) that are present in a specified trusted knowledge source.
es.knowledge_consistency
- Token F1Implemented
Harmonic mean of token precision and recall between the normalized prediction and reference (bag of words, SQuAD); best reference per example, averaged over examples.
es.token_f1
| Metric | Description | Status | API |
|---|---|---|---|
| Benchmark accuracy | Share of benchmark questions answered correctly after extracting the final answer with the benchmark's convention (GSM8K '####' or last number, MATH \boxed{}, multiple-choice letter) and comparing with the gold answer. | Implemented | es.benchmark_accuracy |
| Majority-vote accuracy | Accuracy of the most frequent answer among several samples per question (self-consistency); ties go to the answer seen first. | Implemented | es.majority_vote_accuracy |
| pass@k | Probability that at least one of k samples drawn without replacement from the n generated for a problem passes its tests, estimated without bias from the c passing samples and averaged over problems. | Implemented | es.pass_at_k |
- Benchmark accuracyImplemented
Share of benchmark questions answered correctly after extracting the final answer with the benchmark's convention (GSM8K '####' or last number, MATH \boxed{}, multiple-choice letter) and comparing with the gold answer.
es.benchmark_accuracy
- Majority-vote accuracyImplemented
Accuracy of the most frequent answer among several samples per question (self-consistency); ties go to the answer seen first.
es.majority_vote_accuracy
- pass@kImplemented
Probability that at least one of k samples drawn without replacement from the n generated for a problem passes its tests, estimated without bias from the c passing samples and averaged over problems.
es.pass_at_k