Skip to content
EvalSuite
Documentation menu

Robustness and reliability

Contradiction rate

Implementedrobustness.contradiction_rate

Definition

Share of compared response pairs (repeated samples, or answers to related questions) that an NLI model or judge labels as contradicting each other (SelfCheckGPT-NLI style).

Formula

#contradiction / #pairs

Range: [0, 1]

Inputs and outputs

  • pair_labels: see the signature of es.contradiction_rate

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • ``pair_labels``: per compared pair, ``"contradiction"``, ``"entailment"`` or ``"neutral"`` (or a list of such labels per prompt; all pairs are pooled).

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.contradiction_rate(["entailment", "contradiction", "neutral"])

References

  1. Manakul P, Liusie A, Gales MJF. SelfCheckGPT: zero-resource black-box hallucination detection for generative large language models. EMNLP. 2023:9004-9017.

Implementation status