Skip to content
EvalSuite
Documentation menu

Robustness and reliability

Paraphrase consistency

Implementedrobustness.paraphrase_consistency

Definition

Whether the model gives the same answer to paraphrases of the same question: the share of items where all paraphrases agree, and the mean pairwise agreement (ParaRel consistency).

Formula

mean_items 1[all answers equal]; pairwise = mean over pairs 1[a_i = a_j]

Range: [0, 1]

Inputs and outputs

  • answers: see the signature of es.paraphrase_consistency

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • ``answers``: per item, the answers to each paraphrase (at least two). Strings are compared after lower-casing and whitespace normalisation unless ``normalize`` is given.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.paraphrase_consistency([["Paris", "paris"], ["4", "5"]])

References

  1. Elazar Y, Kassner N, Ravfogel S, et al. Measuring and improving consistency in pretrained language models. TACL. 2021;9:1012-1031.
  2. Ribeiro MT, Wu T, Guestrin C, Singh S. Beyond accuracy: behavioral testing of NLP models with CheckList. ACL. 2020:4902-4912.

Implementation status