Robustness and reliability
Paraphrase consistency
Implemented
robustness.paraphrase_consistencyDefinition
Whether the model gives the same answer to paraphrases of the same question: the share of items where all paraphrases agree, and the mean pairwise agreement (ParaRel consistency).
Formula
mean_items 1[all answers equal]; pairwise = mean over pairs 1[a_i = a_j]
Range: [0, 1]
Inputs and outputs
- answers: see the signature of es.paraphrase_consistency
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``answers``: per item, the answers to each paraphrase (at least two). Strings are compared after lower-casing and whitespace normalisation unless ``normalize`` is given.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.paraphrase_consistency([["Paris", "paris"], ["4", "5"]])References
- Elazar Y, Kassner N, Ravfogel S, et al. Measuring and improving consistency in pretrained language models. TACL. 2021;9:1012-1031.
- Ribeiro MT, Wu T, Guestrin C, Singh S. Beyond accuracy: behavioral testing of NLP models with CheckList. ACL. 2020:4902-4912.