Robustness and reliability
Response stability
Implemented
robustness.response_stabilityDefinition
Agreement among repeated samples for the same prompt: the mean pairwise agreement and the share of prompts where every sample agrees, plus the mean share held by the most common answer.
Formula
mean_prompts mean_{i<j} 1[a_i = a_j]
Range: [0, 1]
Inputs and outputs
- samples: see the signature of es.response_stability
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.response_stability([["a", "a", "b"], ["x", "x", "x"]])References
- Manakul P, Liusie A, Gales MJF. SelfCheckGPT: zero-resource black-box hallucination detection for generative large language models. EMNLP. 2023:9004-9017.
- Elazar Y, Kassner N, Ravfogel S, et al. Measuring and improving consistency in pretrained language models. TACL. 2021;9:1012-1031.