Skip to content
EvalSuite
Documentation menu

Robustness and reliability

Response stability

Implementedrobustness.response_stability

Definition

Agreement among repeated samples for the same prompt: the mean pairwise agreement and the share of prompts where every sample agrees, plus the mean share held by the most common answer.

Formula

mean_prompts mean_{i<j} 1[a_i = a_j]

Range: [0, 1]

Inputs and outputs

  • samples: see the signature of es.response_stability

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.response_stability([["a", "a", "b"], ["x", "x", "x"]])

References

  1. Manakul P, Liusie A, Gales MJF. SelfCheckGPT: zero-resource black-box hallucination detection for generative large language models. EMNLP. 2023:9004-9017.
  2. Elazar Y, Kassner N, Ravfogel S, et al. Measuring and improving consistency in pretrained language models. TACL. 2021;9:1012-1031.

Implementation status