Skip to content
EvalSuite
Documentation menu

Robustness and reliability

Prompt wording sensitivity

Implementedrobustness.prompt_sensitivity

Definition

Spread of task performance across semantically equivalent prompt templates: the range (best − worst accuracy), standard deviation and worst-case accuracy over templates (FormatSpread).

Formula

spread = max_t acc_t − min_t acc_t

Range: [0, 1]

Inputs and outputs

  • correct_by_template: see the signature of es.prompt_sensitivity

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • ``correct_by_template``: mapping template name -> per-example correctness (or scores), the same examples under every template; or a 2-D array (templates × examples).

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.prompt_sensitivity({"t1": [1, 1, 0], "t2": [1, 0, 0]})

References

  1. Sclar M, Choi Y, Tsvetkov Y, Suhr A. Quantifying language models' sensitivity to spurious features in prompt design. ICLR. 2024.
  2. Zhu K, Wang J, Zhou J, et al. PromptRobust: towards evaluating the robustness of large language models on adversarial prompts. arXiv:2306.04528. 2023.

Implementation status