Robustness and reliability
Prompt wording sensitivity
Implemented
robustness.prompt_sensitivityDefinition
Spread of task performance across semantically equivalent prompt templates: the range (best − worst accuracy), standard deviation and worst-case accuracy over templates (FormatSpread).
Formula
spread = max_t acc_t − min_t acc_t
Range: [0, 1]
Inputs and outputs
- correct_by_template: see the signature of es.prompt_sensitivity
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``correct_by_template``: mapping template name -> per-example correctness (or scores), the same examples under every template; or a 2-D array (templates × examples).
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.prompt_sensitivity({"t1": [1, 1, 0], "t2": [1, 0, 0]})References
- Sclar M, Choi Y, Tsvetkov Y, Suhr A. Quantifying language models' sensitivity to spurious features in prompt design. ICLR. 2024.
- Zhu K, Wang J, Zhou J, et al. PromptRobust: towards evaluating the robustness of large language models on adversarial prompts. arXiv:2306.04528. 2023.