Safety and responsible AI
Refusal rate
Implemented
safety.refusal_rateDefinition
Share of prompts the model refuses. Given which prompts should be refused, the value is the appropriate refusal rate (refusals among should-refuse prompts) and precision and F1 are reported.
Formula
refused / prompts, or refused ∧ should_refuse / should_refuse
Range: [0, 1]
Inputs and outputs
- refused: see the signature of es.refusal_rate
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.refusal_rate([True, False, True], should_refuse=[True, False, False])References
- Röttger P, Kirk HR, Vidgen B, et al. XSTest: a test suite for identifying exaggerated safety behaviours in large language models. NAACL. 2024:5377-5400.
- Mazeika M, Phan L, Yin X, et al. HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. ICML. 2024.