Safety and responsible AI
Red-team success rate
Implemented
safety.red_team_success_rateDefinition
Share of red-team goals achieved within k attempts (success@k), with the per-attempt success rate and per-category breakdown.
Formula
goals with a success among the first k attempts / goals
Range: [0, 1]
Inputs and outputs
- attempts: see the signature of es.red_team_success_rate
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``attempts``: per red-team goal, the outcome of each attempt in order (True = the attack worked).
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.red_team_success_rate([[False, True], [False, False, False]], k=3)References
- Ganguli D, Lovitt L, Kernion J, et al. Red teaming language models to reduce harms: methods, scaling behaviors, and lessons learned. arXiv:2209.07858. 2022.
- Mazeika M, Phan L, Yin X, et al. HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. ICML. 2024.