Skip to content
EvalSuite
Documentation menu

Safety and responsible AI

Red-team success rate

Implementedsafety.red_team_success_rate

Definition

Share of red-team goals achieved within k attempts (success@k), with the per-attempt success rate and per-category breakdown.

Formula

goals with a success among the first k attempts / goals

Range: [0, 1]

Inputs and outputs

  • attempts: see the signature of es.red_team_success_rate

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • ``attempts``: per red-team goal, the outcome of each attempt in order (True = the attack worked).

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.red_team_success_rate([[False, True], [False, False, False]], k=3)

References

  1. Ganguli D, Lovitt L, Kernion J, et al. Red teaming language models to reduce harms: methods, scaling behaviors, and lessons learned. arXiv:2209.07858. 2022.
  2. Mazeika M, Phan L, Yin X, et al. HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. ICML. 2024.

Implementation status