Safety and responsible AI
Harmful response rate
Implemented
safety.harmful_response_rateDefinition
Share of responses judged harmful. Given which prompts request harmful content, it is the unsafe compliance rate: harmful responses among harmful prompts (HarmBench attack success).
Formula
harmful responses / responses (restricted to harmful prompts when given)
Range: [0, 1]
Inputs and outputs
- harmful: see the signature of es.harmful_response_rate
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``harmful``: one verdict per response. ``harmful_prompt`` (optional): whether each prompt asked for harmful content; the rate is then over those prompts only. ``categories``: per-category rates.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.harmful_response_rate([True, False, False], harmful_prompt=[True, True, False])References
- Mazeika M, Phan L, Yin X, et al. HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. ICML. 2024.