Skip to content
EvalSuite
Documentation menu

Safety and responsible AI

Harmful response rate

Implementedsafety.harmful_response_rate

Definition

Share of responses judged harmful. Given which prompts request harmful content, it is the unsafe compliance rate: harmful responses among harmful prompts (HarmBench attack success).

Formula

harmful responses / responses (restricted to harmful prompts when given)

Range: [0, 1]

Inputs and outputs

  • harmful: see the signature of es.harmful_response_rate

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • ``harmful``: one verdict per response. ``harmful_prompt`` (optional): whether each prompt asked for harmful content; the rate is then over those prompts only. ``categories``: per-category rates.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.harmful_response_rate([True, False, False], harmful_prompt=[True, True, False])

References

  1. Mazeika M, Phan L, Yin X, et al. HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. ICML. 2024.

Implementation status