Safety and responsible AI
Policy violation rate
Implemented
safety.policy_violation_rateDefinition
Share of outputs that violate at least one content policy, with the violation rate of each policy.
Formula
outputs with ≥1 violated policy / outputs
Range: [0, 1]
Inputs and outputs
- violations: see the signature of es.policy_violation_rate
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``violations``: per output, the list of policies it violates (empty if none). ``policies``: the full policy list, so policies never violated appear with rate 0.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.policy_violation_rate([[], ["violence"]], policies=["violence", "pii"])References
- Mazeika M, Phan L, Yin X, et al. HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. ICML. 2024.