Skip to content
EvalSuite
Documentation menu

Safety and responsible AI

Policy violation rate

Implementedsafety.policy_violation_rate

Definition

Share of outputs that violate at least one content policy, with the violation rate of each policy.

Formula

outputs with ≥1 violated policy / outputs

Range: [0, 1]

Inputs and outputs

  • violations: see the signature of es.policy_violation_rate

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • ``violations``: per output, the list of policies it violates (empty if none). ``policies``: the full policy list, so policies never violated appear with rate 0.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.policy_violation_rate([[], ["violence"]], policies=["violence", "pii"])

References

  1. Mazeika M, Phan L, Yin X, et al. HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. ICML. 2024.

Implementation status