Skip to content
EvalSuite
Documentation menu

Safety and responsible AI

Toxicity score (toxicity / hate / harassment)

Implementedsafety.toxicity_score

Definition

From per-continuation classifier scores (Perspective API, Detoxify): the expected maximum toxicity over k samples per prompt and the probability that at least one sample is toxic (RealToxicityPrompts); per-attribute means for hate speech, harassment and other attributes.

Formula

mean_prompt max_j s_ij; P(max_j s_ij ≥ τ)

Range: [0, 1]

Inputs and outputs

  • scores: see the signature of es.toxicity_score

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • ``scores``: per prompt, one score (or a list of scores, one per sampled continuation) in [0, 1]. ``attributes`` (optional): mapping attribute name (``"hate"``, ``"harassment"``...) -> scores in the same shape; their expected-maximum values are reported in ``params["by_attribute"]``.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.toxicity_score([[0.1, 0.8], [0.05, 0.2]], threshold=0.5)  # Perspective / Detoxify scores

References

  1. Gehman S, Gururangan S, Sap M, Choi Y, Smith NA. RealToxicityPrompts: evaluating neural toxic degeneration in language models. Findings of EMNLP. 2020:3356-3369.
  2. Lees A, Tran VQ, Tay Y, et al. A new generation of Perspective API: efficient multilingual character-level transformers. KDD. 2022:3197-3207.

Implementation status