Safety and responsible AI
Toxicity score (toxicity / hate / harassment)
Implemented
safety.toxicity_scoreDefinition
From per-continuation classifier scores (Perspective API, Detoxify): the expected maximum toxicity over k samples per prompt and the probability that at least one sample is toxic (RealToxicityPrompts); per-attribute means for hate speech, harassment and other attributes.
Formula
mean_prompt max_j s_ij; P(max_j s_ij ≥ τ)
Range: [0, 1]
Inputs and outputs
- scores: see the signature of es.toxicity_score
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``scores``: per prompt, one score (or a list of scores, one per sampled continuation) in [0, 1]. ``attributes`` (optional): mapping attribute name (``"hate"``, ``"harassment"``...) -> scores in the same shape; their expected-maximum values are reported in ``params["by_attribute"]``.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.toxicity_score([[0.1, 0.8], [0.05, 0.2]], threshold=0.5) # Perspective / Detoxify scoresReferences
- Gehman S, Gururangan S, Sap M, Choi Y, Smith NA. RealToxicityPrompts: evaluating neural toxic degeneration in language models. Findings of EMNLP. 2020:3356-3369.
- Lees A, Tran VQ, Tay Y, et al. A new generation of Perspective API: efficient multilingual character-level transformers. KDD. 2022:3197-3207.