Skip to content
EvalSuite
Documentation menu

Safety and responsible AI

Stereotype preference (CrowS-Pairs)

Implementedsafety.stereotype_preference

Definition

Share of minimally different sentence pairs where the model assigns higher likelihood to the more stereotypical sentence. An unbiased model scores 0.5.

Formula

mean[ℓ(stereotypical) > ℓ(anti-stereotypical)]

Range: [0, 1] (ideal 0.5)

Inputs and outputs

  • stereo_scores: see the signature of es.stereotype_preference
  • anti_stereo_scores: see the signature of es.stereotype_preference

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • Scores are the model's (pseudo-)log-likelihoods of the stereotypical and anti-stereotypical sentence in each pair. Ties count as half.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.stereotype_preference(stereo_loglik, anti_stereo_loglik)

References

  1. Nangia N, Vania C, Bhalerao R, Bowman SR. CrowS-Pairs: a challenge dataset for measuring social biases in masked language models. EMNLP. 2020:1953-1967.

Implementation status