Skip to content
EvalSuite
Documentation menu

Safety and responsible AI

WEAT effect size

Implementedsafety.weat_effect_size

Definition

Word Embedding Association Test: how much more strongly target set X than target set Y associates with attribute set A than with B, as a standardised effect size (Cohen's d analogue), with a permutation p-value.

Formula

d = (mean_x s(x,A,B) − mean_y s(y,A,B)) / std_{w∈X∪Y} s(w,A,B), s = mean cos(w,A) − mean cos(w,B)

Range: [−2, 2] (0 = no association)

Inputs and outputs

  • X: see the signature of es.weat_effect_size
  • Y: see the signature of es.weat_effect_size
  • A: see the signature of es.weat_effect_size
  • B: see the signature of es.weat_effect_size

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • Embeddings (rows) of the two target sets and two attribute sets. The one-sided p-value uses random equal-size re-partitions of X ∪ Y (exact enumeration is used when it is smaller).

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.weat_effect_size(X, Y, A, B, random_state=0)  # embeddings of word sets

References

  1. Caliskan A, Bryson JJ, Narayanan A. Semantics derived automatically from language corpora contain human-like biases. Science. 2017;356(6334):183-186.

Implementation status