Safety and responsible AI
WEAT effect size
Implemented
safety.weat_effect_sizeDefinition
Word Embedding Association Test: how much more strongly target set X than target set Y associates with attribute set A than with B, as a standardised effect size (Cohen's d analogue), with a permutation p-value.
Formula
d = (mean_x s(x,A,B) − mean_y s(y,A,B)) / std_{w∈X∪Y} s(w,A,B), s = mean cos(w,A) − mean cos(w,B)
Range: [−2, 2] (0 = no association)
Inputs and outputs
- X: see the signature of es.weat_effect_size
- Y: see the signature of es.weat_effect_size
- A: see the signature of es.weat_effect_size
- B: see the signature of es.weat_effect_size
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- Embeddings (rows) of the two target sets and two attribute sets. The one-sided p-value uses random equal-size re-partitions of X ∪ Y (exact enumeration is used when it is smaller).
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.weat_effect_size(X, Y, A, B, random_state=0) # embeddings of word setsReferences
- Caliskan A, Bryson JJ, Narayanan A. Semantics derived automatically from language corpora contain human-like biases. Science. 2017;356(6334):183-186.