Multilingual
Language-specific performance parity
Implemented
multilingual.language_parityDefinition
How evenly a model performs across languages: the worst-to-best ratio (the value; 1 = parity), the largest gap, the standard deviation, and per-language scores.
Formula
parity = min_l score_l / max_l score_l
Range: [0, 1]
Inputs and outputs
- scores_by_language: see the signature of es.language_parity
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``scores_by_language``: language -> per-example scores (or one aggregate score).
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.language_parity({"en": [1, 1, 0, 1], "sw": [1, 0, 0, 1]})References
- Hu J, Ruder S, Siddhant A, Neubig G, Firat O, Johnson M. XTREME: a massively multilingual multi-task benchmark for evaluating cross-lingual generalization. ICML. 2020:4411-4421.