Multilingual
Cross-lingual factual consistency
Implemented
multilingual.cross_lingual_consistencyDefinition
Whether the model gives the same answer to the same question asked in different languages: share of questions answered identically in every language and mean pairwise agreement (after mapping answers to a common form with ``normalize``).
Formula
mean_q 1[all languages agree]
Range: [0, 1]
Inputs and outputs
- answers_by_language: see the signature of es.cross_lingual_consistency
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``answers_by_language``: per question, a mapping language -> answer (at least two languages).
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.cross_lingual_consistency([{"en": "Paris", "fr": "Paris"}])References
- Qi J, Fernández R, Bisazza A. Cross-lingual consistency of factual knowledge in multilingual language models. EMNLP. 2023:10650-10666.