Multilingual
Cultural appropriateness
Implemented
multilingual.cultural_appropriatenessDefinition
Mean rating of how appropriate responses are for the target culture, rescaled to [0, 1] from the rating scale, overall and per culture / region (the gap between the best and worst region is reported).
Formula
mean (rating − lo) / (hi − lo)
Range: [0, 1]
Inputs and outputs
- ratings: see the signature of es.cultural_appropriateness
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.cultural_appropriateness([5, 3, 4], regions=["IN", "IN", "BR"])References
- Myung J, Lee N, Zhou Y, et al. BLEnD: a benchmark for LLMs on everyday knowledge in diverse cultures and languages. NeurIPS Datasets and Benchmarks. 2024.