Multilingual
Multilingual semantic similarity (bitext mining)
Implemented
multilingual.bitext_mining_accuracyDefinition
For aligned sentence pairs in two languages, the share whose nearest neighbour by cosine (or margin) similarity is the true translation, averaged over both directions (Tatoeba / BUCC style), with the mean cosine of the true pairs.
Formula
acc = ½ (mean_i 1[argmax_j cos(s_i, t_j) = i] + mean_j 1[argmax_i cos(s_i, t_j) = j])
Range: [0, 1]
Inputs and outputs
- source_embeddings: see the signature of es.bitext_mining_accuracy
- target_embeddings: see the signature of es.bitext_mining_accuracy
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``scoring="margin"`` uses the ratio margin of Artetxe & Schwenk with ``k`` neighbours, which removes hub effects.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.bitext_mining_accuracy(src_emb, tgt_emb, scoring="margin")References
- Artetxe M, Schwenk H. Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond. TACL. 2019;7:597-610.
- Hu J, Ruder S, Siddhant A, Neubig G, Firat O, Johnson M. XTREME: a massively multilingual multi-task benchmark for evaluating cross-lingual generalization. ICML. 2020:4411-4421.