Semantic similarity
MoverScore
Implemented
semantic.moverscoreDefinition
One minus the earth mover's distance between the IDF-weighted token embeddings of prediction and reference, with Euclidean transport cost between L2-normalised embeddings (MoverScore v2 style).
Formula
1 − EMD(w_ref, w_pred; ‖r̂_i − p̂_j‖₂)
Range: (−∞, 1]
Inputs and outputs
- reference_embeddings: embeddings from your encoder
- prediction_embeddings: embeddings with the same dimension
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.moverscore(ref_token_embeddings, pred_token_embeddings)References
- Zhao W, Peyrard M, Liu F, Gao Y, Meyer CM, Eger S. MoverScore: text generation evaluating with contextualized embeddings and earth mover distance. EMNLP. 2019:563-578.