Text generation
METEOR
Implemented
text-generation.meteorDefinition
Unigram alignment between prediction and reference by exact, stemmed and (optionally) synonym matches; harmonic mean weighted towards recall, penalised for fragmented alignments. Best reference per example, averaged over examples (NLTK's meteor_score).
Formula
(1 − γ (chunks/m)^β) · P·R / (α P + (1 − α) R)
Range: [0, 1]
Inputs and outputs
- references: one reference string (or a list of references) per example
- predictions: one model output per example
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.meteor(references, predictions) # synonyms=... to add WordNetReferences
- Banerjee S, Lavie A. METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. ACL Workshop on Evaluation Measures. 2005:65-72.
- Lavie A, Agarwal A. METEOR: an automatic metric for MT evaluation with high levels of correlation with human judgments. WMT. 2007:228-231.