Skip to content
EvalSuite
Documentation menu

Text generation

METEOR

Implementedtext-generation.meteor

Definition

Unigram alignment between prediction and reference by exact, stemmed and (optionally) synonym matches; harmonic mean weighted towards recall, penalised for fragmented alignments. Best reference per example, averaged over examples (NLTK's meteor_score).

Formula

(1 − γ (chunks/m)^β) · P·R / (α P + (1 − α) R)

Range: [0, 1]

Inputs and outputs

  • references: one reference string (or a list of references) per example
  • predictions: one model output per example

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.meteor(references, predictions)             # synonyms=... to add WordNet

References

  1. Banerjee S, Lavie A. METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. ACL Workshop on Evaluation Measures. 2005:65-72.
  2. Lavie A, Agarwal A. METEOR: an automatic metric for MT evaluation with high levels of correlation with human judgments. WMT. 2007:228-231.

Implementation status