Skip to content
EvalSuite
Documentation menu

Text generation

BLEU

Implementedtext-generation.bleu

Definition

Corpus-level geometric mean of clipped n-gram precisions (n = 1..4) times a brevity penalty, computed exactly as sacreBLEU (13a tokenization, exponential smoothing).

Formula

BLEU = BP · exp(Σ_n (1/N) log p_n), BP = min(1, exp(1 − r/c))

Range: [0, 100]

Inputs and outputs

  • references: one reference string (or a list of references) per example
  • predictions: one model output per example

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.bleu(references, predictions)  # corpus BLEU, sacreBLEU conventions

References

  1. Papineni K, Roukos S, Ward T, Zhu WJ. BLEU: a method for automatic evaluation of machine translation. ACL. 2002:311-318.
  2. Post M. A call for clarity in reporting BLEU scores. WMT. 2018:186-191.
  3. Chen B, Cherry C. A systematic comparison of smoothing techniques for sentence-level BLEU. WMT. 2014.

Implementation status