Skip to content
EvalSuite
Documentation menu

Text generation

TER

Implementedtext-generation.ter

Definition

Translation edit rate: minimum number of insertions, deletions, substitutions and block shifts to turn the prediction into the closest reference, divided by the average reference length (Tercom, as in sacreBLEU). Lower is better.

Formula

TER = Σ edits / Σ average reference length × 100

Range: [0, ∞)

Inputs and outputs

  • references: one reference string (or a list of references) per example
  • predictions: one model output per example

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.ter(references, predictions)

References

  1. Snover M, Dorr B, Schwartz R, Micciulla L, Makhoul J. A study of translation edit rate with targeted human annotation. AMTA. 2006:223-231.
  2. Post M. A call for clarity in reporting BLEU scores. WMT. 2018:186-191.

Implementation status