Text generation
Sentence BLEU
Implemented
text-generation.sentence_bleuDefinition
BLEU computed for each example separately (exponential smoothing, effective order) and averaged; use for per-example scores, not for reporting corpus quality.
Formula
mean_i BLEU(prediction_i, references_i)
Range: [0, 100]
Inputs and outputs
- references: one reference string (or a list of references) per example
- predictions: one model output per example
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.sentence_bleu(references, predictions, average=None)References
- Papineni K, Roukos S, Ward T, Zhu WJ. BLEU: a method for automatic evaluation of machine translation. ACL. 2002:311-318.
- Chen B, Cherry C. A systematic comparison of smoothing techniques for sentence-level BLEU. WMT. 2014.