Skip to content
EvalSuite
Documentation menu

Text generation

Self-BLEU

Implementedtext-generation.self_bleu

Definition

Average sentence BLEU of each prediction against all other predictions as references; higher means the outputs are more alike (less diverse).

Formula

mean_i BLEU(p_i, {p_j : j ≠ i})

Range: [0, 100]

Inputs and outputs

  • predictions: one model output per example

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.self_bleu(predictions)

References

  1. Zhu Y, et al. Texygen: a benchmarking platform for text generation models. SIGIR. 2018:1097-1100.

Implementation status