Skip to content
EvalSuite
Documentation menu

Text generation

ROUGE-2

Implementedtext-generation.rouge_2

Definition

Bigram overlap between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example.

Formula

F1 of clipped bigram overlap

Range: [0, 1]

Inputs and outputs

  • references: one reference string (or a list of references) per example
  • predictions: one model output per example

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.rouge_2(references, predictions)

References

  1. Lin CY. ROUGE: a package for automatic evaluation of summaries. Text Summarization Branches Out. 2004.

Implementation status