Text generation
ROUGE-Lsum
Implemented
text-generation.rouge_lsumDefinition
Summary-level union LCS over sentences (one sentence per line) between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example.
Formula
union-LCS over newline-separated sentences
Range: [0, 1]
Inputs and outputs
- references: one reference string (or a list of references) per example
- predictions: one model output per example
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.rouge_lsum(references, predictions) # one sentence per lineReferences
- Lin CY. ROUGE: a package for automatic evaluation of summaries. Text Summarization Branches Out. 2004.