Long context and summarization
Summary factual consistency / coverage / completeness
Implemented
long-context.summary_coverageDefinition
Coverage: share of the reference key points (pyramid content units) the summary covers, weighted by importance when weights are given; factual consistency (share of summary claims supported by the source) is reported when claim verdicts are passed.
Formula
coverage = Σ w·covered / Σ w; consistency = supported claims / claims
Range: [0, 1]
Inputs and outputs
- key_points_covered: see the signature of es.summary_coverage
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``key_points_covered``: per summary, one boolean per reference key point. ``weights``: per summary, the importance of each key point (pyramid weights). ``claims_supported``: per summary, one boolean per claim.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.summary_coverage([[True, False, True]], weights=[[3, 2, 1]])References
- Nenkova A, Passonneau R. Evaluating content selection in summarization: the pyramid method. HLT-NAACL. 2004:145-152.
- Min S, Krishna K, Lyu X, et al. FActScore: fine-grained atomic evaluation of factual precision in long form text generation. EMNLP. 2023:12076-12100.