Long context and summarization
Cross-document consistency
Implemented
long-context.cross_document_consistencyDefinition
Agreement of answers to the same question across different (orderings or subsets of) documents or truncation levels: share of questions answered identically and mean pairwise agreement.
Formula
mean_q 1[all answers equal]
Range: [0, 1]
Inputs and outputs
- answers: see the signature of es.cross_document_consistency
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.cross_document_consistency([["A", "A", "B"]])References
- Liu NF, Lin K, Hewitt J, et al. Lost in the middle: how language models use long contexts. TACL. 2024;12:157-173.
- Bai Y, Lv X, Zhang J, et al. LongBench: a bilingual, multitask benchmark for long context understanding. ACL. 2024:3119-3137.