Skip to content
EvalSuite
Documentation menu

Long context and summarization

Cross-document consistency

Implementedlong-context.cross_document_consistency

Definition

Agreement of answers to the same question across different (orderings or subsets of) documents or truncation levels: share of questions answered identically and mean pairwise agreement.

Formula

mean_q 1[all answers equal]

Range: [0, 1]

Inputs and outputs

  • answers: see the signature of es.cross_document_consistency

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.cross_document_consistency([["A", "A", "B"]])

References

  1. Liu NF, Lin K, Hewitt J, et al. Lost in the middle: how language models use long contexts. TACL. 2024;12:157-173.
  2. Bai Y, Lv X, Zhang J, et al. LongBench: a bilingual, multitask benchmark for long context understanding. ACL. 2024:3119-3137.

Implementation status