Long context and summarization
Lost-in-the-middle sensitivity
Implemented
long-context.lost_in_the_middleDefinition
How much worse the model does when the relevant passage is in the middle of the context than at its edges: mean accuracy at the start and end positions minus accuracy in the middle (Liu et al. U-curve).
Formula
(acc_start + acc_end)/2 − acc_middle
Range: [−1, 1] (0 = position-invariant)
Inputs and outputs
- correct: see the signature of es.lost_in_the_middle
- positions: see the signature of es.lost_in_the_middle
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.lost_in_the_middle(correct, positions)References
- Liu NF, Lin K, Hewitt J, et al. Lost in the middle: how language models use long contexts. TACL. 2024;12:157-173.