Long context and summarization
Position-dependent retrieval accuracy
Implemented
long-context.position_accuracyDefinition
Accuracy as a function of where the relevant information sits in the context (relative position binned into equal-width bins), with the spread between the best and worst bin.
Formula
acc_b = mean[correct | position ∈ bin b]; spread = max_b − min_b
Range: [0, 1]
Inputs and outputs
- correct: see the signature of es.position_accuracy
- positions: see the signature of es.position_accuracy
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``positions``: relative position of the relevant passage in [0, 1] (0 = start).
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.position_accuracy(correct, positions, n_bins=5)References
- Liu NF, Lin K, Hewitt J, et al. Lost in the middle: how language models use long contexts. TACL. 2024;12:157-173.