Long context and summarization
Long-context retrieval accuracy
Implemented
long-context.retrieval_accuracy_by_lengthDefinition
Accuracy of retrieving or answering from long contexts, overall and per context-length level (RULER reports the accuracy at each length; the effective length is the longest one above a threshold).
Formula
acc_L = mean[correct | length = L]
Range: [0, 1]
Inputs and outputs
- correct: see the signature of es.retrieval_accuracy_by_length
- context_lengths: see the signature of es.retrieval_accuracy_by_length
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.retrieval_accuracy_by_length(correct, context_lengths, threshold=0.85)References
- Hsieh CP, Sun S, Kriman S, et al. RULER: what's the real context size of your long-context language models? COLM. 2024.
- Bai Y, Lv X, Zhang J, et al. LongBench: a bilingual, multitask benchmark for long context understanding. ACL. 2024:3119-3137.