Skip to content
EvalSuite
Documentation menu

Long context and summarization

Long-context retrieval accuracy

Implementedlong-context.retrieval_accuracy_by_length

Definition

Accuracy of retrieving or answering from long contexts, overall and per context-length level (RULER reports the accuracy at each length; the effective length is the longest one above a threshold).

Formula

acc_L = mean[correct | length = L]

Range: [0, 1]

Inputs and outputs

  • correct: see the signature of es.retrieval_accuracy_by_length
  • context_lengths: see the signature of es.retrieval_accuracy_by_length

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.retrieval_accuracy_by_length(correct, context_lengths, threshold=0.85)

References

  1. Hsieh CP, Sun S, Kriman S, et al. RULER: what's the real context size of your long-context language models? COLM. 2024.
  2. Bai Y, Lv X, Zhang J, et al. LongBench: a bilingual, multitask benchmark for long context understanding. ACL. 2024:3119-3137.

Implementation status