Long context and summarization
Needle-in-a-haystack accuracy
Implemented
long-context.needle_in_haystackDefinition
Retrieval of a planted fact (the needle) from contexts of varying length and insertion depth: overall accuracy and the length × depth accuracy grid (the usual heat map).
Formula
acc = mean[correct]; grid[L, d] = mean[correct | L, d]
Range: [0, 1]
Inputs and outputs
- correct: see the signature of es.needle_in_haystack
- context_lengths: see the signature of es.needle_in_haystack
- depths: see the signature of es.needle_in_haystack
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``depths``: where the needle was placed, as a fraction of the context in [0, 1].
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.needle_in_haystack(correct, context_lengths, depths)References
- Kamradt G. Needle in a haystack: pressure testing LLMs. GitHub repository gkamradt/LLMTest_NeedleInAHaystack. 2023.
- Hsieh CP, Sun S, Kriman S, et al. RULER: what's the real context size of your long-context language models? COLM. 2024.