Skip to content
EvalSuite
Documentation menu

Long context and summarization

Needle-in-a-haystack accuracy

Implementedlong-context.needle_in_haystack

Definition

Retrieval of a planted fact (the needle) from contexts of varying length and insertion depth: overall accuracy and the length × depth accuracy grid (the usual heat map).

Formula

acc = mean[correct]; grid[L, d] = mean[correct | L, d]

Range: [0, 1]

Inputs and outputs

  • correct: see the signature of es.needle_in_haystack
  • context_lengths: see the signature of es.needle_in_haystack
  • depths: see the signature of es.needle_in_haystack

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • ``depths``: where the needle was placed, as a fraction of the context in [0, 1].

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.needle_in_haystack(correct, context_lengths, depths)

References

  1. Kamradt G. Needle in a haystack: pressure testing LLMs. GitHub repository gkamradt/LLMTest_NeedleInAHaystack. 2023.
  2. Hsieh CP, Sun S, Kriman S, et al. RULER: what's the real context size of your long-context language models? COLM. 2024.

Implementation status