Robustness and reliability
Long-context robustness / truncation sensitivity
Implemented
robustness.truncation_sensitivityDefinition
How performance changes with context length or truncation: accuracy per length bucket and the least-squares slope of accuracy against log2 length (negative = degrades as context grows).
Formula
slope of acc_b on log2(len_b)
Range: (−∞, ∞)
Inputs and outputs
- correct: see the signature of es.truncation_sensitivity
- context_lengths: see the signature of es.truncation_sensitivity
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.truncation_sensitivity([1, 1, 0, 0], [4000, 4000, 32000, 32000])References
- Hsieh CP, Sun S, Kriman S, et al. RULER: what's the real context size of your long-context language models? COLM. 2024.