Long context and summarization
Context utilization / retention
Implemented
long-context.context_utilizationDefinition
Share of the relevant facts provided in the context that the output actually uses (or recalls later in a conversation), pooled over examples.
Formula
Σ |used ∩ provided| / Σ |provided|
Range: [0, 1]
Inputs and outputs
- used_facts: see the signature of es.context_utilization
- provided_facts: see the signature of es.context_utilization
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.context_utilization([["f1"]], [["f1", "f2"]])References
- Bai Y, Lv X, Zhang J, et al. LongBench: a bilingual, multitask benchmark for long context understanding. ACL. 2024:3119-3137.