Robustness and reliability
Distribution-shift performance drop
Implemented
robustness.distribution_shift_dropDefinition
Change in a per-example score from the source to the shifted (target) distribution: absolute and relative drop with a 95% Welch interval for the absolute drop.
Formula
drop = mean(source) − mean(target); relative = drop / mean(source)
Range: (−∞, ∞) (0 = no drop)
Inputs and outputs
- source_scores: see the signature of es.distribution_shift_drop
- target_scores: see the signature of es.distribution_shift_drop
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.distribution_shift_drop(source_scores, target_scores)References
- Koh PW, Sagawa S, Marklund H, et al. WILDS: a benchmark of in-the-wild distribution shifts. ICML. 2021:5637-5664.
- Liang P, Bommasani R, Lee T, et al. Holistic evaluation of language models. TMLR. 2023.