Skip to content
EvalSuite
Documentation menu

Robustness and reliability

Distribution-shift performance drop

Implementedrobustness.distribution_shift_drop

Definition

Change in a per-example score from the source to the shifted (target) distribution: absolute and relative drop with a 95% Welch interval for the absolute drop.

Formula

drop = mean(source) − mean(target); relative = drop / mean(source)

Range: (−∞, ∞) (0 = no drop)

Inputs and outputs

  • source_scores: see the signature of es.distribution_shift_drop
  • target_scores: see the signature of es.distribution_shift_drop

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.distribution_shift_drop(source_scores, target_scores)

References

  1. Koh PW, Sagawa S, Marklund H, et al. WILDS: a benchmark of in-the-wild distribution shifts. ICML. 2021:5637-5664.
  2. Liang P, Bommasani R, Lee T, et al. Holistic evaluation of language models. TMLR. 2023.

Implementation status