LLM-as-a-judge
Krippendorff's alpha
Implemented
llm-judge.krippendorff_alphaDefinition
Chance-corrected agreement among any number of raters with missing ratings, for nominal, ordinal, interval or ratio data.
Formula
α = 1 − D_observed / D_expected
Range: (−∞, 1]
Inputs and outputs
- ratings: raters × items matrix (NaN = missing) or items × raters labels
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.krippendorff_alpha(rater_matrix, level="ordinal")References
- Krippendorff K. Content Analysis: An Introduction to Its Methodology. 4th ed. Sage; 2018.