LLM-as-a-judge
Fleiss' kappa
Implemented
llm-judge.fleiss_kappaDefinition
Chance-corrected agreement of a fixed number of raters assigning items to nominal categories.
Formula
κ = (P̄ − P̄_e) / (1 − P̄_e)
Range: (−∞, 1]
Inputs and outputs
- ratings: raters × items matrix (NaN = missing) or items × raters labels
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.fleiss_kappa(item_ratings)References
- Fleiss JL. Measuring nominal scale agreement among many raters. Psychol Bull. 1971;76(5):378-382.