Skip to content
EvalSuite
Documentation menu

LLM-as-a-judge

Fleiss' kappa

Implementedllm-judge.fleiss_kappa

Definition

Chance-corrected agreement of a fixed number of raters assigning items to nominal categories.

Formula

κ = (P̄ − P̄_e) / (1 − P̄_e)

Range: (−∞, 1]

Inputs and outputs

  • ratings: raters × items matrix (NaN = missing) or items × raters labels

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.fleiss_kappa(item_ratings)

References

  1. Fleiss JL. Measuring nominal scale agreement among many raters. Psychol Bull. 1971;76(5):378-382.

Implementation status