LLM-as-a-judge
Judge–human agreement
Implemented
llm-judge.judge_agreementDefinition
Agreement between an automatic judge and human ratings: Cohen's kappa for categories, quadratic-weighted kappa for ordinal scores (with Spearman's ρ), Pearson / Spearman correlation for continuous scores; raw agreement is reported alongside.
Formula
kappa or correlation, by data type
Range: [-1, 1]
Inputs and outputs
- judge: judge labels or scores
- human: human labels or scores
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.judge_agreement(judge_scores, human_scores, kind="ordinal")References
- Zheng L, Chiang WL, Sheng Y, et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks. 2023.