Skip to content
EvalSuite
Documentation menu

LLM-as-a-judge

Judge–human agreement

Implementedllm-judge.judge_agreement

Definition

Agreement between an automatic judge and human ratings: Cohen's kappa for categories, quadratic-weighted kappa for ordinal scores (with Spearman's ρ), Pearson / Spearman correlation for continuous scores; raw agreement is reported alongside.

Formula

kappa or correlation, by data type

Range: [-1, 1]

Inputs and outputs

  • judge: judge labels or scores
  • human: human labels or scores

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.judge_agreement(judge_scores, human_scores, kind="ordinal")

References

  1. Zheng L, Chiang WL, Sheng Y, et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks. 2023.

Implementation status