Skip to content
EvalSuite
Documentation menu

LLM-as-a-judge

Pairwise win rate

Implementedllm-judge.win_rate

Definition

Share of pairwise comparisons in which a system's response is preferred to the baseline; ties count half (or are excluded). A Wilson interval is reported.

Formula

(wins + ½ ties) / comparisons

Range: [0, 1]

Inputs and outputs

  • outcomes: win / loss / tie against the baseline

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.win_rate(outcomes)

References

  1. Zheng L, Chiang WL, Sheng Y, et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks. 2023.
  2. Wilson EB. Probable inference, the law of succession, and statistical inference. JASA. 1927;22:209-212.

Implementation status