LLM-as-a-judge
Pairwise win rate
Implemented
llm-judge.win_rateDefinition
Share of pairwise comparisons in which a system's response is preferred to the baseline; ties count half (or are excluded). A Wilson interval is reported.
Formula
(wins + ½ ties) / comparisons
Range: [0, 1]
Inputs and outputs
- outcomes: win / loss / tie against the baseline
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.win_rate(outcomes)References
- Zheng L, Chiang WL, Sheng Y, et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks. 2023.
- Wilson EB. Probable inference, the law of succession, and statistical inference. JASA. 1927;22:209-212.