LLM-as-a-judge
Bradley–Terry scores
Implemented
llm-judge.bradley_terryDefinition
Maximum-likelihood strengths of a Bradley–Terry model fitted to pairwise preferences (ties as half a win for each side), by Hunter's MM algorithm; reported as log-strengths centred at 0 or on the Elo-like Chatbot Arena scale.
Formula
P(i beats j) = π_i / (π_i + π_j)
Range: (−∞, ∞)
Inputs and outputs
- comparisons: (model_a, model_b, outcome) rows
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.bradley_terry(comparisons, scale="elo")References
- Bradley RA, Terry ME. Rank analysis of incomplete block designs. Biometrika. 1952;39(3/4):324-345.
- Hunter DR. MM algorithms for generalized Bradley-Terry models. Ann Stat. 2004;32(1):384-406.
- Chiang WL, Zheng L, Sheng Y, et al. Chatbot Arena: an open platform for evaluating LLMs by human preference. ICML. 2024.