Skip to content
EvalSuite
Documentation menu

LLM-as-a-judge

Bradley–Terry scores

Implementedllm-judge.bradley_terry

Definition

Maximum-likelihood strengths of a Bradley–Terry model fitted to pairwise preferences (ties as half a win for each side), by Hunter's MM algorithm; reported as log-strengths centred at 0 or on the Elo-like Chatbot Arena scale.

Formula

P(i beats j) = π_i / (π_i + π_j)

Range: (−∞, ∞)

Inputs and outputs

  • comparisons: (model_a, model_b, outcome) rows

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.bradley_terry(comparisons, scale="elo")

References

  1. Bradley RA, Terry ME. Rank analysis of incomplete block designs. Biometrika. 1952;39(3/4):324-345.
  2. Hunter DR. MM algorithms for generalized Bradley-Terry models. Ann Stat. 2004;32(1):384-406.
  3. Chiang WL, Zheng L, Sheng Y, et al. Chatbot Arena: an open platform for evaluating LLMs by human preference. ICML. 2024.

Implementation status