LLM-as-a-judge
Elo ratings
Implemented
llm-judge.elo_ratingsDefinition
Sequential Elo ratings from pairwise outcomes in the given order (expected score 1/(1 + 10^((R_b − R_a)/400)), update K·(outcome − expected)); order-dependent, unlike Bradley–Terry.
Formula
R_a ← R_a + K (S_a − E_a)
Range: (−∞, ∞), starting at the initial rating
Inputs and outputs
- comparisons: (model_a, model_b, outcome) rows
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.elo_ratings(comparisons)References
- Elo AE. The Rating of Chessplayers, Past and Present. Arco; 1978.
- Chiang WL, Zheng L, Sheng Y, et al. Chatbot Arena: an open platform for evaluating LLMs by human preference. ICML. 2024.