Skip to content
EvalSuite
Documentation menu

LLM-as-a-judge

Elo ratings

Implementedllm-judge.elo_ratings

Definition

Sequential Elo ratings from pairwise outcomes in the given order (expected score 1/(1 + 10^((R_b − R_a)/400)), update K·(outcome − expected)); order-dependent, unlike Bradley–Terry.

Formula

R_a ← R_a + K (S_a − E_a)

Range: (−∞, ∞), starting at the initial rating

Inputs and outputs

  • comparisons: (model_a, model_b, outcome) rows

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.elo_ratings(comparisons)

References

  1. Elo AE. The Rating of Chessplayers, Past and Present. Arco; 1978.
  2. Chiang WL, Zheng L, Sheng Y, et al. Chatbot Arena: an open platform for evaluating LLMs by human preference. ICML. 2024.

Implementation status