Skip to content
EvalSuite
Documentation menu

LLM-as-a-judge and preferences

Available since v0.4.0

The judge is yours (a model, a panel or human raters). EvalSuite turns its outputs into ratings with intervals and checks how far the judge can be trusted.

Python
import evalsuite as es

comparisons = [("model_a", "model_b", "win"), ("model_b", "model_c", "win"),
             ("model_c", "model_a", "loss"), ("model_a", "model_c", "tie"), ("model_b", "model_a", "loss")]
es.win_rate(["win", "tie", "loss", "win"])          # ties count half; Wilson interval in params
es.bradley_terry(comparisons, scale="elo")          # MLE (same as choix), Chatbot Arena scale
es.elo_ratings(comparisons)                         # online Elo, order-dependent
es.plot.ratings(comparisons)                        # leaderboard with bootstrap intervals
es.plot.win_matrix(comparisons)

Agreement and judge reliability

Python
import numpy as np
ratings = np.array([[1, 2, 3, 3], [1, 2, 3, 4], [2, 2, 3, np.nan]])   # raters × items, NaN = missing
es.krippendorff_alpha(ratings, level="ordinal")
es.fleiss_kappa([["good", "good", "bad"], ["bad", "bad", "bad"], ["good", "bad", "good"]])
es.judge_agreement([1, 2, 3, 4, 5], [1, 2, 3, 5, 4], kind="ordinal")  # weighted kappa and Spearman

es.position_consistency(["A", "B", "A"], ["A", "A", "A"])   # verdicts with the order swapped
es.verbosity_bias(["A", "B", "A"], [120, 80, 300], [90, 100, 150])
es.self_preference_bias([True, True, False], [True, False, False])
es.rubric_score([[5, 4], [3, 4]], criteria=["helpfulness", "fluency"])

Always validate an automatic judge against human labels on a sample of your own data before relying on it, and report position and length controls.

Metrics

  • Maximum-likelihood strengths of a Bradley–Terry model fitted to pairwise preferences (ties as half a win for each side), by Hunter's MM algorithm; reported as log-strengths centred at 0 or on the Elo-like Chatbot Arena scale.

    es.bradley_terry

  • Elo ratingsImplemented

    Sequential Elo ratings from pairwise outcomes in the given order (expected score 1/(1 + 10^((R_b − R_a)/400)), update K·(outcome − expected)); order-dependent, unlike Bradley–Terry.

    es.elo_ratings

  • Fleiss' kappaImplemented

    Chance-corrected agreement of a fixed number of raters assigning items to nominal categories.

    es.fleiss_kappa

  • Agreement between an automatic judge and human ratings: Cohen's kappa for categories, quadratic-weighted kappa for ordinal scores (with Spearman's ρ), Pearson / Spearman correlation for continuous scores; raw agreement is reported alongside.

    es.judge_agreement

  • Chance-corrected agreement among any number of raters with missing ratings, for nominal, ordinal, interval or ratio data.

    es.krippendorff_alpha

  • Share of pairwise comparisons in which a system's response is preferred to the baseline; ties count half (or are excluded). A Wilson interval is reported.

    es.win_rate

  • Share of pairwise judgements that stay the same when the two responses are presented in swapped order; the inconsistent cases that favour whichever response was shown first measure position bias.

    es.position_consistency

  • Rubric scoreImplemented

    Mean of rubric ratings (correctness, helpfulness, relevance, coherence, fluency, completeness, clarity, conciseness, tone, instruction adherence, reasoning quality, ...) rescaled to [0, 1], with the mean per criterion.

    es.rubric_score

  • How much more often a judge prefers its own model's outputs than human raters do on the same comparisons.

    es.self_preference_bias

  • Verbosity biasImplemented

    Share of decisive judgements (no ties, different lengths) won by the longer response, with a two-sided binomial test against 0.5; well above 0.5 suggests a preference for length.

    es.verbosity_bias