LLM-as-a-judge and preferences
Available since v0.4.0The judge is yours (a model, a panel or human raters). EvalSuite turns its outputs into ratings with intervals and checks how far the judge can be trusted.
import evalsuite as es
comparisons = [("model_a", "model_b", "win"), ("model_b", "model_c", "win"),
("model_c", "model_a", "loss"), ("model_a", "model_c", "tie"), ("model_b", "model_a", "loss")]
es.win_rate(["win", "tie", "loss", "win"]) # ties count half; Wilson interval in params
es.bradley_terry(comparisons, scale="elo") # MLE (same as choix), Chatbot Arena scale
es.elo_ratings(comparisons) # online Elo, order-dependent
es.plot.ratings(comparisons) # leaderboard with bootstrap intervals
es.plot.win_matrix(comparisons)Agreement and judge reliability
import numpy as np
ratings = np.array([[1, 2, 3, 3], [1, 2, 3, 4], [2, 2, 3, np.nan]]) # raters × items, NaN = missing
es.krippendorff_alpha(ratings, level="ordinal")
es.fleiss_kappa([["good", "good", "bad"], ["bad", "bad", "bad"], ["good", "bad", "good"]])
es.judge_agreement([1, 2, 3, 4, 5], [1, 2, 3, 5, 4], kind="ordinal") # weighted kappa and Spearman
es.position_consistency(["A", "B", "A"], ["A", "A", "A"]) # verdicts with the order swapped
es.verbosity_bias(["A", "B", "A"], [120, 80, 300], [90, 100, 150])
es.self_preference_bias([True, True, False], [True, False, False])
es.rubric_score([[5, 4], [3, 4]], criteria=["helpfulness", "fluency"])Always validate an automatic judge against human labels on a sample of your own data before relying on it, and report position and length controls.
Metrics
| Metric | Description | Status | API |
|---|---|---|---|
| Bradley–Terry scores | Maximum-likelihood strengths of a Bradley–Terry model fitted to pairwise preferences (ties as half a win for each side), by Hunter's MM algorithm; reported as log-strengths centred at 0 or on the Elo-like Chatbot Arena scale. | Implemented | es.bradley_terry |
| Elo ratings | Sequential Elo ratings from pairwise outcomes in the given order (expected score 1/(1 + 10^((R_b − R_a)/400)), update K·(outcome − expected)); order-dependent, unlike Bradley–Terry. | Implemented | es.elo_ratings |
| Fleiss' kappa | Chance-corrected agreement of a fixed number of raters assigning items to nominal categories. | Implemented | es.fleiss_kappa |
| Judge–human agreement | Agreement between an automatic judge and human ratings: Cohen's kappa for categories, quadratic-weighted kappa for ordinal scores (with Spearman's ρ), Pearson / Spearman correlation for continuous scores; raw agreement is reported alongside. | Implemented | es.judge_agreement |
| Krippendorff's alpha | Chance-corrected agreement among any number of raters with missing ratings, for nominal, ordinal, interval or ratio data. | Implemented | es.krippendorff_alpha |
| Pairwise win rate | Share of pairwise comparisons in which a system's response is preferred to the baseline; ties count half (or are excluded). A Wilson interval is reported. | Implemented | es.win_rate |
| Position consistency | Share of pairwise judgements that stay the same when the two responses are presented in swapped order; the inconsistent cases that favour whichever response was shown first measure position bias. | Implemented | es.position_consistency |
| Rubric score | Mean of rubric ratings (correctness, helpfulness, relevance, coherence, fluency, completeness, clarity, conciseness, tone, instruction adherence, reasoning quality, ...) rescaled to [0, 1], with the mean per criterion. | Implemented | es.rubric_score |
| Self-preference bias | How much more often a judge prefers its own model's outputs than human raters do on the same comparisons. | Implemented | es.self_preference_bias |
| Verbosity bias | Share of decisive judgements (no ties, different lengths) won by the longer response, with a two-sided binomial test against 0.5; well above 0.5 suggests a preference for length. | Implemented | es.verbosity_bias |
- Bradley–Terry scoresImplemented
Maximum-likelihood strengths of a Bradley–Terry model fitted to pairwise preferences (ties as half a win for each side), by Hunter's MM algorithm; reported as log-strengths centred at 0 or on the Elo-like Chatbot Arena scale.
es.bradley_terry
- Elo ratingsImplemented
Sequential Elo ratings from pairwise outcomes in the given order (expected score 1/(1 + 10^((R_b − R_a)/400)), update K·(outcome − expected)); order-dependent, unlike Bradley–Terry.
es.elo_ratings
- Fleiss' kappaImplemented
Chance-corrected agreement of a fixed number of raters assigning items to nominal categories.
es.fleiss_kappa
- Judge–human agreementImplemented
Agreement between an automatic judge and human ratings: Cohen's kappa for categories, quadratic-weighted kappa for ordinal scores (with Spearman's ρ), Pearson / Spearman correlation for continuous scores; raw agreement is reported alongside.
es.judge_agreement
- Krippendorff's alphaImplemented
Chance-corrected agreement among any number of raters with missing ratings, for nominal, ordinal, interval or ratio data.
es.krippendorff_alpha
- Pairwise win rateImplemented
Share of pairwise comparisons in which a system's response is preferred to the baseline; ties count half (or are excluded). A Wilson interval is reported.
es.win_rate
- Position consistencyImplemented
Share of pairwise judgements that stay the same when the two responses are presented in swapped order; the inconsistent cases that favour whichever response was shown first measure position bias.
es.position_consistency
- Rubric scoreImplemented
Mean of rubric ratings (correctness, helpfulness, relevance, coherence, fluency, completeness, clarity, conciseness, tone, instruction adherence, reasoning quality, ...) rescaled to [0, 1], with the mean per criterion.
es.rubric_score
- Self-preference biasImplemented
How much more often a judge prefers its own model's outputs than human raters do on the same comparisons.
es.self_preference_bias
- Verbosity biasImplemented
Share of decisive judgements (no ties, different lengths) won by the longer response, with a two-sided binomial test against 0.5; well above 0.5 suggests a preference for length.
es.verbosity_bias