LLM-as-a-judge
Verbosity bias
Implemented
llm-judge.verbosity_biasDefinition
Share of decisive judgements (no ties, different lengths) won by the longer response, with a two-sided binomial test against 0.5; well above 0.5 suggests a preference for length.
Formula
wins of longer response / decisive comparisons
Range: [0, 1]
Inputs and outputs
- winners: A / B / tie per comparison
- length_a, length_b: response lengths
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.verbosity_bias(winners, length_a, length_b)References
- Zheng L, Chiang WL, Sheng Y, et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks. 2023.