Skip to content
EvalSuite
Documentation menu

LLM-as-a-judge

Verbosity bias

Implementedllm-judge.verbosity_bias

Definition

Share of decisive judgements (no ties, different lengths) won by the longer response, with a two-sided binomial test against 0.5; well above 0.5 suggests a preference for length.

Formula

wins of longer response / decisive comparisons

Range: [0, 1]

Inputs and outputs

  • winners: A / B / tie per comparison
  • length_a, length_b: response lengths

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.verbosity_bias(winners, length_a, length_b)

References

  1. Zheng L, Chiang WL, Sheng Y, et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks. 2023.

Implementation status