Skip to content
EvalSuite
Documentation menu

LLM-as-a-judge

Self-preference bias

Implementedllm-judge.self_preference_bias

Definition

How much more often a judge prefers its own model's outputs than human raters do on the same comparisons.

Formula

judge win rate of own outputs − human win rate of own outputs

Range: [-1, 1]

Inputs and outputs

  • judge_prefers_own: judge preferred its own model's output
  • human_prefers_own: humans preferred the same output

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.self_preference_bias(judge_prefers_own, human_prefers_own)

References

  1. Panickssery A, Bowman SR, Feng S. LLM evaluators recognize and favor their own generations. NeurIPS. 2024.

Implementation status