LLM-as-a-judge
Self-preference bias
Implemented
llm-judge.self_preference_biasDefinition
How much more often a judge prefers its own model's outputs than human raters do on the same comparisons.
Formula
judge win rate of own outputs − human win rate of own outputs
Range: [-1, 1]
Inputs and outputs
- judge_prefers_own: judge preferred its own model's output
- human_prefers_own: humans preferred the same output
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.self_preference_bias(judge_prefers_own, human_prefers_own)References
- Panickssery A, Bowman SR, Feng S. LLM evaluators recognize and favor their own generations. NeurIPS. 2024.