LLM-as-a-judge
Position consistency
Implemented
llm-judge.position_consistencyDefinition
Share of pairwise judgements that stay the same when the two responses are presented in swapped order; the inconsistent cases that favour whichever response was shown first measure position bias.
Formula
mean[verdict(A, B) = verdict(B, A)]
Range: [0, 1]
Inputs and outputs
- verdicts_original: winner with the original order
- verdicts_swapped: winner with the order swapped
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.position_consistency(verdicts_original, verdicts_swapped)References
- Zheng L, Chiang WL, Sheng Y, et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks. 2023.