LLM-as-a-judge
Rubric score
Implemented
llm-judge.rubric_scoreDefinition
Mean of rubric ratings (correctness, helpfulness, relevance, coherence, fluency, completeness, clarity, conciseness, tone, instruction adherence, reasoning quality, ...) rescaled to [0, 1], with the mean per criterion.
Formula
mean((score − min) / (max − min))
Range: [0, 1]
Inputs and outputs
- scores: rubric ratings, (n, n_criteria)
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.rubric_score(rubric, criteria=["helpfulness", "fluency"])References
- Zheng L, Chiang WL, Sheng Y, et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks. 2023.