Skip to content
EvalSuite
Documentation menu

LLM-as-a-judge

Rubric score

Implementedllm-judge.rubric_score

Definition

Mean of rubric ratings (correctness, helpfulness, relevance, coherence, fluency, completeness, clarity, conciseness, tone, instruction adherence, reasoning quality, ...) rescaled to [0, 1], with the mean per criterion.

Formula

mean((score − min) / (max − min))

Range: [0, 1]

Inputs and outputs

  • scores: rubric ratings, (n, n_criteria)

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.rubric_score(rubric, criteria=["helpfulness", "fluency"])

References

  1. Zheng L, Chiang WL, Sheng Y, et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks. 2023.

Implementation status