Semantic similarity
Model-based score (COMET, BLEURT, BARTScore, AlignScore, ...)
Implemented
semantic.model_scoreDefinition
Scores from any learned evaluation model or judge you supply, per example, so they get the same confidence intervals, model comparison and reports as every other metric.
Formula
mean_i scorer(reference_i, prediction_i[, source_i])
Range: that of the scorer
Inputs and outputs
- references: one reference string (or a list of references) per example
- predictions: one model output per example
- scorer: a callable returning one score per example (COMET, BLEURT, a judge)
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.model_score(references, predictions, scorer=my_comet_scorer)References
- Rei R, Stewart C, Farinha AC, Lavie A. COMET: a neural framework for MT evaluation. EMNLP. 2020.
- Sellam T, Das D, Parikh AP. BLEURT: learning robust metrics for text generation. ACL. 2020.