Skip to content
EvalSuite
Documentation menu

Semantic similarity

Model-based score (COMET, BLEURT, BARTScore, AlignScore, ...)

Implementedsemantic.model_score

Definition

Scores from any learned evaluation model or judge you supply, per example, so they get the same confidence intervals, model comparison and reports as every other metric.

Formula

mean_i scorer(reference_i, prediction_i[, source_i])

Range: that of the scorer

Inputs and outputs

  • references: one reference string (or a list of references) per example
  • predictions: one model output per example
  • scorer: a callable returning one score per example (COMET, BLEURT, a judge)

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.model_score(references, predictions, scorer=my_comet_scorer)

References

  1. Rei R, Stewart C, Farinha AC, Lavie A. COMET: a neural framework for MT evaluation. EMNLP. 2020.
  2. Sellam T, Das D, Parikh AP. BLEURT: learning robust metrics for text generation. ACL. 2020.

Implementation status