Skip to content
EvalSuite
Documentation menu

Semantic similarity

BERTScore

Implementedsemantic.bertscore

Definition

Greedy matching of contextual token embeddings by cosine similarity: precision averages, over prediction tokens, the best similarity to any reference token; recall does the reverse; F1 combines them. Optional IDF weights and baseline rescaling as in the original implementation.

Formula

P = Σ_j w_j max_i cos(r_i, p_j) / Σ_j w_j; R = Σ_i w_i max_j cos(r_i, p_j) / Σ_i w_i; F1 = 2PR/(P+R)

Range: [-1, 1] ([0, 1] in practice)

Inputs and outputs

  • reference_embeddings: embeddings from your encoder
  • prediction_embeddings: embeddings with the same dimension

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.bertscore(ref_token_embeddings, pred_token_embeddings)

References

  1. Zhang T, Kishore V, Wu F, Weinberger KQ, Artzi Y. BERTScore: evaluating text generation with BERT. ICLR. 2020.

Implementation status