Semantic similarity
BERTScore
Implemented
semantic.bertscoreDefinition
Greedy matching of contextual token embeddings by cosine similarity: precision averages, over prediction tokens, the best similarity to any reference token; recall does the reverse; F1 combines them. Optional IDF weights and baseline rescaling as in the original implementation.
Formula
P = Σ_j w_j max_i cos(r_i, p_j) / Σ_j w_j; R = Σ_i w_i max_j cos(r_i, p_j) / Σ_i w_i; F1 = 2PR/(P+R)
Range: [-1, 1] ([0, 1] in practice)
Inputs and outputs
- reference_embeddings: embeddings from your encoder
- prediction_embeddings: embeddings with the same dimension
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.bertscore(ref_token_embeddings, pred_token_embeddings)References
- Zhang T, Kishore V, Wu F, Weinberger KQ, Artzi Y. BERTScore: evaluating text generation with BERT. ICLR. 2020.